2019 Keynote)? Traditional Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults 5
2019 Keynote)? Traditional Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? 5
2019 Keynote)? Traditional Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched 5
2019 Keynote)? Traditional Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? 5
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? 6
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) 6
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. 6
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) 6
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles 6
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) 6
a visionary perspective… but what I know the best, which is: • Advances made by COINSE and collaborators recently :) • Acknowledgements: my students, Prof. Robert Feldt, R&D Division at Hyundai Motors Company Today’s Topic James Laughlin (1914-1997) 7
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) 8
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) Non Image Classifiers 8
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) SA & Cost Effectiveness Non Image Classifiers 8
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers 8
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (MODE@FSE’18, Apricot@ASE’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair 8
better Will help reduce the development cost of DNN based systems Should be possible to identify for various DNN models Better if possible to sample or search freely 9
better Will help reduce the development cost of DNN based systems Should be possible to identify for various DNN models Better if possible to sample or search freely Will be also helpful for correcting the unexpected behaviour 9
a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels than its own • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. 10
a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels than its own • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better 10
perturbation to the original image should not result in a vastly different execution of a DNN. • A small perturbation to the original image should not result in a vastly different execution of a DNN. • A small perturbation to the original image should result in a vastly different execution of a DNN. • See what I did there? :) The Continuity Assumption 5 5 five fave 13
perturbation to the original image should not result in a vastly different execution of a DNN. • A small perturbation to the original image should not result in a vastly different execution of a DNN. • A small perturbation to the original image should result in a vastly different execution of a DNN. • See what I did there? :) The Continuity Assumption 5 5 five fave 13
Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to identify for various DNN models 14
Selected test inputs based on DSA in CIFAR-10 Fig. 2: Accuracy of test inputs in MNIST and CIFAR-10 dataset, selected from the input with the lowest SA, increas- ingly including inputs with higher SA, and vice versa (i.e., from the input with the highest SA to inputs with lower SA). Fig. 4 and C If we feed more surprising inputs first, classification accuracy suffer, and vice versa. In other words, we can identify the inputs that induce unexpected behaviours pretty reliably. What can we do with this? J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. 16
Surprising Truth About What it Takes to Build a Machine Learning Product” Josh Logan (Tech Lead, Cloud AI group at Google) https://medium.com/thelaunchpad/the-ml-surprise-f54706361a6c 18
Surprising Truth About What it Takes to Build a Machine Learning Product” Josh Logan (Tech Lead, Cloud AI group at Google) https://medium.com/thelaunchpad/the-ml-surprise-f54706361a6c 18
Surprising Truth About What it Takes to Build a Machine Learning Product” Josh Logan (Tech Lead, Cloud AI group at Google) https://medium.com/thelaunchpad/the-ml-surprise-f54706361a6c 18
Study for Autonomous Driving Jinhan Kim, Jeongil Ju, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2006.00894 Surprise Adequacy Semantic Segmentation ? Surprise Adequacy ? We collaborated with R&D Division at Hyundai Motors Group to investigate two research questions. 19
see paper for details) SA and IoU (Intersection over Union) are negatively correlated: each dot shows average SA of class pixels as well as the class IoU. 21
of images with lowest SA, and accept the inferred segmentation as IoU = 1.0, 2) the actual IoU for a class will be y off from 1.0. 3) 30% saving at the cost of IoU inaccuracy of 0.03! 24
“bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems 25
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify.
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 4
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 5 4
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 5 4 Sorry, I’ve tricked all of you - they are all generated by machine. But I claim that some are more realistic than others.
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 5 4
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 5 4 C&W Pixel Level Optimisation SINVAD SINVAD FGSM
58054 (no longer active). 28 These are all inputs that cause a trained MNIST classifier misbehave. Which of the following MNIST inputs do you think have NOT been generated by a machine? Mark all that you think qualify. 1 2 3 5 4 C&W Pixel Level Optimisation SINVAD SINVAD FGSM
How do we make it more explorable? Occlusion Darken Different Weather Condition Seed Boundary of correct functional behaviour How can we more freely navigate this space? 29
How do we make it more explorable? Occlusion Darken Different Weather Condition Seed Boundary of correct functional behaviour How can we more freely navigate this space? Add Noise 29
How do we make it more explorable? Occlusion Darken Different Weather Condition Seed Boundary of correct functional behaviour How can we more freely navigate this space? Parameterised Model (Riccio & Tonella, 2020) Add Noise 29
How do we make it more explorable? Occlusion Darken Different Weather Condition Seed Boundary of correct functional behaviour How can we more freely navigate this space? Parameterised Model (Riccio & Tonella, 2020) Variational Autoencoder (Kang et al., 2020) Add Noise 29
How do we make it more explorable? Occlusion Darken Different Weather Condition Seed Boundary of correct functional behaviour How can we more freely navigate this space? Parameterised Model (Riccio & Tonella, 2020) Variational Autoencoder (Kang et al., 2020) Add Noise 29
Input Generation Sungmin Kang, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2005.09296 (SBST 2020) Encoder Decoder (a) VAE Raw t=0 0.2 0.4 0.6 0. Trajectory of AT through interpolation green: 4 red: 9 VAE Please watch the SBST presentation (https://www.youtube.com/watch?v=_psDl3wUh-4) and participate in online interactive session - 13:00 UTC, 2nd July 2020. Better if possible to sample or search freely 30
al., FSE 2018) uses GAN to augment training data with inputs similar to those that induce misbehaviour. • Apricot (Zhang and Chan, ASE 2019) trains multiple sub-DNNs, and adjust the weights of the original DNN w.r.t. the sub-DNNs. Data = $ Training = 33
Kang, and Shin Yoo https://arxiv.org/abs/1912.12463 Misbehaviour Localise Understand Patch Test Again Choose weights of neurons that 1) are most responsible for misbehaviour, and 2) are most influential to the outcome Arachne 36
Kang, and Shin Yoo https://arxiv.org/abs/1912.12463 Misbehaviour Localise Understand Patch Test Again Choose weights of neurons that 1) are most responsible for misbehaviour, and 2) are most influential to the outcome Use both positive and negative test cases, as in GenProg Arachne 36
pair (top plot), or multiple pairs (bottom plot). • The readjustment generalises to test data! • When repairing a single “bug” (misclassification label pair), there is relatively minor impact on other labels. Direct Repair seems to work. 3FQBJSJOHUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSPG$*'"3 3FQBJSJOHNVMUJQMFUZQFTPGNJTCFIBWJPVSPG$*'"3 37
more widely impactful and repairs more instances • I still believe there may be a nice use case for a direct repair like this… Direct repair can be more precise. TBNQMFEGSPNUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSJO$*'"3 Ineg "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSUIFSFQBJSCZ"SBDIOF "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSSFUSBJOJOH 38
more widely impactful and repairs more instances • I still believe there may be a nice use case for a direct repair like this… Direct repair can be more precise. TBNQMFEGSPNUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSJO$*'"3 Ineg "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSUIFSFQBJSCZ"SBDIOF "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSSFUSBJOJOH Will be also helpful for correcting the unexpected behaviour 38
Should be possible to identify for various DNN models Better if possible to sample or search freely Will be also helpful for correcting the unexpected behaviour 39
Should be possible to identify for various DNN models Better if possible to sample or search freely Will be also helpful for correcting the unexpected behaviour What next? 39
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models Classification Tradeoff If only images with IoU < 0.5 are “bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models Classification Tradeoff If only images with IoU < 0.5 are “bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation Sungmin Kang, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2005.09296 (SBST 2020) Encoder Decoder (b) (a) VAE Raw t=0 0.2 0.4 0.6 0.8 1.0 15 10 5 0 5 10 15 10 5 0 Trajectory of AT through interpolation green: 4 red: 9 VAE Raw Please watch the SBST presentation (https://www.youtube.com/watch?v=_psDl3wUh-4) and participate in online interactive session - 13:00 UTC, 2nd July 2020. Better if possible to sample or search freely [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models Classification Tradeoff If only images with IoU < 0.5 are “bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation Sungmin Kang, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2005.09296 (SBST 2020) Encoder Decoder (b) (a) VAE Raw t=0 0.2 0.4 0.6 0.8 1.0 15 10 5 0 5 10 15 10 5 0 Trajectory of AT through interpolation green: 4 red: 9 VAE Raw Please watch the SBST presentation (https://www.youtube.com/watch?v=_psDl3wUh-4) and participate in online interactive session - 13:00 UTC, 2nd July 2020. Better if possible to sample or search freely • Retraining is more disruptive $ • However, it is also more widely impactful and repairs more instances % • I still believe there may be a nice use case for a direct repair like this… Direct repair can be more precise. TBNQMFEGSPNUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSJO$*'"3 I neg "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSUIFSFQBJSCZ"SBDIOF "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSSFUSBJOJOH Will be also helpful for correcting the unexpected behaviour [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models Classification Tradeoff If only images with IoU < 0.5 are “bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation Sungmin Kang, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2005.09296 (SBST 2020) Encoder Decoder (b) (a) VAE Raw t=0 0.2 0.4 0.6 0.8 1.0 15 10 5 0 5 10 15 10 5 0 Trajectory of AT through interpolation green: 4 red: 9 VAE Raw Please watch the SBST presentation (https://www.youtube.com/watch?v=_psDl3wUh-4) and participate in online interactive session - 13:00 UTC, 2nd July 2020. Better if possible to sample or search freely • Retraining is more disruptive $ • However, it is also more widely impactful and repairs more instances % • I still believe there may be a nice use case for a direct repair like this… Direct repair can be more precise. TBNQMFEGSPNUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSJO$*'"3 I neg "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSUIFSFQBJSCZ"SBDIOF "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSSFUSBJOJOH Will be also helpful for correcting the unexpected behaviour Road Ahead Random thoughts on what we need… We need more systematic guidelines for retraining. We should start thinking about benchmarks now. Diversify subjects beyond image classifiers. [email protected] https://coinse.io 42
Code DL System Specification Training Data Logic as Control Flow Logic as Data Flow Written Trained Tested Tested For Faults Faults? Patched Retrained? Verification (abstract interpretation, interval analysis, etc…) DNN models are getting bigger and more complicated. Test Adequacy (DeepXplore, DeepGauge, SA…) Input Generation (GAN, AEs, simulation, model-based) Taxonomy of faults in deep learning systems Metamorphic Oracles Systematic Retraining (Apricot - ASE ’19) SA & Cost Effectiveness VAE based SBST Non Image Classifiers Direct Repair Our guide to cost effectiveness • Intuitively speaking, SA is a distance metric that measures out-of-distribution- ness. • An input is more OOD if: • A similar neural activation has been rarely seen • its neural activation pattern is far from the observed mean • its neural activation is closer to inputs in other class labels • More OOD = More likely to misbehave Surprise Adequacy1 1. J. Kim, R. Feldt, and S. Yoo. Guiding deep learning system testing using surprise adequacy. In Proceedings of the 41th International Conference on Software Engineering, ICSE 2019, pages 1039–1049. IEEE Press, 2019. Will reveal unexpected behaviours better Evaluating Surprise Adequacy on Question Answering Seah Kim and Shin Yoo DeepTest 2020 We apply SA analysis to Question Answering task. Stay tuned for the talk, which is right after this keynote :) Should be possible to generate for various DNN models Classification Tradeoff If only images with IoU < 0.5 are “bugs”… Not labelling 55% of images, and we still lose 5% of images whose Road Marker IoUs are less than 0.5. Will help reduce the development cost of DNN based systems SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation Sungmin Kang, Robert Feldt, and Shin Yoo https://arxiv.org/abs/2005.09296 (SBST 2020) Encoder Decoder (b) (a) VAE Raw t=0 0.2 0.4 0.6 0.8 1.0 15 10 5 0 5 10 15 10 5 0 Trajectory of AT through interpolation green: 4 red: 9 VAE Raw Please watch the SBST presentation (https://www.youtube.com/watch?v=_psDl3wUh-4) and participate in online interactive session - 13:00 UTC, 2nd July 2020. Better if possible to sample or search freely • Retraining is more disruptive $ • However, it is also more widely impactful and repairs more instances % • I still believe there may be a nice use case for a direct repair like this… Direct repair can be more precise. TBNQMFEGSPNUIFNPTUGSFRVFOUUZQF PGNJTCFIBWJPVSJO$*'"3 I neg "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSUIFSFQBJSCZ"SBDIOF "WHPGQBUDIFEBOECSPLFOJOQVUTQFSMBCFMBGUFSSFUSBJOJOH Will be also helpful for correcting the unexpected behaviour Road Ahead Random thoughts on what we need… We need more systematic guidelines for retraining. We should start thinking about benchmarks now. Diversify subjects beyond image classifiers. Recommendations for Going Forward …on a higher, strategic level Contextualise firmly in SE practice. Pay close attention to ML literature. Take replicability and transparency seriously [email protected] https://coinse.io 42