Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Training data selection for cross-project defec...
Search
PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
October 09, 2013
Research
160
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Training data selection for cross-project defect prediction
by Steffen Herbold
PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
October 09, 2013
More Decks by PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
See All by PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
Are Comprehensive Quality Models Necessary for Evaluating Software Quality?
promise
0
150
Using Evidential Reasoning to Make Qualified Predictions of Software Quality
promise
0
200
Using code change types in an analogy-based classifier for short-term defect prediction
promise
0
120
The Impact of Parameter Tuning on Software Effort Estimation Using Learning Machines
promise
0
160
Incremental Development Productivity Decline
promise
1
160
An Algorithmic Approach to Missing Data Problem in Modeling Human Aspects in Software Development
promise
0
130
An Analysis of Multi-objective Evolutionary Algorithms for Training Ensemble Models Based on Different Performance Measures in Software Effort Estimation
promise
0
150
Beyond Data Mining; Towards “Idea Engineering”
promise
0
120
A Comparative Evaluation of Static Analysis Actionable Alert Identification Techniques
promise
0
160
Other Decks in Research
See All in Research
Fukui Shibiten 39 - AI Art
butchi
0
190
最先端NLP 2026 論文紹介: Wait, Wait, Wait... Why Do Reasoning Models Loop? / SNLP Paper Review: Wait, Wait, Wait... Why Do Reasoning Models Loop?
tkng
0
210
HackSick vol.7 LT資料【LLMアーキテクチャ入門・事前学習時の躓き所解説】 スパースなAttention・状態空間モデル
rikkabotan7
0
170
Using our influence and power for patient safety
helenbevan
0
410
PGDM: Physically Guided Diffusion Model for L Downscaling
satai
3
460
CVPR2026論文紹介_VLMにとって良いvision encoderとは何か?Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
kobayashi31
1
210
2025年度秋葉原ウォーカブルプロジェクト調査報告 「アキバらしいウォーカブル」とは何か
izumiyama_lab
1
210
適応的スパムフィルタのための軽量な類似メッセージカウンタ / jsai2026-adaptive-spam-filter
monochromegane
0
5.3k
Data Visualization Tools in the Age of AI
flekschas
0
200
セマンティック通信勉強会 6Gに向けたデバイス間効率的な通信の技術紹介・課題・今後展望
satai
3
320
J-STAGEの現況と全文XML登載必須化について
xspa2012
0
200
Model Discovery and Graph Simulation: A Lightweight Gateway to Chaos Engineering
anatolykr
0
290
Featured
See All Featured
Building Flexible Design Systems
yeseniaperezcruz
330
41k
Hiding What from Whom? A Critical Review of the History of Programming languages for Music
tomoyanonymous
3
1.2k
The Impact of AI in SEO - AI Overviews June 2024 Edition
aleyda
6
1.2k
Jess Joyce - The Pitfalls of Following Frameworks
techseoconnect
PRO
1
410
CSS Pre-Processors: Stylus, Less & Sass
bermonpainter
360
30k
The Illustrated Guide to Node.js - THAT Conference 2024
reverentgeek
1
470
Measuring Dark Social's Impact On Conversion and Attribution
stephenakadiri
2
270
Visualization
eitanlees
152
17k
For a Future-Friendly Web
brad_frost
183
10k
Responsive Adventures: Dirty Tricks From The Dark Corners of Front-End
smashingmag
254
22k
Faster Mobile Websites
deanohume
310
32k
Agile Leadership in an Agile Organization
kimpetersen
PRO
0
220
Transcript
Training data selection for cross- project defect prediction Steffen Herbold
1
Outline • Motivation • Training data selection • Case study
• Conclusion 2
3 Every so)ware has failures!
But where? 4
Defect prediction! 5
Software metrics as foundation public class GUIElementFactory implements IGUIElementFactory {
private static GUIElementFactory instance = new GUIElementFactory(); private GUIElementFactory() {} public static synchronized GUIElementFactory getInstance() { return instance; } private Properties mappingsFromConfiguration; @Override public IGUIElement instantiateGUIElement( IGUIElementSpec specification, IGUIElement parent) throws GUIModelConfigurationException { Properties mappings = getMappingsFromConfiguration(); IGUIElement guiElement = null; String[] typeHierarchy = specification.getTypeHierarchy(); int i = 0; String className = null; while ((className == null) && (i < typeHierarchy.length)) { className = mappings.getProperty(typeHierarchy[i]); i++; } if (className != null) { try { GUIElementFactory Lines of Code (LOC) 193 Weighted Methods per Class (WMC) 34 Number of Methods (NOM) 3 … 6 …
Defect prediction 7 Predictor Training Defect Prediction Target Project
Ant 1.3
Cross-project defect prediction 8 Predictor Training Defect Prediction Target
Project Ant 1.3 Available Data arc Xerces 1.4 ...
Training data as subset of available data 9 Training Data
Selection Predictor Training Defect Prediction Target Project Ant 1.3 Training Data Available Data arc Xerces 1.4 ... Based on
Set-wise selection 10 Training Data Selection Predictor Training Defect
Prediction Target Project Ant 1.3 Training Data Version 1 Version k ... Available Data arc Xerces 1.4 ... Based on
Relationship between distributional characteristics and success 11
Distributional characteristics of a project Project Characteris7cs mean(LOC)
110 stddev(LOC) 30 … mean(WMC) 15 stddev(WMC) 5 … 12 … GUIElementFactory Lines of Code (LOC) 193 Weighted Methods per Class (WMC) 34 … GUIElement Lines of Code (LOC) 75 Weighted Methods per Class (WMC) 10 … Project Data
k-Nearest Neighbor Selection 13
k-Nearest Neighbor Selection 14
EM clustering selection 15
EM clustering selection 16
Case study data 17 • 14 Java projects • 44
releases • 20 software metrics • 34% percent defect prone in total
Defect proneness density 18
Weighting to counter bias 19 ∑↑▒↓ =∑↑▒ ↓ ↓ =0.5∙#/# ↓ =0.5∙#/#
Predictor models • Logistic Regression • Naïve Bayes • Bayesian
Networks • SVM with RBF kernel • C4.5 Decision Trees • Random Forest • Multilayer Perceptron 20
Case study workflow 21 Training Data Selection Predictor Training
Defect Prediction Target Project Ant 1.3 Training Data Version 1 Version k ... Candidate Training Data arc Xerces 1.4 ... Available Projects Ant 1.3 Ant 1.7 ... arc Xerces 1.4 ... Based on For each predictor model For each available project
Evaluation criteria 22 =/+ =/+ =(>0.7) ∧(>0.5)
Results success 23
recall and precision 24
Key findings • Set-wise selection improves results • Equal weighting
improves results • Within-project performance still out of reach • SVM performs best 25
Open Issues 26
Advertisement • http://autoquest.informatik.uni-goettingen.de 27