Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Training data selection for cross-project defec...
Search
PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
October 09, 2013
Research
160
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Training data selection for cross-project defect prediction
by Steffen Herbold
PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
October 09, 2013
More Decks by PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
See All by PROMISE'13: The 9th International Conference on Predictive Models in Software Engineering
Are Comprehensive Quality Models Necessary for Evaluating Software Quality?
promise
0
150
Using Evidential Reasoning to Make Qualified Predictions of Software Quality
promise
0
200
Using code change types in an analogy-based classifier for short-term defect prediction
promise
0
120
The Impact of Parameter Tuning on Software Effort Estimation Using Learning Machines
promise
0
160
Incremental Development Productivity Decline
promise
1
160
An Algorithmic Approach to Missing Data Problem in Modeling Human Aspects in Software Development
promise
0
130
An Analysis of Multi-objective Evolutionary Algorithms for Training Ensemble Models Based on Different Performance Measures in Software Effort Estimation
promise
0
150
Beyond Data Mining; Towards “Idea Engineering”
promise
0
120
A Comparative Evaluation of Static Analysis Actionable Alert Identification Techniques
promise
0
160
Other Decks in Research
See All in Research
クラウド・AI 時代の研究開発 DX / R&D Digital Transformation
hariby
0
130
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
satai
3
260
JPA2026_NetworkTutorial_JunKashihara
junkashihara
0
120
シングルチャネルマルチトーカー音声認識の進展
ryomasumura
0
250
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
200
セマンティック通信勉強会 6Gに向けたデバイス間効率的な通信の技術紹介・課題・今後展望
satai
3
320
Cross-Media Information Spaces and Architectures
signer
PRO
0
360
Language and AI
ayaniwa
0
220
NLP colloquium: AI Safety Survey
kanekomasahiro
1
1.1k
東京大学工学部計数工学科、計数工学特別講義の説明資料
kikuzo
0
650
JICA QUEST 共創×革新プログラム Impact Report(海ノ向こうコーヒー)
ontheslope
0
530
MIRU2026 チュートリアル講演2:三次元データ処理の動向
nnchiba
6
4.7k
Featured
See All Featured
The MySQL Ecosystem @ GitHub 2015
samlambert
251
13k
Data-driven link building: lessons from a $708K investment (BrightonSEO talk)
szymonslowik
1
1.3k
Neural Spatial Audio Processing for Sound Field Analysis and Control
skoyamalab
0
460
Sam Torres - BigQuery for SEOs
techseoconnect
PRO
0
530
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
190
brightonSEO & MeasureFest 2025 - Christian Goodrich - Winning strategies for Black Friday CRO & PPC
cargoodrich
3
800
Jamie Indigo - Trashchat’s Guide to Black Boxes: Technical SEO Tactics for LLMs
techseoconnect
PRO
0
660
Crafting Experiences
bethany
1
280
Fireside Chat
paigeccino
42
4k
Agile Actions for Facilitating Distributed Teams - ADO2019
mkilby
0
270
Into the Great Unknown - MozCon
thekraken
41
2.7k
How To Stay Up To Date on Web Technology
chriscoyier
790
250k
Transcript
Training data selection for cross- project defect prediction Steffen Herbold
1
Outline • Motivation • Training data selection • Case study
• Conclusion 2
3 Every so)ware has failures!
But where? 4
Defect prediction! 5
Software metrics as foundation public class GUIElementFactory implements IGUIElementFactory {
private static GUIElementFactory instance = new GUIElementFactory(); private GUIElementFactory() {} public static synchronized GUIElementFactory getInstance() { return instance; } private Properties mappingsFromConfiguration; @Override public IGUIElement instantiateGUIElement( IGUIElementSpec specification, IGUIElement parent) throws GUIModelConfigurationException { Properties mappings = getMappingsFromConfiguration(); IGUIElement guiElement = null; String[] typeHierarchy = specification.getTypeHierarchy(); int i = 0; String className = null; while ((className == null) && (i < typeHierarchy.length)) { className = mappings.getProperty(typeHierarchy[i]); i++; } if (className != null) { try { GUIElementFactory Lines of Code (LOC) 193 Weighted Methods per Class (WMC) 34 Number of Methods (NOM) 3 … 6 …
Defect prediction 7 Predictor Training Defect Prediction Target Project
Ant 1.3
Cross-project defect prediction 8 Predictor Training Defect Prediction Target
Project Ant 1.3 Available Data arc Xerces 1.4 ...
Training data as subset of available data 9 Training Data
Selection Predictor Training Defect Prediction Target Project Ant 1.3 Training Data Available Data arc Xerces 1.4 ... Based on
Set-wise selection 10 Training Data Selection Predictor Training Defect
Prediction Target Project Ant 1.3 Training Data Version 1 Version k ... Available Data arc Xerces 1.4 ... Based on
Relationship between distributional characteristics and success 11
Distributional characteristics of a project Project Characteris7cs mean(LOC)
110 stddev(LOC) 30 … mean(WMC) 15 stddev(WMC) 5 … 12 … GUIElementFactory Lines of Code (LOC) 193 Weighted Methods per Class (WMC) 34 … GUIElement Lines of Code (LOC) 75 Weighted Methods per Class (WMC) 10 … Project Data
k-Nearest Neighbor Selection 13
k-Nearest Neighbor Selection 14
EM clustering selection 15
EM clustering selection 16
Case study data 17 • 14 Java projects • 44
releases • 20 software metrics • 34% percent defect prone in total
Defect proneness density 18
Weighting to counter bias 19 ∑↑▒↓ =∑↑▒ ↓ ↓ =0.5∙#/# ↓ =0.5∙#/#
Predictor models • Logistic Regression • Naïve Bayes • Bayesian
Networks • SVM with RBF kernel • C4.5 Decision Trees • Random Forest • Multilayer Perceptron 20
Case study workflow 21 Training Data Selection Predictor Training
Defect Prediction Target Project Ant 1.3 Training Data Version 1 Version k ... Candidate Training Data arc Xerces 1.4 ... Available Projects Ant 1.3 Ant 1.7 ... arc Xerces 1.4 ... Based on For each predictor model For each available project
Evaluation criteria 22 =/+ =/+ =(>0.7) ∧(>0.5)
Results success 23
recall and precision 24
Key findings • Set-wise selection improves results • Equal weighting
improves results • Within-project performance still out of reach • SVM performs best 25
Open Issues 26
Advertisement • http://autoquest.informatik.uni-goettingen.de 27