Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Measuring Quality Content
Search
Adam Hyland
August 04, 2012
Research
94
2
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Measuring Quality Content
Presentation to Wikimania 2012 on Article Feedback Tool statistics.
Adam Hyland
August 04, 2012
More Decks by Adam Hyland
See All by Adam Hyland
Here Comes (a significant fraction of) Everybody
protonk
0
83
Boston Data Swap: Data Vis Under Uncertainty
protonk
0
60
Why Nate Silver is Famous
protonk
1
130
Data Visualization under Uncertainty
protonk
0
780
Phillips Academy Wikipedia Introduction
protonk
0
94
Other Decks in Research
See All in Research
Cross-Media Human-Information Interaction
signer
PRO
0
190
大規模言語モデルは誰を覚えているか / Who Do Large Language Models Memorize?
upura
0
140
【Zozo Research 技術共有会】三次元領域の現在と展望
mickey_0226
3
570
Visual SLAM未来予測 / Future Prediction in Visual SLAM
koide3
1
920
2026年版中小企業白書・小規模企業白書の概要
ozekinote
0
180
マーケットストリート 社会実験2024 in 秋葉原ジャンク通り 調査報告書
izumiyama_lab
1
120
LLM Compute Infrastructure Overview
karakurist
2
1.6k
2026年 オープンキャンパス 研究室紹介
junkurihara
0
140
PHTalks Bengaluru - SSRF When All Else Fails
dk999
0
1.1k
超効率化への挑戦:1bit LLMの現状と展望
yumaichikawa
0
550
VLMの推論を高速化する視覚トークン削減の仕組み
tattaka
2
270
nlp2026 In-Context Learningに基づく経路案内のための地理的知識の活用方法に関する検討
takashiinui
0
130
Featured
See All Featured
Docker and Python
trallard
47
4.1k
Optimising Largest Contentful Paint
csswizardry
37
3.9k
Navigating Team Friction
lara
192
16k
KATA
mclloyd
PRO
35
15k
The Hidden Cost of Media on the Web [PixelPalooza 2025]
tammyeverts
2
480
Digital Ethics as a Driver of Design Innovation
axbom
PRO
1
380
Tell your own story through comics
letsgokoyo
1
1k
Paper Plane (Part 1)
katiecoart
PRO
1
10k
The Cost Of JavaScript in 2023
addyosmani
55
10k
How to build a perfect <img>
jonoalderson
1
5.9k
How to optimise 3,500 product descriptions for ecommerce in one day using ChatGPT
katarinadahlin
PRO
2
3.8k
Why Your Marketing Sucks and What You Can Do About It - Sophie Logan
marketingsoph
0
390
Transcript
Measuring Article Quality Peer Review and the Article Feedback Tool
Adam Hyland protonk @ en-wp
Look Familiar?
Maybe This Version?
None
Article Feedback Tool • Deployed in 2010 • Version 4
(the current version) ramped up in 2011 • Designed to offer an avenue for reader feedback • High volume of reader feedback
• 6 months of public data • 795,353 articles --
2,487,522 responses
Featured Articles (FA) • 3,599 articles (0.09% of all articles)
• 2,267 Featured Lists (FL) • Most rigorous peer review process on the English Wikipedia • Very sensitive to editor preferences • Some idiosyncrasies
Good Articles (GA) • 15,357 articles • Relatively rigorous peer
review (yes I know reasonable minds may disagree) • Less idiosyncratic than FA in some ways • Perhaps less dependent on editor preference
Data • Article name • Length (in bytes) • GA/FA
status (including former/not- promoted) • Some user data
None
Beyond Summaries • Reader ratings follow pageviews • Predominantly non-editors
• Popular articles: • Call of Duty • Justin Bieber • Jimmy Wales (avg. rating: 1.10585)
Power Laws Everywhere!
Classical(ish) Models • Logistic regression model supports a relationship between
rating and likelihood of FA/GA • Linear model does, but with a twist • Can’t escape Cambridge Endogeneity Police!
None
Data Mining • Predicting featured status from reader ratings and
minimal meta-data. • Bayesian classifier able to roughly predict featured status (with a high false positive rate)
But the system’s changing! • AFT v4 is a multi-category
quantitative measure • AFT v5 is, roughly, YES/NO • Is this a problem? • Frank Harrell and the perils of dichotomization.
Actual Reader Ratings
Another Look
For the skeptics
Information • We can imagine we might not lose information
in shifting to v5 • This is born out by the classifier, to some degree. • We don’t lose a lot of power when dichotomizing individual ratings
A Look Ahead • Really exciting! • Great compliment to
current research methods • Long exposures can help discover reader/editor divergence • Predictive analytics • Need more open data
Questions? • Of course you have questions! • All work
is or soon will be available on github under a free license • Full writeup on en-wp forthcoming