individual and that is not public information: ➤ Membership to a private community ➤ Medical data ➤ Earnings, savings, financial info ➤ Political beliefs, voting choices, ... ➤ Personal habits ➤ Religion ➤ ... Those are meaningless if not linked to someone in particular.
personnal data 2. By working on an anonymized dataset - Swapping personal information for ids: pseudonymisation - Aggregating pseudo-identifiers using k-anonymity - Using differential privacy
House Magical Disease 1 15 M Slytherin Dragon pox 2 19 M Hufflepuff Black cat flu 3 12 F Griffindor Levitation sickness 4 18 F Slytherin Petrification 5 14 M Griffindor Hippogriff bite 6 14 M Griffindor Dragon pox 7 19 M Ravenclaw Black cat flu 8 13 F Ravenclaw Levitation sickness 9 17 F Slytherin Lycanthropy 10 15 M Griffindor Hippogriff bite What can you tell me about Harry, a 15-year-old Griffindor boy ?
some attributes, so that they cannot be used as pseudo-identifiers. ➤ Even in large datasets, some combinations of gender, age and zipcode are often unique. ➤ Eg: Use ranges for age (20-30, 30-40) instead of the value. It's probably enough for your analysis. ➤ Eg: Use larger region rather than zipcodes, or remove zipcodes for less populated areas.
Age Gender House Magical Disease 1 15-20 M Slyth./Griff. Dragon pox 2 15-20 M Huff./Rav. Black cat flu 3 10-14 F Griff./Rav. Levitation sickness 4 15-20 F Slyth./Griff. Petrification 5 10-14 M Slyth./Griff. Hippogriff bite 6 10-14 M Slyth./Griff. Dragon pox 7 15-20 M Huff./Rav. Black cat flu 8 10-14 F Griff./Rav. Levitation sickness 9 15-20 F Slyth./Griff. Lycanthropy 10 15-20 M Slyth./Griff. Hippogriff bite What can you tell me about Luna, a 14-year-old Ravenclaw girl ?
we cannot know for sure that they are in the dataset. ➤ However, it is possible that all k individuals in the same group share the same value for a protected attribute. ➤ l-diversity ensures that each k-anonymous group contains at least l different values of a sensitive attribute
House Magical Disease 3 10-14 F \ Levitation sickness 5 10-14 M \ Hippogriff bite 6 10-14 M \ Dragon pox 8 10-14 F \ Levitation sickness 12 10-14 M \ Common cold 18 10-14 F \ Levitation sickness 21 10-14 M \ Black cat flu 25 10-14 F \ Levitation sickness 28 10-14 F \ Dragon pox 30 10-14 M \ Hippogriff bite What can you tell me about Luna, a 14-year-old girl ?
➤ If 90% of the persons in a l-diverse group have the same value for a protected attribute, then we can infer with high-probability that a person in this group will have that value. ➤ t-closeness ensures that, in each group, the distribution with respect to a sensitive attribute does not differ "too much" from the overall distribution.
Dataset size: 72 people Members of the army: 17 people Thursday: Dataset size: 73 people Members of the army: 18 people What can you tell me about Ginny, who was added on Wednesday night ?
student membership to Dumbledore's army. ➤ It's a highly sensitive information, we cannot just store the true value. ➤ So let's sometimes lie about it Should I store the truth? p 1-p No: what should I store? k 1-k Yes, the truth Yes No Is this person member of Dumbledore's Army ?
can't draw any definitive conclusion just by watching the dataset grow over time ➤ If probabilities p and k are known, we can adjust our estimators to get an unbiaised estimate of aggregated values, with a limited loss of precision
even when the graph looks fine (think JS-based graphs, or SVGs) ➤ Some ML models can leak private information, esp. on minority classes: consider privacy-preserving Machine Learning methods ➤ Sometimes, even just knowing someone is part of a dataset is a privacy leak
levitating lately ? ID Age Gender House OWL Magical Disease 3 10-14 F \ B Levitation sickness 5 10-14 M \ C Hippogriff bite 6 10-14 M \ B- Dragon pox 8 10-14 F \ A- Levitation sickness 12 10-14 M \ B+ Common cold 18 10-14 F \ F Levitation sickness 21 10-14 M \ C+ Black cat flu 25 10-14 F \ A+++ Levitation sickness 28 10-14 F \ A+ Dragon pox 30 10-14 M \ E Hippogriff bite
Andreas Dewes and Katharine Jarmul: https://github.com/KIProtect/data-privacy-for-data-scientists ➤ k-anonymity: a model for protecting privacy, Latanya Sweeney ➤ Mondrian Multidimensional k-Anonymity, K. LeFevre, D.DeWitt, and R. Ramakrishnan ➤ Differential Privacy, Cynthia Dwork ➤ Encrypted statistical machine learning: new privacy preserving methods, L. Aslett, P . Esperança, and C. Holmes ➤ Communication-Efficient Learning of Deep Networks from Decentralized Data, H. Brendan McMahan, E. Moore, D. Ramage, S. Hampson, B. Agüera y Arcas