subreddits (r/Conservative, r/Liberal, etc.) ◦ Midterm Election (November 2018) ◉ Russian Twitter propaganda ◦ 3 million tweets posted by Russian trolls ◦ Labeled tweet content ▪ type of bot (including left and right leaning) ▪ 20 other features not used ◦ Presidential Election ◉ “Normal” tweets ◦ 5k tweets collected by Sanders Analytics Propaganda Tweets Reddit Comments “Normal” Tweets Propaganda dataset Political bias dataset
◦ 10,000 tweets from labeled Russian Twitter propaganda ◉ Two fields ◦ Text ◦ Political leaning (right or left) 4 Propaganda ◉ 10,000 records ◦ 5000 tweets from Russian Twitter propaganda ◦ 5000 tweets from “normal” tweet dataset ◉ Two fields ◦ Tweet ◦ “bot” or “not-bot”
with empty content (NAs) ◉ Replace non-ASCII and non-english characters with spaces ◉ Ensure balance of class labels (base rate of 50%) ◉ Word stemming (Porter algorithm) ◉ Term-frequency inverse document frequency (TF-IDF) ◉ Remove features (words) mentioned very infrequently ◉ Reduced features by about 90% ◉ Significant model performance increase