← All posts

My first text classifier for sentiment analysis

Naive Bayes, a linear SVM and a CNN on 1.6 million tweets, and why the data mattered more than the model.

In spring 2019 I did a course project at KTH with two classmates. It was my first project on text. The question was simple: which classifier is best at telling positive tweets from negative ones? We tried Naive Bayes, a linear SVM and a small convolutional network. I expected the CNN to win, because in 2019 the neural network was supposed to win.

It didn’t. A year and a half later I reread the notebook, and I think the more useful lesson is somewhere other than where we put it at the time.

The data

We used Sentiment140, 1.6 million English tweets from 2009, each labelled positive or negative. The first notebook was exploration: which words show up on each side.

Word cloud of negative tweets: today, work, still, miss, sad, bad, lol, now.

negative tweets

Word cloud of positive tweets: love, today, thank, good, well, lol, awesome, now.

positive tweets

Most frequent words in each class, from our exploration notebook. "Today", "now" and "lol" are large on both sides.

The clouds already hint at the difficulty. The biggest words are the same on both sides. Sentiment sits in the smaller words and in how they combine.

From a tweet to numbers

A model can’t read a tweet; it needs a row of numbers. Everything before the model is about producing that row. Here is the whole pipeline at a glance:

flowchart TB
  T[Raw tweets] --> C[Clean]
  C --> N[N-grams]
  N --> B[Counts or TF-IDF]
  B --> NB[Naive Bayes]
  B --> SVM[Linear SVM]
  C --> E[Word indices]
  E --> CNN[CNN]
  NB --> V[Accuracy on 1M held-out tweets]
  SVM --> V
  CNN --> V

Cleaning. Tweets are messy: mentions, links, HTML entities, inconsistent case, contractions. Our cleaning function handled them in a fixed order:

Cleaning one tweet, step by step A raw tweet with a mention, an HTML entity and a link is cleaned in four steps into the tokens: do, not, love, this, song, anymore, it, hurts. raw tweet · labelled negative from a :( that Sentiment140 already stripped @jess_k I don't love this song anymore & it hurts http://bit.ly/x9 1 · decode HTML, remove mentions and links I don't love this song anymore & it hurts 2 · lowercase, expand contractions i do not love this song anymore & it hurts 3 · keep letters only, drop one-letter words do not love this song anymore it hurts 4 · split into tokens do not love this song anymore it hurts
One made-up tweet through our cleaning function. Orange marks noise that a later step removes; blue marks what a step changed. Expanding "don't" to "do not" was deliberate: it keeps the negation as its own word.

N-grams. Next, each cleaned tweet is split into n-grams: single words (unigrams), pairs (bigrams) and triples (trigrams). Longer n-grams keep a little word order, which matters most for negation.

Unigrams, bigrams and trigrams of one tweet The tweet this is not good split into unigrams (this, is, not, good), bigrams (this is, is not, not good) and trigrams (this is not, is not good). tweet: this is not good unigrams this is not good bigrams this is is not not good negation kept as one feature trigrams this is not is not good
N-grams are runs of consecutive words. With single words only, "not" and "good" are separate clues that pull in opposite directions; the bigram "not good" is one clear clue.

Bag of words and TF-IDF. Finally, every n-gram in the training set becomes a column, and each tweet becomes a row saying how often each one appears. That’s a bag of words: word order beyond the n-gram is thrown away. Plain counts treat every term alike, so common words dominate. TF-IDF (term frequency × inverse document frequency) scales each count down by how many tweets contain the term:

Bag-of-words counts and TF-IDF weights for three tweets Three tweets as rows and six terms as columns. Counts are all 0 or 1. With TF-IDF, good, which appears in every tweet, gets 0.23 to 0.28, while rarer terms such as not, today and morning get 0.40 to 0.48. BAG OF WORDS: HOW MANY TIMES EACH TERM APPEARS good not not good this today morning this is not good 1 1 1 1 0 0 so good today 1 0 0 0 1 0 good morning all 1 0 0 0 0 1 TF-IDF: THE SAME COUNTS, WEIGHTED BY RARITY good not not good this today morning this is not good 0.23 0.40 0.40 0.40 0 0 so good today 0.28 0 0 0 0.48 0 good morning all 0.28 0 0 0 0 0.48
Each tweet becomes a row of numbers, one column per term (only six of the columns are shown). "good" is in every tweet, so TF-IDF gives it the least weight. Values follow scikit-learn's default formula, with each row normalised over all its unigrams and bigrams.

With about 100,000 training tweets and tens of thousands of columns or more, almost every cell is zero. Models built for this kind of sparse data are fast and hard to beat, which is part of why the linear ones did so well.

The models we picked

We picked three models that use those numbers in different ways, plus a baseline that uses none:

  • TextBlob was the baseline. It doesn’t learn anything: it looks words up in a fixed dictionary of positive and negative scores. It showed what we got for free.
  • Multinomial Naive Bayes learns how likely each term is in positive and in negative tweets, then multiplies those likelihoods for a new tweet. It assumes terms are independent, which isn’t true, but it trains in seconds and is the standard first model for text.
  • A linear SVM learns one weight per term and draws the boundary between the classes with as wide a margin as possible. With tens of thousands of sparse features, of which only a few matter in any one tweet, that’s exactly the setting it’s good at.
  • A convolutional network (CNN) skips the counts. Each word becomes a small learned vector, and filters slide over windows of a few words at a time, learning to detect phrases wherever they appear. Ours was small: a 5,000-word vocabulary, 25-dimensional word vectors, and a short stack of convolution layers.
How each model scores this is not good Naive Bayes multiplies per-word likelihood ratios; not favours negative and good favours positive. The linear SVM sums learned weights, and the bigram not good has a large negative weight, so the sum is negative. The CNN turns words into vectors and a filter over a window of words fires on not good. All three output negative. pushes toward negative pushes toward positive numbers are illustrative Naive Bayes how typical each word is of each class this is not good ≈1 ≈1 2.1× 1.8× multiply negative Linear SVM one learned weight per word or phrase this is not good not good 0 0 −0.4 +0.6 −1.1 sum −0.9 negative CNN word vectors, then filters over windows this is not good filter fires on "not good" max, dense negative
Three ways to reach the same answer. Naive Bayes and the SVM only see which words and phrases occur; the CNN sees word order inside each window. The numbers are made up to show the mechanics, not taken from our models.

Naive Bayes and the SVM only see which terms occur. The CNN sees word order within a window, which is why I expected it to win.

What we measured

After cleaning, we held out a balanced test set of one million tweets and trained on a sample of about 100,000. Before splitting, we removed any training tweet whose text also appeared in the test set. That was a good call, since Twitter is full of identical tweets.

On that test set:

Model Twitter accuracy
TextBlob, no training (baseline) 0.59
CNN (Keras, 5,000-word vocabulary) 0.777
Multinomial Naive Bayes, TF-IDF 1–3 grams 0.784
Linear SVM, TF-IDF 1–3 grams 0.820

The linear SVM won, and that was the headline of our report. A fair reading is narrower. Anything trained beat the off-the-shelf lexicon by about twenty points, and the three trained models then sat within about four points of each other.

We ran the same three models on Amazon product reviews. All of them scored between 90% and 94%. Changing the dataset moved accuracy by eleven to fifteen points. Changing the model moved it by about four.

Accuracy by model on Twitter and Amazon reviews CNN 77.7% on Twitter, 92.75% on Amazon. Naive Bayes 78.4% and 89.97%. Linear SVM 82.0% and 93.71%. Twitter (Sentiment140) Amazon reviews 75%80% 85%90% 95% CNN CNN on Twitter: 77.7% CNN on Amazon: 92.75% 77.792.8 Naive Bayes Naive Bayes on Twitter: 78.4% Naive Bayes on Amazon: 89.97% 78.490.0 Linear SVM Linear SVM on Twitter: 82.0% Linear SVM on Amazon: 93.71% 82.093.7
Each line is one model on two datasets. The lines are long, while dots of the same colour sit close together: the data moved the result far more than the choice of model did.

Reviews are longer, the words are more specific, and a star rating is a cleaner label than whatever a tweet implies. If I had to predict how well a sentiment model would do, knowing what text it would see would help me much more than knowing its architecture.

The labels were already a model

Sentiment140 was never labelled by people. Its authors collected tweets containing emoticons, treated :) as positive and :( as negative, and then removed the emoticons from the text. So every one of our models learned to predict whether the author had typed a smiley.

That explains the gap with Amazon better than anything we tried. A tweet like “finally finished my exam :(“ carries a negative label that the remaining words hardly support. Beyond a certain point, no classifier can recover information the labelling process never captured. Asking “why can’t we get past 82%?” was really asking about the data.

The stopword that mattered

The experiment I’m happiest with in hindsight is a small one. With Naive Bayes, we compared three ways of handling stopwords:

  • keep every word: 0.78
  • remove scikit-learn’s standard English stopword list: 0.75
  • remove our own list of the twenty most frequent words, excluding not: 0.77
Notebook plot of validation accuracy against number of features for three stopword settings. Keeping stopwords sits near 0.78, the custom list near 0.77, the standard list near 0.75.
The original notebook plot: Naive Bayes validation accuracy as the vocabulary grows. The standard list (orange) stays lowest at every size.
What each stopword setting leaves of "this is not good" Keeping every word leaves "this is not good", accuracy 0.78. The standard English list removes this, is and not, leaving "good", accuracy 0.75. Our list removes is but keeps not, leaving "this not good", accuracy 0.77. keep every word this is not good 0.78 standard English list this is not good reads as positive 0.75 our list, kept "not" this is not good negation survives 0.77
One example sentence under each setting, with the Naive Bayes validation accuracy on the right.

The standard list includes not. Dropping it turns “not good” into “good”. Our cleaning step had already expanded contractions for this reason (“don’t” became “do not”), and then the stopword list removed the part we had been careful to keep. The custom list skipped not deliberately, and in the notebook that exclusion is a single line: del custom_stop_words[2].

It’s the least impressive-looking line in the project, and it is also the best example of the actual job: find out what a preprocessing step removes before assuming it removes noise.

What I would question a year on

Rereading your own work from eighteen months ago is humbling. A few things I would change.

The CNN comparison was unfair. It had a 5,000-word vocabulary, 25-dimensional embeddings, tweets cut at 50 tokens, and 100,000 training examples. The SVM had 1–3 grams over the full vocabulary. The training curve shows the same thing from another angle:

Notebook plot of CNN loss over 20 epochs. Training loss falls from 0.35 to 0.33 while validation loss rises from 0.52 to 0.55.
CNN loss over 20 epochs, from the notebook. Training loss (blue) falls; validation loss (orange) rises from the first epoch.

Validation loss went up from the very first epoch, so the network was memorising the training set rather than learning anything that transferred. With that curve, the right move was to stop, get more data or regularise more, not report its accuracy next to the others. So “the SVM beat the CNN” really means “a well-fed linear model beat a network that was overfitting”. That can still be the right practical choice, but it’s a different claim.

One normalisation experiment never ran. The lemmatisation and stemming cells build a normalised copy of the data and then split the original data for training. Their results match the unnormalised run exactly, and we concluded that normalisation made no difference. We never tested it.

The validation AUC was computed from hard labels. During the feature search, the ROC curve was built from predicted classes instead of scores, so the “AUC” matched accuracy almost exactly and added nothing. The final test runs used probabilities correctly. Even so, our README pairs the linear SVM’s accuracy (82%) with an AUC of 0.86, which belongs to the RBF kernel. The linear model’s own AUC is 0.88 in the saved notebook and 0.90 in our slides; those come from different runs, and we never wrote down which one the reported accuracy came from.

The RBF SVM was scored on a 5% sample of the test set, because the full set was too slow. It appears in our comparison table next to models scored on the full million.

None of these change the main conclusion. They do show that the numbers we were proudest of were the least carefully checked.

The notebooks, figures and report are on GitHub.

References

  1. Alec Go, Richa Bhayani and Lei Huang. Twitter Sentiment Classification using Distant Supervision. Stanford CS224N project report, 2009. https://www-cs.stanford.edu/people/alecmgo/papers/TwitterDistantSupervision09.pdf — the Sentiment140 emoticon-labelling method.
  2. Christopher D. Manning, Prabhakar Raghavan and Hinrich Schütze. Introduction to Information Retrieval. Cambridge University Press, 2008. https://nlp.stanford.edu/IR-book/ — covers TF-IDF and Naive Bayes text classification.
  3. scikit-learn developers. Feature extraction: text feature extraction. scikit-learn documentation. https://scikit-learn.org/stable/modules/feature_extraction.html#text-feature-extraction
  4. Steven Loria. TextBlob: Quickstart. TextBlob documentation. https://textblob.readthedocs.io/en/dev/quickstart.html
  5. Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning 20, 1995. https://link.springer.com/article/10.1007/BF00994018
  6. Yoon Kim. Convolutional Neural Networks for Sentence Classification. EMNLP, 2014. https://arxiv.org/abs/1408.5882