1.2. Data science research perspectives
1.2.4. Preprocessing of Texts: Stop words
So, we have now selected our population and acquired a corpus of texts that we consider representative of that population. Depending on our research question and the meta-data available in our structured dataset, we have also selected a sample that fits our inquiry and our study's time constraints. We can now move on to working directly with our pile of texts.
Most of the time, the first step of analysis is to preprocess our texts to fit our selected method and research question. Depending on what we are interested in looking at, the preprocessing often involves some kind of cleaning the texts and removing things like punctuation and formatting. These things often cause issues with analytic tools. While preprocessing is usually done in corpus linguistic research, the approach is somewhat different, with linguists being more concerned about removing textual content.
We need to notice here that preprocessing will depend entirely on the other steps we have planned and that there is no "correct" way of going about it. For example, if we want to break the texts into sentences to study collocations within the sentence only, we will need to keep the punctuation. Incorrect preprocessing can end up causing us untold amounts of problems further down the line, so we must sit down and compose a workflow with thought because otherwise, we may cause quite a mess.
Common types of text preprocessing include stop word removal, stemming and lemmatization. Stop word removal relies on a list of words deemed irrelevant to our interests and often includes conjunctions and articles, or what linguists would call function words. Removing these words from the data allows us to, for instance, look at a word list for highly frequent words and have the results come out a little cleaner and more readable. On the other hand, linguists are often particularly interested in grammatical structures, so it is prudent to keep in mind when to remove stop words and when not to! It is also good to know that there is no single list of stop words. There are some widely used ones, such as the one used in the Natural Language Tool Kit (NLTK). Still, different programming environments and individual research projects also use and distribute their own. The stop words in the NLTK list are given below.

The differences between the lists can be pretty dramatic. The figure below visually compares the number of words in some well-known stop-word lists.

Depending on the stop word list size, the text may look very different after removing the stop words. This can have a significant effect on the subsequent analysis. For topic modelling, stop word removal generally works quite well because it removes function words that do not carry topical information. On the other hand, sentiment analysis results are likely to be affected much more. If we remove personal pronouns, negation words like "no", "not", and "never", or modal verbs like "should" or "must", we might turn a sentence like "we must never harm another human being" to "harm another human being"! The example below shows the effect of stop-word removal using a well-known literary quote.
