3.1. Text analytics and language

3.1.2. Machine learning and pattern identification

Machine learning is a field of research that most corpus linguists probably hadn't heard much about ten or fifteen years ago. However, today, it is mentioned more and more often, especially in the more computational and quantitative end of linguistics. Until recently, most corpus tools did not include machine learning functions.

By contrast, data scientists working in text analytics are no doubt very familiar with machine learning. Indeed, it is fair to say that machine learning is one of the key concepts in text analytics.

So, let's begin by dispelling one misconception straight away. Machine learning does not really mean that a machine, in this case, a computer, is actually learning anything. Machine learning refers to a wide variety of statistical and algorithmic methods that perform grouping, clustering and classification tasks. What does this mean in practice? And how can linguists make use of machine learning?

Machine learning methods allow us to identify patterns in data. A pattern, in this case, would be a number of variables, each of which can take various values. For example, the author of a short story will use a variety of different words and phrases, and we can calculate the frequency at which each of those words gets used. A second story by the same author will naturally use somewhat different words but probably also many of the same words; the same applies to a third, fourth, and so on. Individually, the stories simply contain varying amounts of different words. Still, when we look at them closely, we may discover that the author appears to have a distinct tendency to use certain words particularly often. That is a pattern. A machine learning algorithm can identify these patterns as the tell-tale sign of a particular author. Subsequent stories could then be analysed to see whether they match the same patterns. For more on this type of analysis, see 3.4.3 Authorship attribution.

The example above actually introduces two fundamental concepts in machine learning. When we look for distinct patterns of word use by authors, we are carrying out unsupervised learning. This simply means that we ask a computer program to identify similarities between a dataset's units (such as texts in a corpus). If groups of units with similar features appear, they may be meaningful in some way, and we can name them as we see fit. Once we have identified group features, such as the distinct lexical use of an author, we can carry out supervised learning, where we look for patterns in new data that match the already identified features. A well-known example of supervised learning is spam email detection, where known characteristics of spam email are used when identifying which new emails in your mailbox are spam.

Unsupervised machine learning methods are primarily clustering and dimension reduction methods. Dimension reduction means that we take a large number of variables and attempt to find co-occurrence patterns between them. If we know that there is a group of variables that always behave the same way, we can either remove the variables and replace them with one new variable that represents the cluster, or we can identify the variable that is the best representative of the cluster and remove the others.

The most fundamental supervised machine learning methods are linear regression and logistic regression modelling (the two YouTube videos are by Josh Starmer of StatQuest) and more sophisticated regression models. Other methods include conditional inference trees, decision trees and random forests, support vector algorithms, nearest neighbour clustering, etc.

If you are interested in learning more about the details of machine learning in an accessible way, we warmly recommend The StatQuest Illustrated Guide To Machine Learning (2022) by Josh Starmer.