3.3. Topic Modelling
3.3.1. Topic modelling
Topic modelling brings us back to the idea of supervised and unsupervised ML. As previously mentioned, supervised ML makes use of a training dataset that teaches an algorithm to identify our pre-existing categories. On the other hand, unsupervised ML will attempt to find any patterns in the material without relying on a training data set to distinguish between pre-existing categories. Topic modelling falls into the category of unsupervised ML, and its practical use case is detecting the topics of a set of documents and then clustering the documents along with those topics. There are a few different ways we can go about topic modelling, and in this lesson, we will look at a few of the common applications and uses. We will then move on to have a look at how we can use KNIME for topic modelling.
The general idea of topic modelling is providing the tool with a set of documents and then letting the tool figure out the features of the documents and place them into categories. More information about how automated categorization can be useful in the humanities can be found in OER 7, where it is applied to knowledge organisation. Drawing on what we already know about corpus linguistics, it makes sense that this could be done by keyword analysis. In practice, this means that a model is built by seeing which words are frequent in the documents and then sorting the documents according to the high-frequency items found in each. When we do topic modelling, we accomplish this through two separate steps.
The first step is figuring out which topics are covered by our collection of documents. As previously hinted, one way to do this could be by identifying keywords that appear frequently in the documents. For instance, we could apply pre-processing to remove function words and stopwords and then simply look at the frequency of remaining items for each document to obtain our list of topics. We recognize this approach from the Bag of Words approach to sentiment analysis. However, much like with the Bag of Words approach to sentiment analysis, more accurate methods are available to us here.
The underlying assumption that drives topic modelling is that documents are made up of distinguishable topics and that the topics can be defined as words. The approach described in the previous paragraph functions on this assumption but will miss topics created from several words. As we remember from our corpus linguistics unit, multi-word items are referred to as N-grams. Many tools used for topic modelling include N-grams in their frequency lists and can identify topics made up of multiple words because of it.
As we have acquired our list of topics in our dataset, we can move on to the second step. The second step is where we take our list of topics and figure out in which individual document which topic appears. There are many approaches, but this lesson will focus on Latent Dirichlet Allocation. The former relies on the Bag of Words approach and looks for the topics as individual items devoid of context, while the latter takes the words surrounding the topic item into account and provides a more context-sensitive topic model.
As we now have a good idea of how topic modelling works under the hood, we will move on to having a look at how topic modelling can be used in KNIME.