3.3. Topic Modelling

Discussion from linguistics perspective

3.3.2. Topic modelling in KNIME

In KNIME, there are several options for topic modelling, using both the LSA and the LDA methods discussed in the previous lesson. For our example, we will focus on the LDA method, and we will be making use of the Topic Extractor node from NodePit. The Topic Extractor node makes use of multi-threading, which might cause some issues on your machine. If this is the case, simply select to run the node on a single thread in its settings. The topic modelling resource used in the node comes from the Machine Learning for Language Toolkit (MALLET).

Think of the node as a wrapper for the MALLET resource, which provides us with a graphic user interface for what is originally a tool used from the command line or as an external resource in programming. The node allows us access to the tool's settings, such as the number of topics extracted and the length of N-grams considered as topics.

Why would we want to limit the number of topics extracted from our dataset? Partially, this is to remove clutter, as considering every recurring token (word) would render the entire process useless and end up providing us with what is essentially a type list, meaning a list of every unique word in our dataset. The MALLET documentation recommends the use of 200 - 400 topics, depending on the dataset size, and states that this will result in a fine-grained yet functional model. In reality, this should be understood in relation to our purposes. Will our project benefit from larger clusters containing a higher number of documents with broad topic definitions? Or would we benefit more from a fine-grained topic model with narrow topics and fewer documents in each cluster?

We should take this perspective with us when deciding on the number of words per topic. Setting the Topic Extractor node at a high N-gram length would, in theory, mean that we can catch more complex topics. However, much like with the number of topics extracted, a balance must be kept in order for our results to be useful.

The most important thing when deciding on the length of our N-grams and the number of topics to extract from our dataset is having some idea of what our dataset is made from. If we are looking for complex topics that we know will include multiple words and where we know that those extra words are central to the accuracy of the model we want to create, a longer N-gram setting makes sense. Having some idea about the document's content can also clue us into the appropriate number of topics. For instance, a dataset with a narrow focus could do with a lower number of topics while a dataset with a lot of variety would benefit from a longer list of topics to portray the content correctly.

Looking at this example of a KNIME workflow, we can see how the Topic Extractor node is used. The top example makes use of the node in a straightforward manner that matches the procedure we have already discussed, while the bottom example uses a more intricate method in order to calculate an appropriate number of topics for a given dataset. The second example would be an excellent way to develop your topic model if you are creating a workflow that you intend to reuse frequently or if the end goal is a tool for visualizing different datasets with differing parameters.