3.1. Text analytics and language
3.1.1. Text analytics and language
Moving on from the corpus linguistic perspective on language, we will now have a look at how text analytics may approach written materials. As previously mentioned, data scientists are typically interested in carrying out computational operations that allow us to do something with a specific set of texts, such as classifying them into groups based on similar content and by pre-determined characteristics. In this unit, we will explore four ways data scientists may engage with their text collections: Sentiment Analysis, Classification, Topic Modelling, and Authorship Attribution. In each lesson, we will explain how the methods work and how to carry out the analysis in Knime.
Sentiment analysis, as the name might suggest, is an approach where the sentiment (positive or negative) of a text or excerpt is annotated. This relies on the text material being tagged as either expressing a positive or negative opinion or emotion, relying on different factors depending on the method being applied. In this unit, we will learn about lexicon-based sentiment analysis and machine learning-based sentiment analysis.
Text classification involves the automatic classification of a text as belonging to one or several of a previously defined set of categories. This can be done based on multiple attributes of the text, such as keywords or meta-data, and will often rely on training a machine learning model. Topic modelling approaches the idea of categorizing texts from the other side. It is instead intended to create a model of the topics in the text material from the contents of the material itself.
Finally, authorship attribution is the employment of the full analytical repertoire in order to ascertain the originator of a text, often within a collection of texts. Popular uses of this that you have likely encountered previously would be things like plagiarism detection. However, the concept can be taken much further than the sentence-matching often employed in those scenarios. By using language patterning focused on sentence lengths, word choices and other features of the text, a far more useful type of authorship attribution can be arrived at.
As hinted at previously, the change of perspective here towards data science will also include a more practice-oriented influence in the methods we are exploring. While corpus linguistics is often employed in research to explore language, text analytics often serve more practical purposes in business organizations. Throughout this unit, we will have a look at these uses of the methods as a way of starting a discussion on the differences (and similarities) between corpus linguistics and text analytics.
It is also notable that machine learning will start to play a larger role in the methods we are going to look at. While this might initially sound a bit scary, machine learning in practice is not as difficult and complicated to make use of as some believe. For those of us who are engaging with it for the first time, watching this YouTube video by WIRED might help demystify it and will hopefully make the idea more approachable.