3.1 How automatic methods work
3.1.4 Common approaches
Now that you have discovered the elemental magic behind how computers decide on documents' themes let us look into three key approaches to the whole package of automatic assignment of subject terms.
1. TEXT CATEGORISATION
Often named text categorisation or text classification, the most common approach to automated subject indexing is applying supervised machine learning algorithms.
In this approach, the algorithm 'learns' about characteristics of target index terms based on characteristics of documents (e.g., frequently occurring tokens) that had already been indexed manually with those terms; these documents are called training documents. The algorithm is then tested with a new set of documents from the collection, which are called test documents. Between these two steps, another set of documents is often used to select the best settings (e.g., application of which stemming approach, if any), known as validation documents or the development set.
There are many ways to build supervised machine learning classifiers, such as support vector machines (SVM), artificial neural networks, random forest learning, adaptive boosting, linear models, tree-based methods, and, most recently, deep learning approaches such as transformer models. Also, two or more different classifiers can be combined to make a classification decision. In the literature, these approaches are called multiple classifier systems or ensembles.
The supervised machine learning approach requires a relatively large number of training documents per each target index term or class. However, for many document collections, too few training documents will be available for the classifier to perform well. In these situations, semi-supervised learning can be adopted. However, semi-supervised learning has not been made clear to what degree it could help with large KOSs such as those used in libraries. Instead, it has been reported that it is hard to apply supervised machine learning on large KOSs, one of the reasons being that in large KOS such as Dewey Decimal Classification (DDC), the distribution of existing documents will be very skewed across the classes; this is because we may have a lot of books in the class 'metaphysics' but not so many in the narrowest branch of that class, 'number and quantity'. (For more examples, see the Swedish version of DDC and works classified by it at the Swedish DDC browsing interface, an excerpt shown in the figure below).
The Swedish DDC browsing interface figure shows that class Metaphysics (110) has 6739 works while class Number and Quantity (119) only one works. Too few works in a class make it hard to train the machine learning algorithm.
2. STRING-MATCHING APPROACH
3. DOCUMENT CLUSTERING
Of relevance is also an approach known as document clustering, an unsupervised machine learning approach in which the target KOS, meaning both its terms and relationships between the terms, are derived automatically. Although this comes with challenges, the approach could be relevant for document collections with insufficient training documents and where no appropriate KOS is available to support implementing a lexical approach. However, the problems arising from difficulties in automatically deriving names of groups of related documents and the relationships between them may outweigh the actual benefit. Furthermore, as new documents are added, the names and relationships change, which is not user-friendly, so document clustering may be more appropriate for other tasks like organising web search engine results.
Finally, for an overview of text classification from a linguistic perspective, you may visit the "Text Analysis: Linguistic Meets Data Science" course rather than the knowledge organisation of cultural heritage resources. Its Unit III discusses sentiment analysis, topic modelling and text classification. Also, the "Digital Historical Research on European Historical Newspapers with the NewsEye Platform" presents the application of automated approaches for finding and analysing historical newspapers. In Unit II, you can also learn about Named Entity Recognition (NER) and Document Semantic Enrichment (DSE), two processes which may also be helpful in our context. Automated approaches to knowledge organisation may need to resort to such digital tools, showing how knowledge organisation is essential for Digital Humanities. Still, digital tools like the NewsEye Platform are also helping with knowledge organisation.
Exercise 1: Which approaches would you choose when adding an unindexed document collection to an existing, indexed collection (click all that apply)?
Exercise 2: Which approaches would you choose when your document collection is independent of another collection (click all that apply)?
For more information on different approaches to automatically deriving and assigning a theme and other application areas, please see Unit III (Text Analytics and Language) of the Text Analysis: Linguistic Meets Data Science OER. For a more detailed overview of natural language processing, please see Unit III (Information Extraction and Document Understanding) of the OER titled Digital Research on European Historical Newspapers with the NewsEye Platform. It describes how to automatically extracted named entities (like names of people or cities).
FURTHER READING
-
Golub, K. (2021). Automated Subject Indexing : An Overview. Cataloging & Classification Quarterly, 59(8), 702–719. https://doi.org/10.1080/01639374.2021.2012311