3.1 How automatic methods work

3.1.1 Why automatic methods?

We have already talked about problems of online searching, including web search engines that rely on automatic methods. Let us look at some examples of why this is problematic and why there is also a need for knowledge organization systems and human experts. At the same time, automatic methods could help -- but to what degree and how exactly?

Let us start with a video overview of problems faced by the searcher for humanities information resources today. The video is a 30:45 min long webinar recording titled "Searching for humanities -- The promise of the digital" from June 2020. The talk was given at the "Enriching Metadata -- Enriching Research" webinar series organized by the Swedish National Heritage Board and Digital Humanities Uppsala at Uppsala University, supported by the Europeana Research Grants Programme. You will find the following topics covered: background (starting at 1:05), the challenge of searching (4:26), need for KOS (6:26), bibliographic objectives (10:26), finding humanities resources (11:40), KOS at search interface (15:23), creating own metadata/collection (20:12), choosing KOS (22:50), alternatives to KOS (25:00), concluding remarks (28.25).  


Building on the video, let us look at common retrieval approaches and associated problems used to discover an increasing number of information resources: books, music, films, images, blogs, websites, data, digitized historical newspapers, digitized museum objects, 3D objects, etc. Once we have them on the Internet, how can we find them? Different search systems have different mechanisms. Most commonly used are web search engines, which automatically compare the text we input in their search box and what they have in their databases. However, these automated methods often fail for several key reasons:

  1. We (users who type text in the search box) may use different search terms than the authors of the documents in the databases, which could lead to incomplete search results. For example, when we write 'trip', the search engine will not find documents that do not use the word 'trip' but are also about travels and use, for example, words like 'journey' or 'travel'). This is due to the synonymy of the natural language, meaning different terms are used to refer to the same or very similar concept.

  2. Our use of the search term could mean something other than the authors of the documents. For example, our search term 'bow' will find documents on ships, ribbons, string instruments, and weapons. This phenomenon is known as homonymy, whereby the same term refers to two or more unrelated concepts. 

  3. We often get too many results, often in the hundreds of thousands, and most of us do not take the time to go through them, which may lead to missing crucial documents. This is known as high recall.

  4. At the same time, we may get many irrelevant results and miss important ones, especially in the cases of images, audio, video and other types of documents with little text. Too few relevant results are known as low precision.

  5. Also, search engines often assume that the user knows which term to use instead of providing an overview of subjects covered by the database.

Google search engine interface with the

Google search engine interface with the "I'm Feeling Lucky" button that would take the user to the one top-ranked result. The problems listed above make it hard for this feature to deliver the best result consistently.

On the other hand, databases of libraries, for example, use KOS such as thesauri, subject headings systems and classification systems (see figure below from Zeng 2008). Using subject terms from KOS provides numerous benefits compared to the free-text indexing of commercial search engines. Subject index terms from KOS are crucial for describing non-textual information resources. The benefits include:

  1. KOS terms provide consistency in indexing through uniformity in term format and assignment of terms. This helps ensure the user gets only relevant and complete results (high precision and high recall).

  2. KOS incorporate semantic relationships among terms, allowing the user to express the right level of specificity of their search term and disambiguate homonyms. For example, the homonym 'table' can be an 'HTML table' or a piece of 'furniture', and the user can decide which one they are looking for; then, if looking for a table as a piece of furniture, the system can check with the user what they mean more specifically: 'coffee table', 'dining table' or maybe even a 'desk' etc. to retrieve only relevant results.

  3. Support for hierarchical browsing through consistent and clear hierarchies, which allow the user to learn about the topics covered by the database and support the user in navigating to the best term.

Types of KOS and their significant functions (Zeng 2008).

Types of KOS and their significant functions (Zeng 2008).

However, despite all the potential benefits of KOS, the intellectual organization of information based on KOS is very resource intensive. Just creating, managing, maintaining and updating KOS requires resources. Furthermore, subject indexing and classification based on the KOS for every document take lots of time for libraries, museums, archives, digital humanities projects, etc., where we often face the reality that we need to do more with less. This leads us to the sad situation we have today: much of the information resources remain undiscovered due to increasing reliance on purely automatic methods, and even when an institution or a project uses a good KOS, they are rarely visible at the level of the search interface (will talk more about this in Unit IV). With this, cultural heritage institutions often fail to meet user needs in their online systems, especially their ability to find everything on a specific topic in their database and retrieve only relevant information.

Then we must ask ourselves, can we use automatic means to support KOS-based subject indexing? The simple answer is that we can, but only if we control the automated suggestions. This indexing approach is known as machine-aided indexing (MAI) or computer-assisted indexing (CAI), where the human indexer makes the final choice based on suggestions provided by the computer. A prominent example is the Medical Text Indexer used by the U.S. National Library of Medicine.

A well-tested and implemented machine-aided indexing software at the U.S. National Library of Medicine.

A well-tested and implemented machine-aided indexing software at the U.S. National Library of Medicine. The image is taken from https://lhncbc.nlm.nih.gov/ii/tools/MTI.html.


REFERENCES
  1. Zeng, M. L. (2008). Knowledge Organization Systems (KOS). Knowledge Organization, 35(2/3), 160–182. https://doi.org/10.5771/0943-7444-2008-2-3-160