2.1 KO in digital humanities
2.1.1 Identifying relevant concepts
In Unit I, you studied the meaning of KO and KOS and the benefits of using KOS to achieve better results in information retrieval. You also had the opportunity to learn about the types of KOS and some examples. The use of KOS is usually done by information professionals working in libraries, archives, and research centres. Still, it can also be done by general users trained for this purpose. KOS should be introduced as subjects in the cataloguing records of a wide range of resources, from text to sound and still or moving images. In such a manner, they can form subject access points in databases, allowing for better information retrieval. Before selecting the KOS terms to be included in the records, it is necessary to determine what the subjects are. This is the process you will learn about in this lesson.
Determining the subject of an information resource can be a challenging task. It can be done through an intellectual process of determining the subject of a text, image, sound, or video. In these cases, an information professional looks and listens to determine the subject. The result may vary from person to person, depending on their experience in the process, their knowledge of the subject area they are analysing, or the different perspectives they take when analysing an image, for example.
The process of intellectually identifying relevant concepts requires uniformity. ISO 5963:1985 provides a set of principles for an information professional to follow when conducting content analysis and mentions the creation of subject analysis grids to achieve uniformity in both the content analysis and the conversion of the selected concepts into the controlled vocabulary (KOS).
In the case of written documents, some parts help to understand their subject: title, abstract, table of contents, introduction, chapter and paragraph headings, conclusion, illustrative material, and typographical highlights. Non-written documents, such as audiovisual, visual, or sound documents, require different procedures because analysing the entire recording (e.g., a film projection) is demanding. It is usually done based on the title or summary, with the possibility of seeing or hearing the resource if the description is inadequate or inaccurate.
Content analysis and relevant concept selection can be done using text analysis techniques. To learn these techniques, please see the course:
The OER2 aims to bridge the gap between linguistic and data science approaches to text analysis. This OER seeks to provide an introduction not only to methods and techniques but primarily to the conceptual and theoretical differences and overlaps between the two fields of study.
In this course, see in particular the following lessons:
1.1 Linguistic research perspectives (study of linguistic features, corpus as sample)
1.3.6 Our Tool of Choice: KNIME (a general data science tool that allows us to analyse numerical data, text, and images)
2.2.6 Keywords (corpus linguistic approaches to keyword analysis)
3.3 Topic Modelling (the first step is to find out which topics are covered by our collection of documents. The second step is to take our list of topics and find out in which individual document each topic appears)
3.4 Text Classification (defining a set of categories and then training a model to recognise these categories in documents so that they can be classified accordingly)
The correct selection of concepts and their conversion into a controlled vocabulary is crucial to avoid natural language peculiarities such as synonymy (different terms for the same concepts), homonymy (same terms for different concepts), and polysemy (a term with several meanings).
In the following pages, you will study different KOS applicable to the digital humanities, which will help you practise converting selected natural language concepts into controlled vocabularies.
REFERENCES
- International Organization for Standardization. (1985). Documentation—Methods for examining documents, determining their subjects, and selecting indexing terms. (ISO Standard No. 5963:1985). https://www.iso.org/standard/12158.html