2.1. Concordances, tagging, and annotation

2.1.1. What are concordances?

The most basic operation in corpus linguistics is to search for all occurrences of a word (or phrase) of interest in the corpus. This is similar to searching through a webpage or a word-processing document by clicking Control+F. Still, in the linguistic context, we want to retrieve all instances of the query word at once and to be able to view all of them in the most convenient way.

A concordance list is a list of all relevant instances of the query word, often presented alphabetically or in order of appearance in the text or corpus. This list is made from all the concordance lines, which present the query word and a span of contextual words on either side. These lines are also called Keyword in Context lines or KWIC lines. While this might seem simple, concordancing is an integral part of much corpus linguistic research, and the resulting KWIC lines will, in many situations, be our way of inspecting the texts we are working on. When we are interested in looking at a specific word, we will often use the concordance to get an overview of how that word is used in our corpus. Concordance lines are usually extracted from the corpus tool and then inspected in a spreadsheet application. Linguists will often classify the concordance lines according to the linguistic features of each individual line. The example shown below uses the previously mentioned tool AntConc by Lawrence Anthony.

xxx


We will look at two commonly used standalone tools for concordancing. Both tools offer many other functions beyond concordance lines, and we will return to those functions later in the unit.

The first tool we will introduce is called AntConc. As a freely available and relatively versatile piece of software for all the major operating systems, Antconc is one of the most commonly used corpus tools when newcomers are introduced to corpus linguistics during their studies; an extensive set of guides is available through its creator's YouTube channel. The tool's main function is to allow us to perform queries and then provide us with all the KWIC concordance lines containing the word we searched for. AntConc also has the functionality to let us look at which words appear around our keyword (collocation) and can be combined with Part of Speech (POS) tagged materials, which we will explore in a future lesson. 

Another excellent example of a concordance tool is CasualConc. Unlike AntConc, it is Mac OS exclusive and can not be used on Windows machines. The tool covers much of the same ground as AntConc, such as concordancing, collocations, and frequency lists, but it also includes functionality for creating visualizations. Some examples of uses can be found on CasualConc's Google site. 

So why do linguists use these types of tools? To a certain extent, it has to do with the linguistics research perspective. These tools allow us to explore single words, phrases or syntactic structures in their context to see how they function in our sample. Furthermore, with the addition of pre-processing, such as Part-of-speech-tagging, we can extend our queries beyond just words and phrases to explore specific grammatical structures or patterns. In addition, the ability to create frequency lists is very useful and can often tell us a lot about our texts. Knowing which words are common, especially in comparison to other texts, can provide valuable insights into how a language, or a specific register of a language, works. For example, comparing spoken and written English, or fiction and academic writing, will reveal many differences, some of them predictable and others surprising.