2.1. Concordances, tagging, and annotation
2.1.4. Manual pruning and classification of corpus results
One of the great benefits of corpus queries is that it is quick and efficient to find all occurrences of a word, phrase or syntactical pattern, even in a large corpus. However, it is very rare that a query's results would be immediately ready for analysis. Typically, the concordance list will include lines that are not relevant to the present study, which need to be discarded or pruned before the analysis can move on to the next stage. Likewise, it may be necessary to run multiple queries and combine the concordance lines together.
Much of the time, the research question will require that some additional classifications be carried out on the data based on some linguistic parameters. For example, it may be necessary to classify syntactical structures or the sentence's voice, tense or polarity. Linguists typically do this by opening the concordance lines in a spreadsheet application and manually classifying the lines by adding new analytical columns for each parameter included.
This may sound old-fashioned and labour-intensive for someone from a data analytics background. Couldn't these classification tasks be done automatically? Yes and no. While some of these classification tasks could indeed be automated using machine-learning solutions, many others are currently impossible to automate reliably, and even if they were, doing so might take longer than carrying out the classification task by hand. There is also an important distinction to be kept in mind when it comes to linguistic analysis and general computational text analysis. Whilst the latter typically aims to train a classifier to perform at an acceptably high level of accuracy, most linguists would insist on manual classification in order to achieve perfect accuracy. For linguists, the most interesting data observations are often difficult to classify because they reveal weaknesses in the existing theory. For the text analyst, by contrast, the small number of classification errors is often meaningless because the overall performance of the model is sufficiently high.
Something else that might be even more surprising is that there are numerous different theoretical perspectives on almost every linguistic phenomenon. In text analytics, tasks such as part-of-speech tagging are often treated as solved problems; when deemed useful, a POS-tagger or a parser is added to the preprocessing workflow and the word-class tags and syntactic annotations are added by a tagging library. However, many linguists would want to know exactly what grammatical theory the tools are based on. For a linguist, there are generative grammars, phrase-structure grammars, transformational grammars, dependency grammars, systemic functional grammars, etc. The tagging tool available might be based on an entirely different grammatical theory than what the linguist is interested in and, therefore, produce unusable metadata.
There are also many other annotation tasks that are still quite difficult to operationalise using software — though developments in neural networks and machine learning methods such as BERT (Bidirectional Encoder Representations from Transformers) hold much promise. However, at the moment, tasks such as pragmatic and semantic annotation, which rely on a human interpretation of the textual context, are still very challenging to carry out algorithmically.