2.4. Frequency, distribution and inferential statistics

2.4.1. Frequency

Corpus linguistics is often described as a quantitative field of linguistics. While this does not mean that corpus linguists would not engage in qualitative analysis — both corpus queries and the analysis of query results almost always involve plenty of qualitative tasks — it is true that quantitative analysis plays a key role in corpus-based and corpus-driven methods.

Inference and representation by Jukka Tyrkkö


One of the key concepts in corpus linguistics is frequency, simply understood as the number of times a query term occurs in a text or a corpus. However, since texts and corpora come in all different sizes, there is usually little sense in comparing frequencies across corpora directly. Instead, most of the time, frequencies of linguistic features are discussed in terms of standardised frequency. In practice, standardised frequency usually refers to frequency relative to a pre-defined word count, such as 1000 words. Therefore, the frequency of a word in a corpus would be reported as being, for example, 0.05/1000 words or 12.3/1000 words. This would mean that for every 1000 words, the word of interest occurs 0.5 times or 12.3 times respectively. When the same denominator is used for each corpus being compared, the comparison makes sense regardless of the overall word counts of the corpora in question.

xxx


Depending on the research question, linguists may sometimes prefer to focus on proportional distributions between two or more features that are alternatives to each other. For example, a linguist might want to examine the choice between two different verb complementation patterns. In English, we can say either afraid + to + INFINITIVE (e.g. afraid to go) or afraid + of + ING VERB (e.g. afraid of going). The linguist would query the corpus for all instances of both patterns and then examine the proportional distribution between the two choices. Importantly, in this analytical model, we are always interested in both the proportional breakdown, which can be expressed as a percentage, and the absolute number of occurrences, which indicates the amount of data we have supporting our conclusions.

xxx