3.2 Natural Language Processing (NLP) programming basics
3.2.2 Natural Language Processing (NLP) advanced
In this lesson, we will explore examples of advanced text-processing techniques in the field of digital humanities. These techniques enhance text analysis performance and extract meaningful insights from textual data. For instance, preprocessing the text becomes crucial when visualizing the most frequently used words in a text. Otherwise, the word frequency analysis may be skewed by overrepresenting articles, adverbs, conjunctions, or other grammatical elements. Researchers can focus on the more significant content words like nouns, verbs, and adjectives by applying preprocessing techniques such as stop word removal, part-of-speech tagging, or named entity recognition. This improves the accuracy and relevance of text analysis results, enabling deeper explorations of language patterns, semantic relationships, sentiment analysis, or even comparative studies across different corpora or historical periods in digital humanities research.
Collocation
A collocation is a sequence of words that occur together unusually often. To understand collocations, we start by extracting from the
sentence a list of word pairs, also known as bigrams. The bigrams are lists of two words (bi), while n-grams are lists of n words (e.g. ..). This is easily accomplished with the function bigrams():from nltk import bigrams
from nltk.tokenize import word_tokenize, sent_tokenize
sentence = word_tokenize("The blizzard which has held New York and the surrounding country in its grip for the past forty-eight hours continues without abatement. The seas continue to lash the coasts, doing millions of dollars worth of damage and imperilling all shipping within one hundred miles of this port.")
list_of_bigrams = list(bigrams(sentence))
print(list_of_bigrams)
[('The', 'blizzard'), ('blizzard', 'which'), ('which', 'has'), ('has', 'held'), ('held', 'New'), ('New', 'York'), ('York', 'and'), ('and', 'the'), ('the', 'surrounding'), ('surrounding', 'country'), ('country', 'in'), ('in', 'its'), ('its', 'grip'), ('grip', 'for'), ('for', 'the'), ('the', 'past'), ('past', 'forty-eight'), ('forty-eight', 'hours'), ('hours', 'continues'), ('continues', 'without'), ('without', 'abatement'), ('abatement', '.'), ('.', 'The'), ('The', 'seas'), ('seas', 'continue'), ('continue', 'to'), ('to', 'lash'), ('lash', 'the'), ('the', 'coasts'), ('coasts', ','), (',', 'doing'), ('doing', 'millions'), ('millions', 'of'), ('of', 'dollars'), ('dollars', 'worth'), ('worth', 'of'), ('of', 'damage'), ('damage', 'and'), ('and', 'imperilling'), ('imperilling', 'all'), ('all', 'shipping'), ('shipping', 'within'), ('within', 'one'), ('one', 'hundred'), ('hundred', 'miles'), ('miles', 'of'), ('of', 'this'), ('this', 'port'), ('port', '.')]
Now, if we apply the same method in the filtered text (without stopwords), we can clearly see the collocation of words.
from nltk.corpus import stopwords
stop_words = set(stopwords.words('english'))
filtered_sentence = [w for w in sentence if not w.lower() in stop_words]
list_of_bigrams = list(bigrams(filtered_sentence))
print(list_of_bigrams)
[('blizzard', 'held'), ('held', 'New'), ('New', 'York'), ('York', 'surrounding'), ('surrounding', 'country'), ('country', 'grip'), ('grip', 'past'), ('past', 'forty-eight'), ('forty-eight', 'hours'), ('hours', 'continues'), ('continues', 'without'), ('without', 'abatement'), ('abatement', '.'), ('.', 'seas'), ('seas', 'continue'), ('continue', 'lash'), ('lash', 'coasts'), ('coasts', ','), (',', 'millions'), ('millions', 'dollars'), ('dollars', 'worth'), ('worth', 'damage'), ('damage', 'imperilling'), ('imperilling', 'shipping'), ('shipping', 'within'), ('within', 'one'), ('one', 'hundred'), ('hundred', 'miles'), ('miles', 'port'), ('port', '.')]
Collocation analysis is one of the methods for studying changes in discourse over time. An example of such analysis is developed in section 4 of this OER.
POS tagging
In digital humanities, Part-of-Speech (POS) tagging is valuable as it assigns grammatical categories to words in a text, enabling analysis of language structure, semantic relationships, and cultural aspects. POS tagging aids in tasks such as text analysis, information extraction, sentiment analysis, language modelling, machine translation, and corpus linguistics. It provides insights into word usage patterns, syntactic structures, and linguistic phenomena, facilitating understanding of language variation, change, and meaning within digital texts.
The part of speech explains how a word is used in a sentence. There are eight main parts of speech - nouns, pronouns, adjectives, verbs, adverbs, prepositions, conjunctions and interjections.
- Noun (N)- Daniel, London, table, dog, teacher, pen, city, happiness, hope
- Verb (V)- go, speak, run, eat, play, live, walk, have, like, are, is
- Adjective(ADJ)- big, happy, green, young, fun, crazy, three
- Adverb(ADV)- slowly, quietly, very, always, never, too, well, tomorrow
- Preposition (P)- at, on, in, from, with, near, between, about, under
- Conjunction (CON)- and, or, but, because, so, yet, unless, since, if
- Pronoun(PRO)- I, you, we, they, he, she, it, me, us, them, him, her, this
- Interjection (INT)- Ouch! Wow! Great! Help! Oh! Hey! Hi!
A part-of-speech tagger, or POS-tagger, processes a sequence of words and attaches a part-of-speech tag to each word:
from nltk import pos_tag
pos_tags_sentence = pos_tag(sentence)
pos_tags_sentence
[('The', 'DT'),
('blizzard', 'NN'),
('which', 'WDT'),
('has', 'VBZ'),
('held', 'VBN'),
('New', 'NNP'),
('York', 'NNP'),
('and', 'CC'),
('the', 'DT'),
('surrounding', 'VBG'),
('country', 'NN'),
('in', 'IN'),
('its', 'PRP$'),
('grip', 'NN'),
('for', 'IN'),
('the', 'DT'),
('past', 'JJ'),
('forty-eight', 'JJ'),
('hours', 'NNS'),
('continues', 'VBZ'),
('without', 'IN'),
('abatement', 'NN'),
('.', '.'),
('The', 'DT'),
('seas', 'JJ'),
('continue', 'VBP'),
('to', 'TO'),
('lash', 'VB'),
('the', 'DT'),
('coasts', 'NNS'),
(',', ','),
('doing', 'VBG'),
('millions', 'NNS'),
('of', 'IN'),
('dollars', 'NNS'),
('worth', 'IN'),
('of', 'IN'),
('damage', 'NN'),
('and', 'CC'),
('imperilling', 'VBG'),
('all', 'DT'),
('shipping', 'NN'),
('within', 'IN'),
('one', 'CD'),
('hundred', 'VBD'),
('miles', 'NNS'),
('of', 'IN'),
('this', 'DT'),
('port', 'NN'),
('.', '.')]
If we need more information or examples for each of these tags, we can use
nltk help. The following example explains the meaning of NNS tag.!python -m nltk.downloader tagsets
/opt/anaconda3/envs/jupyter/lib/python3.9/runpy.py:127: RuntimeWarning: 'nltk.downloader' found in sys.modules after import of package 'nltk', but prior to execution of 'nltk.downloader'; this may result in unpredictable behaviour warn(RuntimeWarning(msg)) [nltk_data] Downloading package tagsets to /Users/csu/nltk_data... [nltk_data] Package tagsets is already up-to-date!
from nltk import help
help.upenn_tagset("NNS")
NNS: noun, common, plural
undergraduates scotches bric-a-brac products bodyguards facets coasts
divestitures storehouses designs clubs fragrances averages
subjectivists apprehensions muses factory-jobs ...
For more information on POS tagging, OER2 proposes a dedicated lesson.
Chunking
Chunking segments texts and depends on PoS tagging. Like tokenization, which omits whitespace, chunking usually selects a subset of the tokens. Instead of just simple tokens that may not represent the text's actual meaning, it's advisable to use phrases such as South Africa as a single word instead of South and Africa as separate words. Also, like tokenization, the pieces produced by a chunker do not overlap in the source text. To find the chunk structure for a given sentence, the
RegexpParser chunker begins with a flat structure in which no tokens are chunked. The chunking rules are applied, successively updating the chunk structure. Once all the rules have been invoked, the resulting chunk structure is returned.The following example shows a simple chunk of grammar consisting of two rules. The first rule matches an optional determiner (
DT) or possessive pronoun (PP), zero or more adjectives (JJ) then a noun (NN) or more nouns (NNS). The second rule matches one or more proper nouns (NNP). We will use the previous sentence with the already detected parts of speech and run the chunker on this input.import nltk
pos_tags_sentence = pos_tag(word_tokenize("The blizzard which has held New York and the surrounding country continues without abatement."))
pos_tags_sentence
grammar = r"""
NP: {<DT|PP\$>?<JJ>*<NN|NNS>} # chunk determiner/possessive, adjectives and noun or nouns
{<NNS>+} # chunk sequences of proper nouns
"""
chunker = nltk.RegexpParser(grammar) # chunker
chunks_sentence = chunker.parse(pos_tags_sentence)
print(chunks_sentence)
(S (NP The/DT blizzard/NN) which/WDT has/VBZ held/VBN New/NNP York/NNP and/CC the/DT surrounding/VBG (NP country/NN) continues/VBZ without/IN (NP abatement/NN) ./.)
Of course, we can also visualize it directly with the help of Jupyter Notebook:
chunks_sentence
The following example shows how to do the same with name entities (New York in our case)
!python -m nltk.downloader maxent_ne_chunker
!python -m nltk.downloader words
from nltk import ne_chunk
ne_chunks_sentence = ne_chunk(pos_tags_sentence)
print(ne_chunks_sentence)
(S The/DT blizzard/NN which/WDT has/VBZ held/VBN (GPE New/NNP York/NNP) and/CC the/DT surrounding/VBG country/NN continues/VBZ without/IN abatement/NN ./.)
ne_chunks_sentence