3.2 Natural Language Processing (NLP) programming basics

3.2.1 Natural Language Processing (NLP) basics

For a more profound introduction to NLP programming techniques, we continue using NLTK. NLTK is a widely used standard Natural Language Processing (NLP) and Computational Linguistics (CL) Python library with prebuilt functions and utilities for ease of use and implementation.

To use nltk, you must install it via pip. The following lines install requirements in Jupyter Notebooks (notice the ! in front of each instruction).

!pip install nltk
!python -m nltk.downloader punkt

In the previous lesson, we learned that it was possible in Python to transform Strings into Lists. There are simpler and faster methods to manipulate long texts and extract both sentences and words.

We need to apply an NLP technique called tokenisation to obtain the list of words from a sentence. Tokenisation is splitting the text (which elsewise is just a long stream of characters) into tokens, our basic processing unit. This is quite close to what English speakers would call words. Before separating a sentence into words, we first need to obtain the sentences from a text, a process called sentence splitting.

Thus, for separating a text into sentences, we can use the NLTK method, sent_tokenize, and for obtaining the tokens (words) in a sentence, we use word_tokenize as the following:

from nltk.tokenize import word_tokenize, sent_tokenize
sentence = "The blizzard which has held New York and the surrounding country in its grip for the past forty-eight hours continues without abatement. The seas continue to lash the coasts, doing millions of dollars worth of damage and imperilling all shipping within one hundred miles of this port."
sentences = sent_tokenize(sentence)
Here we can observe that the text provided by the sent_tokenizemethod consists of 2 sentences. The original text has been split into a list of two sentences.
print(len(sentences))
sentences
2
['The blizzard which has held New York and the surrounding country in its grip for the past forty-eight hours continues without abatement.',
 'The seas continue to lash the coasts, doing millions of dollars worth of damage and imperilling all shipping within one hundred miles of this port.']
Now we can extract each token (word) of one of these sentences.
words = word_tokenize(sentences[0])
print(words)
['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement', '.']
Note that this method is not only more straightforward but also more effective than those we have seen before. Here, punctuation (like the . at the end of the sentence) is recognized as a token.

Clean up the text.

When working with text mining applications, we often use the term stopwords. Stopwords are basically a set of commonly used words in any language (not just English).

Stopwords are critical to many applications because if we remove the words that are very commonly used in a given language, we can focus on the important words instead.

For example, in the context of a search engine, let’s consider that your search is how to develop text mining applications.

Suppose the search engine tries to find web pages that contain the terms how, to, develop, text, mining, and applications. In that case, the search engine will find many more pages that contain the terms how and to than pages that contain information about developing text mining applications because the terms how and to are very commonly used in English.

If we disregard these two terms, the search engine can focus on retrieving pages that contain the keywords: develop, text, mining, and applications, which would bring up pages that are actually of interest.

In the same way, when you try to extract information from text corpora, like historical newspapers text, you should be able to focus on the words that make the most sense and remove the stopwords.

In the following sentence, try to find all irrelevant words: 


In NLP, there are some generally developed pre-set lists with stopwords for many languages. First, we download them using nltk.downloader, and then we look at the English pre-set list.

Do not forget to install stopwords from nltk.

!python -m nltk.downloader stopwords
from nltk.corpus import stopwords
stop_words = set(stopwords.words('english'))
For example, for English, NLTK contains a list of 179 most commonly used words. You can also try other languages by replacing 'english' it with one of your choice (written in lowercase, e.g., 'turkish').
print(len(stop_words))
print(stop_words)
179
{'you', 'about', 'this', 'after', 'before', 'am', 'are', 'aren', 'that', 'into', 'so', "wouldn't", 'isn', 'by', 'further', 'not', 'hasn', "you're", 'here', 'when', 'his', 'our', 've', 'won', 'ours', 'a', 'it', 'having', 'why', "won't", "haven't", 'no', 'hers', 'an', 'of', 'i', 'shan', "isn't", "shouldn't", 'where', 'now', "couldn't", 're', 'nor', 'whom', 'who', 'same', 'hadn', 'their', "didn't", "wasn't", 'they', 'does', "you'd", 'above', "weren't", 'himself', 'both', 'how', "should've", 'which', 'very', 'and', 'the', 'them', 'me', 'while', 'out', 'during', 'o', 'be', 'do', 'then', 'needn', 'herself', 'these', 'until', 'what', 'below', "don't", "you've", 'as', 'mustn', 'over', "aren't", 'she', 'each', 'will', 'been', 'its', 'all', 'than', "shan't", 'him', 'is', 'yourselves', 'her', 'if', 'itself', 'any', 'own', 'yours', 'were', 'only', 'haven', 'because', "it's", 'up', 'he', 'we', 'from', 'll', 'being', "doesn't", 'your', 'wouldn', 'through', "you'll", 'in', 'couldn', 'mightn', 'themselves', "hasn't", 'd', 'had', 'weren', 'again', 'did', 'down', 'theirs', 'for', 'to', 'too', "mightn't", 'have', 'on', 'ain', 'most', 'myself', 'between', 'has', 'was', 'shouldn', 'm', 'wasn', 'some', "she's", 's', 'should', 'doesn', "needn't", "mustn't", "hadn't", 'don', 'yourself', 't', 'or', 'at', 'but', 'such', 'can', "that'll", 'once', 'doing', 'with', 'more', 'there', 'few', 'those', 'ourselves', 'under', 'off', 'ma', 'didn', 'other', 'my', 'against', 'y', 'just'}
Once we have the tokens, we can check the count and frequency distributions of words in a text. Computing these word frequencies in extensive collections of texts can also reveal the most commonly used words (or stopwords). For example, if we take the previous sentence, we obtain:
from nltk.probability import FreqDist

fdist = FreqDist(words)

print(fdist.most_common(10))
[('the', 2), ('The', 1), ('blizzard', 1), ('which', 1), ('has', 1), ('held', 1), ('New', 1), ('York', 1), ('and', 1), ('surrounding', 1)]
When we first invoke FreqDist, we pass the name of the text as an argument. The expression most_common(10) gives us a list of the ten most frequently occurring types in the text (in this case, it will print all the words in the sentence since the length of the sentence is less than 10).
word_tokens = word_tokenize(sentence)

filtered_sentence = [w for w in word_tokens if not w.lower() in stop_words]

print("Original sentence: ")  
print(word_tokens)
print("Filtered sentence: ")
print(filtered_sentence)
Original sentence: 
['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement', '.', 'The', 'seas', 'continue', 'to', 'lash', 'the', 'coasts', ',', 'doing', 'millions', 'of', 'dollars', 'worth', 'of', 'damage', 'and', 'imperilling', 'all', 'shipping', 'within', 'one', 'hundred', 'miles', 'of', 'this', 'port', '.']
Filtered sentence:
['blizzard', 'held', 'New', 'York', 'surrounding', 'country', 'grip', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement', '.', 'seas', 'continue', 'lash', 'coasts', ',', 'millions', 'dollars', 'worth', 'damage', 'imperilling', 'shipping', 'within', 'one', 'hundred', 'miles', 'port', '.']
As you can observe, NLP tasks required the same preprocessing work as the indexing processes we have studied in section 2 of this OER. In some platforms, like the NewsEye platform, it is also possible to download preprocessed text.