3.1 Document Understanding and Information Extract : Overview and Programming Basics
3.1.3 Text Pre-processing : Programming Basics
In this lesson, we will study the basics of text manipulation with the computer language Python. You don't need to know Python, but you should be familiar with the Python environment and Jupyter Notebook if you want to run the example code yourself.
This section of the course required some prior Python knowledge. The OER1 will provide you with all the knowledge you need to get started with Python.
The text examples we use in this lesson are from historical newspapers extracted from the NewsEye Platform. You can easily adapt the examples with other texts if you wish.
To practice and test the code examples provided here, you can use online Jupyter Notebook services such as Google Colab for example. In such services, if you need to install any code library, you can do it with !pip install library_you_need in any code block.
Programming basics
A text is nothing more than a sequence of words and punctuation. In the following example, sentence followed by the = sign, and then some quoted words, separated with commas and surrounded with brackets, which is known as a list in Python. In the following example, sentence and sentence_list represent the same text.
sentence = "The blizzard which has held New York and the surrounding country in its grip for the past forty-eight hours continues without abatement"
sentence_list = ['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip',
'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement']
Next, we can print the contents of the sentence and ask for its length to find how many words are in the sentence:
print(sentence_list)
len(sentence_list)
['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement'] 22For the purpose of this exercise, we define two sentences. Adding two lists creates a new list with everything from the first list, followed by everything from the second list:
sentence1 = ['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding']
sentence2 = ['country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues',
'without', 'abatement']
sentence = sentence1 + sentence2
print(sentence)
['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement']
Python lists offer many useful methods. For example, when we append() to a list, the list itself is updated as a result of the operation.
sentence.append(".")
print(sentence)
['The', 'blizzard', 'which', 'has', 'held', 'New', 'York', 'and', 'the', 'surrounding', 'country', 'in', 'its', 'grip', 'for', 'the', 'past', 'forty-eight', 'hours', 'continues', 'without', 'abatement', '.']
The advantages of manipulating text in list form are that: - we can identify the elements of a Python list by their order of occurrence in the list, - we can also extract a slice or sublist from sentence.
print(sentence[0])
print(sentence[1])
print(sentence[1:5]) # extract a slice
The blizzard ['blizzard', 'which', 'has', 'held']
Any individual word in the sentence list is a word of type String (short str):
print(sentence[1])
print(type(sentence[1]))
blizzard <class 'str'>
Important takeaway: In most programming languages, Strings are implemented as sequences (lists, arrays) of characters that act like a single object. Indeed, manipulating sequences is generally simpler and more efficient. For Text Mining methods, this is also the case. It is, therefore, necessary to understand the basic techniques of list manipulation and how to transform the text into a list.
RegEx Basics
Regular expressions or RegEx is a sequence of characters mainly used to find or replace patterns embedded in the text. Regex is a very powerful tool in programming that is frequently used. They are useful in many domains and essential in data science. Here is the list of patterns we can use to match expressions in text :
-
\b returns a match where the specified pattern is at the beginning or at the end of a word.
-
\d returns a match where the string contains digits (numbers from 0-9).
-
\D returns a match where the string does not contain any digit. It is basically the opposite of
-
\w helps in the extraction of alphanumeric characters only (characters from a to Z, digits from 0-9, and the underscore _ character)
-
\W returns a match at every non-alphanumeric character. Basically opposite of (.)
-
(.) matches any character (except a newline character)
-
(^) checks whether the string starts with the given pattern or not.
-
($) checks whether the string ends with the given pattern or not.
-
(*) matches for zero or more occurrences of the pattern to the left of it
-
(+) matches one or more occurrences of the pattern to the left of it
-
(?) matches zero or one occurrence of the pattern left to it.
-
(|) checks whether any of the two patterns, to its left and right, is present in the text or not.
Like many other languages, Python has a built-in module to work with regular expressions called re. Some common methods from this module are:
- re.match() - re.search() - re.findall()
Here is an example of how it works :
-
The
re.match(pattern, string)the function returns a match object on success and none on failure.
-
import re
result = re.match('the blizzard', 'the blizzard which has held New York and the surrounding country ...')
result.span()
(0, 12)
-
re.search(pattern, string)matches the first occurrence of a pattern in the entire string (and not just at the beginning).
result = re.search('country', 'the blizzard which has held New York and the surrounding country ...')
print(result.group())
country
-
re.findall(pattern, string)will return all the occurrences of the pattern from the string. It is recommended to use re.findall() always. It can work like bothre.search()andre.match().
result = re.findall('the','the blizzard which has held New York and the surrounding country ...')
print(result)
['the', 'the']
![]()
By using Python or one of the numerous online regex tester, try to find all the occurrences of the word “damage” in the following text (extracted from historical newspapers).
(BT SPECIAL CABLE TO THE HERALD.) New Tonk, Friday.—The blizzard ghich has hel1 New Fork and the suricunding country in its grip for tho past fortz-eight hours coütinnes withoüt abatement. The seas continng to lash thie cousts, doing millions of dollars worth of damage and imperilling all shipping within one huindred miles of this port. The Dominion liner Princess Anno, formerly the Norfolk, which was due here Wednesday, is ashore at Rockaway and has sent out S.O.S. signals. Tugs are being rushed to take off thirty-two passengers and the crew of seventy-two. Heavy scas are hammering the vessel and the work o rescue will be difficult. No other calls for help have been received from seaward, but all vessels in their midnight reports told of heavy seas, gales and driving snow which made progress diflicult. The scas continue to hammer Coney and Rockaway with increasing damage. New Fork business houses opened late to-day due to the impossibility of employés reaching their appeinted places on account of the interruption of street traffic. Commuters from nearby towns were much later in reaching the city and many did not make the attempt