3.3 Extracting information : Named Entity Recognition (NER)

3.3.3 Named Entity Recognition with Python

Now, we explore the task of Named Entity Recognition (NER) sentence tagging. Tagging (or labelling) means the detection of entities in text and the correct assignment of an entity type of them. In NLP, an entity is a sequence of one or more words (tokens). Thus, the task is to tag each token in a given sentence with an appropriate tag such as Person, Location, etc.

For detecting entities with NLP programming techniques, we continue using two libraries: NLTK and spaCy. NLTK is a widely used standard Natural Language Processing (NLP) and Computational Linguistics (CL) Python library with prebuilt functions and utilities for ease of use and implementation. spaCy is an open-source software Python library for advanced natural language processing that covers multiple NLP tasks (part-of-speech tagging, named entity recognition, etc.).

Using NLTK

To use NLTK for NER recognition, you will have to install the NLTK library and some components :

# Use "!" if you're using Jupyter Notebook environment.
 
!pip install nltk
!python -m nltk.downloader punkt
!python -m nltk.downloader averaged_perceptron_tagger
Then use NLTK in a list of sentences preprocessed beforehand. Refer to section 3.2.1 for the preprocessing.

1
2
3
4
5
6
for sentence in sentences:
    print("Sentence:", sentence)
    words = word_tokenize(sentence)
    pos_tags_sentence = pos_tag(words)
    ne_chunks_sentence = ne_chunk(pos_tags_sentence)
    print(ne_chunks_sentence)

This code loops through a list of phrases. It displays each sentence, breaks it into tokens and parses each word. Then, the ne_chunk() function recognizes entities of type PERSON, ORGANIZATION and GPE (location). As you can see, NLTK is quite limited and only supports a few types of named entities.
Sentence: Beginning in February 1965, there were 8 weeks of unbroken bombing by U.S. forces in North Vietnam.
(S
  Beginning/VBG
  in/IN
  February/NNP
  1965/CD
  ,/,
  there/EX
  were/VBD
  8/CD
  weeks/NNS
  of/IN
  unbroken/JJ
  bombing/NN
  by/IN
  (GPE U.S./NNP)
  forces/NNS
  in/IN
  (GPE North/NNP Vietnam/NNP)
  ./.)
Sentence: Over the next three years, the Unites States dropped more bombs than were dropped over Asia and Europe during World War II.
(S
  Over/IN
  the/DT
  next/JJ
  three/CD
  years/NNS
  ,/,
  the/DT
  (GPE Unites/NNP States/NNPS)
  dropped/VBD
  more/JJR
  bombs/NNS
  than/IN
  were/VBD
  dropped/VBN
  over/IN
  (GPE Asia/NNP)
  and/CC
  (GPE Europe/NNP)
  during/IN
  World/NNP
  War/NNP
  II/NNP
  ./.)

Using spaCy

SpaCy is a powerful open-source Python software library for automatic language processing. SpaCy currently provides support for the following languages. We need to download a spaCy NER model for English to apply NER for an English text. We choose to download en_core_web_sm because it is the smallest one regarding download size (12 MB).

To use SpaCy, install the library in your Python environment:

!python -m spacy download en_core_web_sm

Then, for the purpose of this lesson, define two sentences :
1
2
3
phrase1 = "In 1979, after more than a century of vaccination campaigns around the planet, the World Health Organization certified that smallpox had been eradicated"
 
phrase2 = "Beginning in February 1965, there were 8 weeks of unbroken bombing by U.S. forces in North Vietnam. Over the next three years, the Unites States dropped more bombs than were dropped over Asia and Europe during World War II."

Of course, you can write something else to test the library with other examples.

Then, we load the SpaCy pre-trained model in English and print recognized entities.
1
2
3
4
5
6
7
8
9
10
import spacy
from spacy import displacy
from collections import Counter
 
nlp = spacy.load("en_core_web_sm")
 
doc = nlp(phrase1)
 
for entity in doc.ents:
    print('Entity:', entity.text, '---', 'Entity Type (tag/label):', entity.label_)

The output of this code should be as follows :
Entity: 1979 --- Entity Type (tag/label): DATE
Entity: more than a century --- Entity Type (tag/label): DATE
Entity: the World Health Organization --- Entity Type (tag/label): ORG

As you can see, SpaCy is much more accurate and efficient than NLTK.

If you wish, you can display the results in a graphic form:
1
2
3
4
5
from spacy import displacy
 
for sent in doc.sents:
    if len(sent.ents) > 0:
        displacy.render(nlp(sent.text), style='ent')

marked with colors

Modify your code to do the same thing in sentence 2. You should get a result identical to this one:
Another colored text

Of course, to automate the processing of sentences or long texts, standardized output formats exist. For example, The IOB format (short for inside, outside, beginning) is a common tagging format for tagging tokens in a chunking task in computational linguistics (ex., named entity recognition) to describe the entity boundaries. In IOB, the I- prefix before a tag indicates that the tag is inside a chunk. An O tag indicates that a token belongs to no chunk. The B-prefix before a tag indicates that the tag is the beginning of a chunk that immediately follows another chunk without O tags between them.
image with table

This is an example of the IOB format :
image with table

Another tagging scheme is BIOES/BILOU, where ‘E’ and ‘L’ denotes the Last or Ending token in such a sequence and ‘S’ denotes a Single element or ‘U’ Unit element.
Image with table

The examples proposed function, even if they are straightforward. You can quickly adapt them to handle longer texts and exploit the different output formats for other tasks. In the next lesson, we will see examples of how to exploit the extracted named entities.