3.1 Document Understanding and Information Extract : Overview and Programming Basics

3.1.1 Overview

Document understanding (DU) is the ability of a system to process documents automatically. It includes technologies that can interpret and extract text and meaning from a wide range of document types, including structured, semi-structured and unstructured that can interpret and extract text and meaning from a wide range of document types including structured, semi-structured and unstructured.

DU models take in documents and segment pages of documents into useful parts (i.e. regions corresponding to a specific table or property), often using optical character recognition (OCR) with some level of document layout analysis. These methods use this information to understand the contents of the document at large, e.g. that this region or bounding box corresponds to an address or a news article.

Some examples of DU topics are the following research tasks:

  • Document Layout Analysis (DLA) — A computer-vision-based document layout analysis module which partitions each document page into distinct content regions. This model not only delineates between relevant and irrelevant regions but also serves to categorize the type of content it identifies.

  • Optical Character Recognition (OCR) — This step aims to extract text from images or PDF files by faithfully transcribing all written text in the document.

  • Information extraction (IE) is the process of converting unstructured text into a structured database containing selected information from the text. It is an essential step in making the information content of the text usable for further processing. These models generally use the output of OCR or document layout analysis to comprehend and identify relationships between the information conveyed in the document. Usually specialized to a particular domain and task, these models provide the structure necessary to make a document machine-readable, providing utility in document understanding.

  • Document Semantic Enrichment (DSE) - This process uses the extracted data from the previous step. These annotations enrich the document metadata and provide new types of visualizations in an information retrieval context.