Skip to main content
logo
Shibboleth Login   Login as Guest  
Forgotten your username or password?
HomeContact
About
  • The #dariahTeach Project
  • The IGNITE Project
Testimonials
  • Introduction
  • Unit I: Introduction to historical newspaper and the NewsEye platform
  • Unit II: Searching with the NewsEye Platform
  • Unit III: Information extraction and document understanding
  • Unit IV: Discover and use a corpus of historical newspapers
  • Unit II: Searching with the NewsEye Platform

    • 2.1. Presentation of the NewsEye Platform Lesson
    • 2.2. Understanding how the platform works Lesson
    • Results page overview Interactive Content
    • Datasets show page overview Interactive Content
    • Use case SERP Interactive Content
    • Compound mode Interactive Content
    • Exemple of a processed text Interactive Content
    • Quality of OCR extracted text Interactive Content
    • Example of XML file generated by OCR process Interactive Content
    • Matrix of our collection Interactive Content
◄Unit I: Introduction to historical newspaper and the NewsEye platformUnit III: Information extraction and document understanding►
Skip Course Custom Menu
Course Custom Menu
Introduction
  What is Historical Newspapers Analysis and Why Bother?
  A Dimpah OER
  Instructions
Unit I: Introduction to historical newspaper and the NewsEye platform
  1.1 Introduction to Historical Newspaper Analysis and Digitisation
  • 1.1.1 Historical newspapers: an infinite source of data
  • 1.1.2 What is NewsEye?
  • 1.1.3 From paper to digital
  • 1.1.4 From data to information: overview
  original newspaper vs OCR text
Unit II: Searching with the NewsEye Platform
  2.1. Presentation of the NewsEye Platform
  • 2.1.1. The search page
  • 2.1.2. The search result page
  • 2.1.3. The document show page
  • 2.1.4. The datasets pages
  • 2.1.5. Creating a dataset
  2.2. Understanding how the platform works
  • 2.2.1. Demonstration of blackboxes
  • 2.2.2. Indexing textual content
  • 2.2.3. Opening the blackbox : the ranking problem
  Results page overview
  Datasets show page overview
  Use case SERP
  Compound mode
  Exemple of a processed text
  Quality of OCR extracted text
  Example of XML file generated by OCR process
  Matrix of our collection
Unit III: Information extraction and document understanding
  3.1 Document Understanding and Information Extract : Overview and Programming Basics
  • 3.1.1 Overview
  • 3.1.2 Information Extraction History
  • 3.1.3 Text Pre-processing : Programming Basics
  3.2 Natural Language Processing (NLP) programming basics
  • 3.2.1 Natural Language Processing (NLP) basics
  • 3.2.2 Natural Language Processing (NLP) advanced
  3.3 Extracting information : Named Entity Recognition (NER)
  • 3.3.1 What are named entities ?
  • 3.3.2. Why are named entities important?
  • 3.3.3 Named Entity Recognition with Python
  Stopwords interactive activity
  NER interactive activity
Unit IV: Discover and use a corpus of historical newspapers
  4.1 Approaches for analysing newspaper collection
  • 4.1.1 Reading the corpus
  • 4.1.2 Exploring data using wordclouds
  • 4.1.3 Focusing on meaningful: adding POS tagging
  • 4.1.4 Extracting named entities
  • 4.1.5 Mapping locations on a map
  Spanish Flu dataset
Skip Navigation
Navigation
  • Home

    • Site pages

      • Tags

      • Search

      • Calendar

    • Courses

      • Digital Research on European Historical Newspapers...

        • Participants

        • Introduction

        • Unit I: Introduction to historical newspaper and t...

        • Unit II: Searching with the NewsEye Platform

          • Lesson2.1. Presentation of the NewsEye Platform

          • Lesson2.2. Understanding how the platform works

          • Interactive ContentResults page overview

          • Interactive ContentDatasets show page overview

          • Interactive ContentUse case SERP

          • Interactive ContentCompound mode

          • Interactive ContentExemple of a processed text

          • Interactive ContentQuality of OCR extracted text

          • Interactive ContentExample of XML file generated by OCR process

          • Interactive ContentMatrix of our collection

        • Unit III: Information extraction and document unde...

        • Unit IV: Discover and use a corpus of historical n...

Skip Course Custom Fields
Course Custom Fields

Show All

Main Language

en

Other Languages

en

Publisher Name

dariahTeach

Authors

Antoine Doucet, Cyrille Suire, Axel Jean-Caurant, Nicolas Sidère

Licence

CC-BY-SA

Publication Date

2023-05-22 12:11:00

ECTS

5

Back

Guest (Log in)