OCR Tools
Transkribus, Abbyy, tesseract/ocropus. Possibly Transkribus (https://transkribus.eu/Transkribus/) Issues of characters/font styles not handled by mainstream engines; trainability; maybe postcorrection, crowd postcorrection etc.
Which OCR?
OCR products can be divided into two main types:
- 'omnifont' engines, which know about many existing typefaces, but are difficult to adapt to a new one; and the
- 'trainable' egines, which can learn to recognize new fonts, but may be very hard to teach.
There are four main OCR products available that you could consider for your project:
- ABBYY FineReader (commercial, omnifont)
- Omnipage (commercial, omnifont)
- Tesseract (open source, trainable)
- Ocropus (open source, trainable)
If you have a script (for instance, Glagolitic) which is not supported by the omnifont engines, trainable OCR is the only option, but the process is far from obvious.
It is beyond the scope of this course to teach you how to train an OCR engine, but there are resources out there that could help you in this area:
- information on training Tesseract in the EMOP project with the Franken+ tool can be found at: http://emop.tamu.edu/node/47, http://emop.tamu.edu/node/54.
- More information on the Franken+ tool is at: http://dh-emopweb.tamu.edu/Franken+/.
- There is also a paper on this subject: Early Modern OCR Project (eMOP) at Texas A&M University: Using Aletheia to Train Tesseract , http://dl.acm.org/citation.cfm?id=2494304
For a discussion of the differences between (limited) training with an omnifont engine and extensive font training with Tesseract, you could also consult this study: http://lib.psnc.pl/dlibra/doccontent?id=358.