OCR Tools
Transkribus, Abbyy, tesseract/ocropus. Possibly Transkribus (https://transkribus.eu/Transkribus/) Issues of characters/font styles not handled by mainstream engines; trainability; maybe postcorrection, crowd postcorrection etc.
To OCR or not to OCR?

As we have learned in the previous lesson, text capture is an essential step in the digitization workflow. We need to capture the text in scanned or photographed images in order to create searchable editions of legacy dictionaries. OCR is one option that we have for text capture.
Before you invest time and effort into OCRing your dictionary, you should consider:
- whether OCR is at all feasible
- which OCR program should you use: trainable or omnifont?
- what should be the output format of your OCR?
Some texts are totally un-ocr-able, while some can in principle be recognized, but the quality will be too bad to make it worthwhile to even try.

Fortunately, there are also numerous dictionaries, especially the more modern ones, which can be handled well by OCR.

For the dictionaries that are good candidates for OCR, the particular requirements of the project will influence the decision whether to OCR or not, and what OCR engine to use.
If you are planning to put the scanned images of your dictionary online and use the underlying OCR for full-text search, a less than perfect text capture might not be an issue. If, however, you are planning to implement extensive searching in an XML-encoded dictionary, you will need good quality OCR, with post-correction.