Add info about text parsers and pdftotext to docs

This commit is contained in:
Roberto Rosario
2012-07-27 03:36:08 -04:00
parent bd4f25df15
commit e1bd1b6f55
3 changed files with 19 additions and 1 deletions

View File

@@ -17,5 +17,15 @@ concatenated and shown to the user. All newly uploaded documents will be
queued automatically for OCR, if this is not desired setting the :setting:`OCR_AUTOMATIC_OCR`
option to ``False`` would stop this behavior.
---------------------
Document text parsers
---------------------
When checking queued documents, **Mayan EDMS** will first try to extract
text using one of the registered parsers corresponding to the document
MIME type. Only when failing to extract any text using a parser,
**Mayan EDMS** will fallback to process the document's image representation
using the OCR engine Tesseract_ and the OCR preprosessor unpaper_.
.. _Tesseract: http://code.google.com/p/tesseract-ocr/
.. _unpaper: http://unpaper.berlios.de/