orderanalyzer (version 1.0.0)
Extracting Order Position Tables from PDF-Based Order Documents
Description
Functions for extracting text and tables from
PDF-based order documents. It provides an n-gram-based approach for identifying
the language of an order document. It furthermore uses R-package 'pdftools' to
extract the text from an order document. In the case that the PDF document is
only including an image (because it is scanned document), R package 'tesseract'
is used for OCR. Furthermore, the package provides functionality for identifying
and extracting order position tables in order documents based on a clustering approach.