orderanalyzer (version 1.0.0)

Extracting Order Position Tables from PDF-Based Order Documents

Description

Functions for extracting text and tables from PDF-based order documents. It provides an n-gram-based approach for identifying the language of an order document. It furthermore uses R-package 'pdftools' to extract the text from an order document. In the case that the PDF document is only including an image (because it is scanned document), R package 'tesseract' is used for OCR. Furthermore, the package provides functionality for identifying and extracting order position tables in order documents based on a clustering approach.

orderanalyzer (version 1.0.0)

Extracting Order Position Tables from PDF-Based Order Documents

Description

Copy Link

Version

Install

Monthly Downloads

Version

License

Maintainer

Last Published

Functions in orderanalyzer (1.0.0)