Skip to content

Reading born-digital documents into the document record, without OCR

The digital-convert workflow. Tool: atrium-digital-convert. The tool was part of llm-enrich until 1 October 2026; it continues in a new repository under its own name, and the keyword extraction that llm-enrich also did moved to keyword-extract.

Scope

This page gives the stable core of the workflow: its purpose, its steps, the formats it reads and writes, and the licence floor of its output. The converter's engines, its options and the formats still being added are documented with the code — in the tool's README — because they change while the service is being built.

Purpose

A PDF or DOCX file that was never printed and scanned already carries its text; running OCR on it would only add errors. This workflow reads such a file directly and writes the ATRIUM document record from what it finds: the pages, the lines with their boxes, the blocks they belong to, the headings, running headers and footers, footnotes and tables. It also checks the text layer. A page whose layer does not decode to readable text is flagged needs_ocr, so that the orchestration can send exactly those pages to OCR; a PDF whose text layer is itself the result of an earlier OCR run is refused, because it belongs to ocr-postprocess.

At a glance

In PDF, DOCX, ODT, ODS, XLSX and RTF files, and DOC/XLS through LibreOffice; optionally the AMČR seed record, whose source.sha512 must match the file
Out the ATRIUM document record (JSON); on request the same document as Markdown; a paradata log (JSON)
Runs as an HTTP service (-api: POST /reformat, and POST /describe for a per-page assessment) and a command-line tool (-digital), since v1.1.0-beta
Compute CPU; the optional layout-model engine for complex PDFs is a local, opt-in build
Network none for /reformat and the command line; /describe calls page-classification and ocr-postprocess when their URLs are set; the optional layout-model engine fetches its models at build time
Code licence MIT
Output licence MIT with the default engine, which uses permissively licensed libraries; the optional layout-model engine adds weights under CDLA-Permissive-2.0, which leaves the output unchanged
Record none of its own yet

Steps

# Step What happens In → out Activity
1 Read the file A light engine reads the PDF or DOCX structure directly; for complex PDFs a layout-model engine can be selected instead. PDF, DOCX → IR Converting
2 Check the text layer Every page's text layer is tested and its verdict recorded as text_layer (digital, garbled, ocr, none, blank). A page without a usable layer is marked needs_ocr with its reason; a PDF whose layer is mostly an earlier OCR run is refused with a registered reason and no record. IR → IR —
3 Build the record Pages, lines with boxes and block ids, line styles (heading level, running header or footer, footnote), tables and the source's identity are written as the blocks the tool owns. IR → JSON Converting
4 Render, when asked for The record is turned into Markdown in reading order — the form in which the keyword stage shows a document to a language model. JSON → Markdown —

Provenance and licence

Every run writes a paradata log. The record's licence is computed from the components that ran, as declared in the tool's para_config.txt, and the most restrictive one wins; with the default engine that is MIT. The source document's own licence applies to its text.

In the document record the tool is one of the two originators of the positional layer — pages, content, lines and tables — for a born-digital document, as ocr-postprocess is for a scanned one; the document's source.origin decides which of the two writes, and never both. The program id in the record is digital-convert, as before the move.

Where it sits

  • First stage of the born-digital route — Pipelines → W2. Scans and images go through OCR and ocr-postprocess instead; the route step of the AMČR pipeline chooses by the file type AMČR detected.
  • The OCR hand-off. Pages flagged needs_ocr go to the ATR service. ocr-postprocess then merges the ATR ALTO of each such page back into the same record, page by page; the born-digital pages stay as the converter wrote them.
  • What follows it. The record goes on to keyword-extract, which reads the converter's lines like any other record's.

On other platforms

SSH Open Marketplace. No record of its own yet.

Galaxy. Inputs map to pdf and docx; the output to json. Galaxy's own grobid and markitdown tools convert born-digital documents but do not write the ATRIUM record. The container is the service image ghcr.io/ufal/atrium-digital-convert-api:<version> (POST /reformat), released since v1.1.1-beta.

Sources

Read from ufal/atrium-llm-enrich at release v0.8.0 (the converter as a command-line tool) and from its plan for the repository's move to ufal/atrium-digital-convert (2026-10-01). This table records provenance, not a build instruction.

Source What was taken from it
README.md §§ intro, the converter purpose, the steps, the inputs and outputs
api_util/digital_to_json.py, api_util/doc_to_visual_md.py the converter's route, its checks and its output
requirements_digital.txt, requirements_digital_docling.txt, para_config.txt the engines and the licence components
DARIAH-ERIC/atrium-galaxy-tools; bgruening/galaxytools the Galaxy analogues