Full-text digitization

Full text in the formats
your systems expect

Searchable, reusable full text from your existing scans and digitized collections - for libraries, archives, service bureaus and publishers. Delivered in the formats your systems expect, up to METS/ALTO, with contractually agreed, sample-tested accuracy. Processing exclusively in the EU.

ALTO v4 · PAGE-XML · TEI · JATS Coordinates + per-word confidence DFG-guideline acceptance testing Processed in the EU

Services by buyer

Libraries and archives

  • Full text as ALTO XML (v4), embedded as fileGrp USE="FULLTEXT" per the METS application profile - enabling full-text search in the DFG-Viewer. For prints from 1850, full text is mandatory under the DFG Practical Guidelines.
  • Integration with Kitodo.Production workflows: METS/MODS import, write-back of full-text references, stable identifiers; ingestion also directly via IIIF/OAI-PMH.
  • Delivery packages fit for long-term preservation (METS/MODS/ALTO, checksums, schema-valid) per binding specifications such as SLUBArchiv.digital.
  • Full text and underlying XML handed over without rights reservation - CC0- and CC-BY-ready, as state programs and the guidelines expect.

Digitization bureaus

  • White-label component inside your existing workflow: REST API or batch handover, input TIFF/JPEG/JP2/PDF - ideally 300 ppi and up, older legacy scans and microfilm too - from a single volume to large-scale projects.
  • Per-word confidence values enable confidence-driven post-correction: targeted proofreading instead of full review - your correction hours drop, your margin stays yours.
  • Layout analysis and segmentation: multi-column newspaper pages, marginalia, tables, footnote apparatus, reading order, article segmentation.
  • Fixed per-page cost, contractually agreed target accuracy.

Publishers with backfiles

  • From scan PDFs to real full text: article-level splitting - including reviews and miscellanea - with per-article metadata, ready for Crossref DOI registration under your prefix.
  • Full-text XML as JATS (articles) and BITS (books), plus searchable PDFs with a correct text layer - platform-ready for your eLibrary; your platform stays.
  • Accessibility under the German BFSG: tagged PDFs (PDF/UA) and EPUB 3 - a scan-image PDF cannot meet the BFSG accessibility requirements; real full text is the precondition.
  • Metadata feeds as MARC 21, ONIX, CSV and KBART; archive packages for Portico/CLOCKSS.

Source material

  • Fraktur, antiqua and mixed typefaces - the antiqua passages inside Fraktur text are captured as markup, not swallowed.
  • Newspapers and journals - multi-column layouts, marginalia, tables, footnote apparatus, review sections with italicized work titles.
  • Microfilm scans - including low-contrast, stained or edge-damaged film (16/35 mm).
  • Multilingual sources - German, English, French and further languages in the same volume; accents, special characters, Greek and Hebrew passages.
  • Special cases - genres with complex typography and multilingual settings, such as funeral sermons (Leichenpredigten), which standard mass OCR is not built for.

Our core business is printed text, including Fraktur. Handwritten material (Kurrent, Sütterlin) we assess on request via free test pages - just include a sample page.

Delivery formats and standards

ALTO XML v4

Line and word coordinates, per-word confidence values; embeddable in METS (fileGrp USE="FULLTEXT") for the DFG-Viewer and presentation systems.

PAGE-XML

Per the OCR-D specification - compatible with ground-truth workflows, evaluation tools and post-correction environments.

TEI P5

On request DTABf-oriented, with typographic markup: <hi rendition="#i"> (italic), #b (bold), #aq (antiqua passages), #g (letter-spaced).

JATS / BITS

Article and book XML for publisher platforms, with per-article metadata down to the individual review.

PDF/A · PDF/UA · EPUB 3

Searchable PDFs with a correct text layer; accessible, tagged versions (PDF/UA, EPUB Accessibility) for offerings subject to the BFSG.

Encoding and validity

Unicode (UTF-8 without BOM), special characters per MUFI/OCR-D; every XML delivery schema-valid. Optionally faithful to the source with long ſ and ligatures (ground-truth level 2) or normalized (level 1).

Quality and acceptance

  • Measured, not asserted. Character accuracy is measured per material class and contractually agreed as the target; acceptance testing follows the sampling procedure of the DFG Practical Guidelines on Digitisation (2022 edition).
  • CER/WER per the OCR-D evaluation specification. On request we evaluate against your own ground truth - Unicode-NFC-normalized, compatible with the usual evaluation tools.
  • Faithful to the source, no silent normalization. Historical orthography, printing errors and punctuation stay as they stand; nothing is modernized.
  • Illegible passages are marked, never invented. Whatever the scan does not support appears as <gap reason="illegible"> or [illegible] in the text - never as a plausible guess.
  • Line-bound and image-referenced. Every text line stays tied to the source via its coordinates and is machine-checkable - no omissions, no unsupported additions.
  • Raw and final version. You receive the unedited first pass and the final version; every correction stays traceable. Processing is frozen per project - the same page yields the same output, even years later.

Work sample

A journal page from 1904, from microfilm - the source on the left, the captured full text on the right: italic work titles preserved as markup, French passages with correct accents.

Microfilm scan of a 1904 journal page with italic book titles and French quotations
Source: microfilm scan, 1904 volume (public domain), with visible film damage along the left edge.

meanwhile rendered a real service to historical scholarship by his careful comparison of the Viennese texts with those already published by the two imperial commissions and by MM. Lecestre and Brotonne. […]

Correspondence, No. 7745. - Ayez soin d'envoyer par votre courrier des numéros du Moniteur depuis quinze jours, soit à Berlin, soit à Saint-Pétersbourg.

Viennese Text. - Ayez soin d'envoyer par vos courriers, soit à Berlin, soit à Saint-Pétersbourg, des exemplaires du 'Moniteur' depuis 15 jours. […]

The Corrispondenza inedita dei Cardinali Consalvi e Pacca (Torino : Unione tipografico-editrice, 1903), edited by P. Ilario Rinieri, is a bulky and valuable addition to the collection of diplomatic despatches […]

Captured full text (excerpt) - italics and accents faithful to the source.

How a project runs

  1. 1Sample conversion

    Up to 50 sample pages from your holdings - your hardest volume, preferably. Free of charge and without obligation; results within 5 working days, including an error-rate measurement on a sample.

  2. 2Pilot

    A delimited part of the collection at a fixed price, with a contractually agreed accuracy target and acceptance criteria. On the buyer's side, a pilot can usually be structured below the direct-award threshold.

  3. 3Production

    Ongoing processing via API or batch handover, with agreed throughput and continuous sample checks.

  4. 4Acceptance

    Sampling procedure per the DFG Practical Guidelines, error-rate report per material class, named responsibility for rework.

We document sample-conversion and pilot results so that they can be made available to all bidders in a later tender.

Processing and location

  • German GmbH: HexWorld Solutions GmbH, registered at Amtsgericht Dresden (HRB 48068).
  • Processing exclusively in the EU. Your scans never leave the EU and are not stored after processing.
  • No model training on your collections. Your contact person is based in Germany.
  • Details on data processing: privacy policy and service-provider list.

Your contact

Alaa Mroue

Managing Director, HexWorld Solutions GmbH

+49 351 25065907 info@hexworld.eu

Request a sample conversion

Free of charge and without obligation: describe your material and we will get back to you within one working day. The sample pages - up to 50, ideally your hardest volume - are exchanged by email or transfer link; you receive the result with an error-rate measurement within 5 working days.

We use your details solely to answer this inquiry.