Full-text digitization

Full text in the formats
your systems expect

Searchable, reusable full text from your existing scans and digitized collections - for libraries, archives, service bureaus and publishers. Delivered in the formats your systems expect, up to METS/ALTO, with contractually agreed, sample-tested accuracy. Processing exclusively in the EU.

ALTO v4 · PAGE-XML · TEI · JATSCoordinates + per-word confidenceDFG-guideline acceptance testingProcessed in the EU

Services by buyer

Libraries and archives

  • Full text as ALTO XML (v4), embedded as fileGrp USE="FULLTEXT" per the METS application profile - enabling full-text search in the DFG-Viewer. For prints from 1850 onwards, full text is mandatory under the DFG Practical Guidelines.
  • Integration with Kitodo.Production workflows: METS/MODS import, write-back of full-text references, stable identifiers; ingestion also directly via IIIF/OAI-PMH.
  • Delivery packages fit for long-term preservation (METS/MODS/ALTO, checksums, schema-valid) per binding specifications such as SLUBArchiv.digital.
  • Full text and underlying XML handed over without rights reservation - CC0- and CC-BY-ready, as state programs and the guidelines expect.

Read more

Digitization bureaus

  • White-label component inside your existing workflow: batch handover by transfer link or physical media, input TIFF/JPEG/JP2/PDF - ideally 300 ppi and up, older legacy scans and microfilm too. We work collection by collection: one volume year, one series, one delimited holding, with throughput agreed per project rather than an open capacity commitment.
  • Per-word confidence values enable confidence-driven post-correction: targeted proofreading instead of full review - your correction hours drop, your margin stays yours.
  • Layout analysis and segmentation: multi-column newspaper pages, marginalia, tables, footnote apparatus, reading order, article segmentation.
  • Fixed per-page cost, contractually agreed target accuracy.

Read more

Publishers with backfiles

  • From scan PDFs to real full text: article-level splitting - including reviews and miscellanea - with per-article metadata, ready for Crossref DOI registration under your prefix.
  • Full-text XML as JATS (articles) and BITS (books), plus searchable PDFs with a correct text layer - platform-ready for your eLibrary; your platform stays.
  • Accessibility under the German BFSG: tagged PDFs (PDF/UA) and EPUB 3 - a scan-image PDF cannot meet the BFSG accessibility requirements; real full text is the precondition.
  • Metadata feeds as MARC 21, ONIX, CSV and KBART; archive packages for Portico/CLOCKSS.

Read more

Source material

  • Fraktur, antiqua and mixed typefaces - the antiqua passages inside Fraktur text are captured as markup, not swallowed.
  • Newspapers and journals - multi-column layouts, marginalia, tables, footnote apparatus, review sections with italicized work titles.
  • Microfilm scans - including low-contrast, stained or edge-damaged film (16/35 mm).
  • Multilingual sources - German, English, French and further languages in the same volume; accents, special characters, Greek and Hebrew passages.
  • Special cases - genres with complex typography and multilingual settings, such as funeral sermons (Leichenpredigten), which standard mass OCR is not built for.

Our core business is printed text, including Fraktur. Handwritten material (Kurrent, Sütterlin) we assess on request via free test pages - just include a sample page.

Free sample conversion

Up to 50 sample pages from your own holdings, ideally your hardest volume year. No obligation, result with an error-rate measurement within 5 working days.

Request a sample conversionHow we measure