AI agent index: /llms.txtFull content index for AI agents: /llms-full.txt
AI

Structured extraction from digital PDFs and scans

Define document classes, output requirements and evaluation boundaries before choosing tools.

Two Words/18 September 2026/6 min read/Extraction
Document regions and linked form fields are translated into a structured record through explicit mappings.

Before approving “PDF support,” require the team to specify what that commitment covers: digital documents or scans, text capture or structured records, tables or linked form fields. Make those distinctions part of the acceptance criteria before comparing tools.

Scope PDF and form extraction as a set of testable requirements. Select candidate methods by document class and required output, then evaluate them against the receiving workflow’s needs. Bound approval by those tests rather than extending one successful demonstration to the whole document estate.

01 — Start with the document estate

Start with the document estate

The PdfTable authors illustrate why modality belongs in the initial scope. They describe Camelot and pdfnumber as limited to extracting tables from digital PDFs, while reporting that PP-StructureV2 can handle image-based PDFs and tables in pictures. This is an author-reported capability comparison between named tools, not an independent accuracy test or a guarantee about current implementations.

Treat digital and image-based inputs as separate evaluation classes. For each, write an extraction contract identifying the required structure and receiving system. Test whether the implementation produces an acceptable output, rather than stopping at whether it can ingest the file.

  • — Define the extraction contract:
  • — Input modality: digital PDFs, image-based documents, or both, evaluated separately.
  • — Required output: plain text, table structure or linked form fields.
  • — Document variation: the document types and layout variations within scope.
  • — Evaluation status: which classes and variations have been tested, and which remain untested.
  • — Receiving-system requirements: the schema and submission conditions the output must meet.
  • — Disposition: when to submit a record, withhold it or route it as an exception.

Keep unevaluated document classes outside the initial approval. Use the contract to decide whether a newly encountered document belongs within the existing scope or requires another acceptance test.

02 — Compare table-extraction methods against the same contract

Compare table-extraction methods against the same contract

The SPARTAN paper, published on 25 March 2026, groups automatic table-detection and extraction methods into heuristic-based, hybrid and deep-learning approaches. Use these categories to organize a shortlist, not to rank performance. The cited categorization offers no quantified basis for choosing one family over another.

Examine what the team would need to maintain. For a heuristic-based candidate, ask which rules it would own and how changes would be tested. For a hybrid candidate, identify component responsibilities and failure boundaries. For a deep-learning candidate, examine how model changes would be evaluated on the target document population. Apply all three lines of questioning wherever relevant; a method label should not exempt any implementation from scrutiny.

Compare candidates on the same required table structure, evaluation documents and acceptance conditions. Record integration effort and maintenance ownership alongside results, so the selection decision addresses both output quality and the responsibility the team is taking on.

03 — Separate scanned-form recognition from understanding

Separate scanned-form recognition from understanding

FUNSD’s May 2019 dataset paper describes 199 real, fully annotated noisy scanned forms that vary widely in appearance. Its intended tasks distinguish recognizing text from assigning structure. That task coverage can inform requirements, but the dataset does not establish system performance or representativeness across form domains; detailed domain-distribution information is unavailable in its abstract.

Two Words / comparison
Proposed task-level acceptance questions
01
Text detection

Was all text required for the target record located?

02
OCR

Were required values transcribed correctly?

03
Spatial layout analysis

Were the spatial groupings needed to interpret the record preserved?

04
Entity labeling

Were recognized entities assigned the required field labels?

05
Entity linking

Were the correct entities connected to form the required relationships?

Engineering review questions organized by FUNSD’s task categories, not a FUNSD-validated protocol or prescribed execution order.

Evaluate recognition, layout and relationships separately, even when one implementation handles several tasks.

Keep task-level results alongside complete-record acceptance. Use the former to identify which responsibility needs attention and the latter to decide whether the output meets the receiving system’s requirements. Do not let success at recognition stand in for acceptance of the resulting record.

04 — Evaluate against the workflow’s documents and risks

Evaluate against the workflow’s documents and risks

A reported score and a reference corpus answer different questions. SPARTAN reports performance for a stated evaluation; NIST Special Database 2 defines a structured-form population. Assess each against the document classes and tasks in your extraction contract.

Two Words / evidence summary
Two kinds of evaluation evidence
01
SPARTAN: reported performancePrecision 0.94; recall 0.91; F1 0.93; OCR character accuracy 96.7%

March 2026 paper: more than 20,000 pages of PCN-480 product-change and product-discontinuance notifications, scientific papers, certificates and datasheets.

02
NIST SD2: corpus composition5,590 images; 900 simulated tax submissions

Synthetic binary black-and-white images based on 12 US IRS 1040 Package X forms from 1988. Dataset counts, not accuracy results.

Neither establishes accuracy for a contemporary mixed-PDF workflow or provides a common cross-system benchmark.

Keep reported performance separate from dataset composition.

SPARTAN’s figures apply only to its reported evaluation. Full methodology and error analysis are unavailable in the reported summary assessed here, so treat the scores as bounded findings rather than a basis for predicting local performance. In particular, do not translate OCR character accuracy into a structured-record success rate. For NIST SD2, assess whether historical, synthetic US tax forms match the purpose of the intended test.

  • — For a local evaluation, require the team to:
  • — Name the document classes and variations included, with untested classes explicitly excluded.
  • — Define the scored unit and success condition for each task, separating character recognition from structured-record acceptance.
  • — Report results by document class and task, not only as a pooled score.
  • — Review failures against downstream consequences and agree acceptance criteria with the operational owner.
  • — Document the test procedure so it can be repeated when the implementation or supported population changes.

Base selection on results for the intended workflow, with unresolved coverage gaps visible alongside the scores. Use published evidence to shape the test, not replace it.

05 — Make the production decision explicit

Make the production decision explicit

Before approving production use, turn the extraction contract and evaluation results into an operating decision. The following is proposed engineering guidance, not a control design validated by the cited studies. Require a decision record covering:

  • — Target schema: required outputs and how missing or unresolved values will be represented.
  • — Supported scope: approved document classes and the route for inputs outside that scope.
  • — Acceptance criteria: the evaluation set, task-level requirements and conditions that block release.
  • — Exception handling: what happens when no acceptable record can be produced, including whether processing stops or enters an agreed review path.
  • — Integration boundary: what the receiving system accepts and which failures must prevent submission.
  • — Ownership: responsibility for classification, extraction changes, evaluation maintenance and operational exceptions.

If the design relies on schema validation, confidence thresholds or human review, justify and test those controls separately. Do not credit them with an assumed accuracy improvement or treat extraction scores as evidence of their reliability.

Approve a defined extraction capability: named document classes, required outputs, acceptance evidence and accountable owners. Treat every expansion of that scope as a new acceptance decision. That is a commitment the team can assess and maintain, rather than an open-ended promise of “PDF support.”

Keep reading

All posts
Start here

Some of our best projects started with a two-line email.

Most of our work starts with a conversation. No deck required.

Start yours