Illustration of a scanned geology report passing through text extraction into a ground truth dataset: the extracted text flags merged two-column pages, an invisible page break made visible as a page marker, OCR damage such as the letter l in place of the digit 1, and replaced names, and each dataset entry cites its source page
The benchmark never sees the PDF, only the text extracted from it

TL;DR

  • To test an AI on old PDF reports, the PDFs must first become plain text. The AI and the test’s answer key both read that text, so any conversion mistake ends up in the test.
  • We found four common mistakes: jumbled two-column pages, missing page breaks, scanning errors that swap letters and digits, and files broken by hiding names.
  • Our examples come from Groundtruth, Eigenform’s open-source benchmark for geology. A script checks every reference in it; on one build it caught 41 errors.

Key Takeaways

  • A PDF you can search is not always one a computer can read well. We scored 190 reports for text quality before choosing any.
  • Tables and two-column pages need different handling. One conversion method keeps tables intact but jumbles columns; the other does the reverse.
  • Scanning errors hide inside the text. In one 1920s report, 18% of the pages were too damaged to use.
  • Even cleaning the data can break it. Replacing names with pseudonyms made 207 files unreadable to software.

Every claim in a rubric for Groundtruth, Eigenform’s benchmark for geological reasoning, carries an evidence locator: a pointer to the passage that supports it. Together, those locators make up the benchmark’s ground truth dataset. Neither the rubric nor the candidate model reads the PDF; both read the extracted text. So the extraction step decides what the benchmark can ask, what counts as evidence, and whether a correct answer can be found at all.

We measured these traps on three corpora: NI 43-101 technical reports, four 1920s USGS district studies, and WAMEX exploration reports from Western Australia. The same traps apply to any ground truth machine learning or ground truth deep learning project built on scanned documents. For the wider picture, see our guide to ground truth data for AI.

Measure text quality before building a ground truth dataset

For the USGS corpus we downloaded 190 candidate reports. Instead of opening 190 scans, we scored every extraction. The script records three signals per report:

  • words per page, where fewer than 100 suggests an image-only scan with no usable text layer;
  • the share of junk tokens, such as words with no vowel or a capital letter in the middle, which is typical OCR noise;
  • the number of corrupted numerals, digit strings containing l, O or I, such as “41.9l per cent”.

All 190 reports had a usable text layer, just over eight million words in total. Ranking them by noise produced the shortlist. A metadata field saying a PDF is searchable tells you none of this. Measure legibility on the text you will actually serve.

PDF text extraction for AI: single-column and two-column pages

PDF text extraction for AI has one tool doing two opposite jobs. The page layout that makes tables readable also destroys prose set in two columns.

Single-column pages and tables: keep the layout

pdftotext -layout keeps each character in its visual position. That keeps assay and resource tables readable. Without it, columns of numbers collapse into a stream.

Two-column pages: follow reading order

On a two-column page, the same behaviour is destructive. Every output line joins a fragment of the left column to a fragment of the right. Those pages need pdftotext -raw, which follows reading order.

How to detect a column merge

Looking for a gutter is unreliable, because a merged line can join the two columns with a single space. Line length works better. Prose from a single column has a line-length distribution with one peak. A two-column merge leaves a second peak at roughly double that width.

The main report in the USGS set showed both outcomes in one document. Its line-length histogram peaked at 50 to 59 characters with no second peak, so the prose had reflowed correctly. Its tables had not survived, so no question takes a tabulated value from that report. Eight of its pages were still column-merged. One page was clean in its upper half and merged from about line 135 downward. A page-level check alone would have passed that page.

The rule we took from this is blunt. Where extraction destroyed a structure, that whole class of evidence in that document is unusable, however clearly the surrounding prose reads.

Make page breaks visible

pdftotext separates pages with a form-feed character. grep does not display it, yet every locator needs to cite a page. The extraction step therefore rewrites each form feed as a visible marker, === page N ===. It then checks the resulting page count against pdfinfo. Sometimes the page number printed on the scan differs from the PDF page. In that case locators cite the marker, because a script and a candidate can both see it.

OCR damage in a ground truth dataset

Scanned reports often come with a text layer produced by somebody else’s OCR. That damage becomes part of the ground truth dataset unless it is found and fenced off.

OCR errors live in the text layer

Re-reading that text verifies nothing, because the error is in the text. Checking a damaged string means rendering the page image and reading it.

The scale of the damage is easy to underestimate. The USGS report above arrived with a note listing 24 bad numerals. We then measured the out-of-vocabulary word rate per page. By that measure, 75 of its 419 text pages (18%) exceeded our threshold. Each of the other three reports had only 3 to 4%. Whole words were wrong as well as digits. On one page, OCR swapped several ordinary words for unrelated ones, so the sentence says something the authors never wrote. Pages above the threshold were quarantined, and no anchor may point into them.

Counting damage without counting identifiers

Counting damage needs care too. In the WAMEX archive, digit-dominant tokens with a letter-for-digit confusion numbered 4,172. A naive count of every token mixing letters and digits came out about twenty times larger. Drill hole and sample identifiers legitimately contain both, so that count measured the identifiers, not the damage.

Rules for load-bearing figures

Two rules came out of this work. First, no question rests a load-bearing claim on a damaged string. For example, a production figure with a lowercase l in place of the digit 1 is excluded from scoring. The relevant rubric says so in its do-not-credit list. Second, we check every load-bearing grade or intercept against a second occurrence in the corpus. If there is none, the figure is dropped.

Damage that is not OCR

The WAMEX corpus is stored as JSON chunk files. When the corpus was prepared, names in it were replaced with pseudonyms. Where a replacement name followed a newline, the substitution turned the \n into an invalid escape sequence. As a result, 207 chunk files across 46 reports no longer parse as JSON.

A retrieval layer that loads the corpus as JSON cannot see those chunks. A candidate running grep over the raw files can. A question anchored in one of them would partly measure which tool the candidate happened to use. One item’s anchor was in exactly that position, and we re-anchored it on complete passages.

Verifying evidence anchors in a ground truth dataset

An anchor is only useful if the text it points to is really there. Groundtruth’s authoring flow checks each locator by script, for every item. The locator must be a verbatim substring of the extracted text. This check is one of the scripted gates described in Generating a 50-Question Benchmark With One Approval Step. One build needed 41 anchor corrections found this way. Three of them were locators on the wrong page, an error that re-reading would not have surfaced.

Normalise before comparing

The verifier has to normalise whatever extraction introduced before it compares strings. That includes soft hyphens (U+00AD), words hyphenated across line breaks, ligatures, curly quotes, non-breaking spaces, runs of whitespace, and running heads. Without normalisation it reports false failures. Worse, it invites “corrections” to anchors that were already right.

Span audit

A substring match misses two things, and two further audits catch them. The first is the span audit. The most common anchor error is a claim whose sentence crosses a page boundary. A running head can also land in the middle of that sentence.

Retrievability

The second is retrievability. A fact stated only in a split or garbled passage cannot be found by an agent that searches. That makes the item harder than its label says. The main USGS report has a median raw line of 52 characters, so searching for a full sentence fails there. Each evidence record in that benchmark therefore carries a short search handle: an exact substring of a single raw line. 313 of its 430 locators match verbatim. The other 117 are claims spanning several sentences, each with a verified handle.

All 308 anchors in the Technical benchmark match verbatim. The scores those anchors support are on the Groundtruth leaderboard. Anchors sit inside single lines so a candidate’s grep returns them. Changing the extraction moves every line boundary, so anchor verification is re-run whenever a corpus is re-extracted.

Why figures stay out of a text ground truth dataset

A rubric author can open the PDF and look at a geological map, a cross-section or a drill plan. The candidate reads text. So no Groundtruth question may require reading a map, section, stratigraphic column, drill plan or photograph. Every item is verified against the text files instead of the PDF. In scanned reports, plates and figures show up as sparse pages with little more than a caption. That makes them easy to count and exclude.

That restriction is one reason maps are a separate problem for us, handled by the Geological Map Processing Suite.

Keep navigation aids out of the evidence

Each corpus gets an INDEX.md alongside the text. It covers document structure, metadata, known traps and attribution hazards, with no geological claims. An index that summarised the geology would become a low-authority source that authors and candidates both lean on. Generated summaries and filenames never count as evidence in a Groundtruth rubric.

Age matters as well. For the USGS set, the corpus as written is the authority for what it claims. No question asks a candidate to correct a 1920s source against present-day understanding. Sometimes a report reasons its way to a conclusion that was later revised. In that case, the question tests the reasoning as the report gives it.

The corpus preparation steps are written up in AUTHORING.md in the Groundtruth repository. For how the questions built on these corpora are generated and graded, see Dynamic Benchmarking: How We Did It.

FAQs

What is a ground truth dataset?

A ground truth dataset is the verified reference that a model’s outputs are compared against. In a source-grounded benchmark, it holds questions, reference answers and evidence locators. Each locator points to an exact passage in a document corpus. The dataset is only as good as those passages: they must really say what the reference claims.

How do you extract text from PDFs for machine learning?

Use a text extractor such as pdftotext and choose the mode per page. Layout mode suits tables and single-column pages; raw mode suits two-column prose. Mark page breaks visibly. Measure words per page and junk-token rates before trusting a source, and render page images to check any string the OCR may have damaged.

How do OCR errors affect ground truth in machine learning?

OCR errors sit in the text layer itself, so anything built on that text inherits them. A wrong digit becomes a wrong reference value. A garbled sentence becomes evidence nobody can find by search. Ground truth in machine learning is reliable only when damaged pages are measured, quarantined and kept out of scoring.

What is an evidence anchor in an AI benchmark?

An evidence anchor ties a claim in a grading rubric to the exact passage that supports it, usually a page marker plus a verbatim string. A script checks every anchor as an exact substring of the extracted text. That check catches wrong pages and broken quotes that human review misses.