
TL;DR
- Ground truth data is the set of correct answers used to test an AI, like the answer key a teacher uses to mark an exam. If the answer key has a mistake, the AI gets marked wrong even when its answer is right.
- It is built in four steps: collect the documents, turn them into text, write the correct answers, and test the grading. Most mistakes creep in early and only show up at the end.
- Our examples come from Groundtruth, Eigenform’s open-source benchmark that tests AI models on real geology reports: 150 questions, eight models.
Key Takeaways
- A bad answer key ruins every score. A wrong reference answer marks a correct model wrong, every time.
- Turning documents into text quietly adds mistakes. In one 1920s report we used, 18% of the pages were too damaged by scanning errors to rely on.
- Test the grading, don’t just re-read it. Reading the grading rules found 2 mistakes; practice answers with known scores found 15 more.
- Every correct answer should point to its source page. A script checks each link to the source; one build needed 41 fixes.
Ground truth data for AI is the reference that decides whether an answer counts as right. Every accuracy figure, leaderboard position and “the new model is better” claim is a comparison against it, so its errors pass straight into every result built on it.
We learned this building Groundtruth, Eigenform’s benchmark that asks AI models geological questions about real exploration reports. The hard part was never the scoring code. It was making sure the reference answers, and the evidence behind them, were actually true to the documents.
What is ground truth data for AI?
Ground truth data is the set of verified, correct answers that an AI model’s outputs are compared against during training or evaluation. It can be class labels, bounding boxes, transcriptions or reference answers with cited evidence. What is ground truth in machine learning, in one line? The best available record of what is actually true.
The term comes from remote sensing, where “ground truth” meant observations made on the ground to check what a satellite image seemed to show. The meaning carried over: ground truth is the check, not the thing being checked. You will also see it written as one word, as in “groundtruth machine learning”; the meaning is the same. (Our benchmark, Groundtruth, uses that spelling as its product name.) It is also only as good as the process that produced it. In Groundtruth, each reference answer is tied to an exact passage in the source reports, and a geology expert vetted every question and rubric before any model was scored.
How ground truth data for AI is built
Building ground truth data takes four steps. The table shows what each one decides and where it usually goes wrong; the sections below show what each looked like for us.
| Step | What it decides | Typical failure |
|---|---|---|
| Sourcing | Which documents or examples the reference can draw on | Material the model already memorised; illegible sources |
| Extraction | What text or data the reference is written against | OCR damage, merged columns, broken files |
| Labeling | What the correct answer is and how it is graded | Inconsistent judgment calls, rewarding the same point twice |
| Validation | Whether the reference and its grading behave as intended | Checks that only re-read, instead of testing |
Sourcing
Sourcing decides what can be asked. For our USGS document set we downloaded 190 candidate reports, just over eight million words, and ranked them by measured legibility before choosing any. We also favour obscure sources on purpose. A benchmark built from material a model saw in training measures recall, not reasoning, a problem covered in Dynamic Benchmarking: How We Did It.
Extraction
Extraction decides what the reference is written against, because both the rubric’s evidence and the candidate’s answers come from extracted text, not from the original PDF. It is where the quietest errors enter. Two-column pages merge into nonsense under one extraction mode and tables collapse under the other. OCR damage sits inside the text layer, where re-reading cannot see it. Preparing one corpus even broke 207 JSON files across 46 reports without any OCR involved. The measured traps, and the rules we adopted for each, are in From PDF to Ground Truth Dataset, our deep dive on building a ground truth dataset from scanned reports.
Labeling
Labeling decides the correct answer and how partial answers earn credit. In a benchmark of open questions, a label is a reference answer plus a rubric: required points, a hard gate that zeroes an answer missing the core fact, and a list of claims that must not be credited. We derive grading criteria from the underlying geological principle, not from the wording of the source sentence, so that a correct answer shows reasoning rather than paraphrase.
Consistency is the other half of labeling. When people reviewed questions batch by batch, inconsistent judgment calls between batches produced incompatible rubric structures. We now generate all 50 questions of a set from one blueprint, approved once by a person before any drafting, as described in Generating a 50-Question Benchmark With One Approval Step. Planning up front matters: on one build, 9 of the last 22 questions collided with earlier ones before we did.
Validation
Validation checks that the reference and its grading behave as intended, and it has to test rather than re-read. We run fabricated answers with known expected scores through the production judge before any real model is graded. Across 200 such answers on two projects, 14 and 16 scored higher than they should have, because the rubric rewarded the same idea twice. The method is in Testing the Rubric Before You Test the Model.
Scripts check the rest. Every evidence locator must be a verbatim substring of the extracted text, and the judge must return a structured breakdown that software can verify. A breakdown that fails the check is rejected and retried up to three times; after that, the question stays unscored instead of getting an invented number. The full authoring and validation procedure is public in AUTHORING.md in the Groundtruth repository.
Common ground truth errors and how they propagate
An error in ground truth data for AI does not stay in one place. It passes into every score computed against it, for every model, and it looks exactly like a model failure.
| Error | How it propagates | How we catch it |
|---|---|---|
| Damaged source text | A correct answer disagrees with a corrupted reference value | Page-level damage scores; quarantined pages; a second occurrence required for load-bearing figures |
| Wrong evidence anchor | Graders credit or penalise against the wrong passage | Script-verified verbatim anchors, span and retrievability audits |
| Double-counted rubric points | Weak answers score too high; rankings shift | Fixture answers with known scores run through the production judge |
| Leaked reference | Agents read the answer instead of reasoning | Reference files moved outside the candidate’s reach; transcript audits |
| Too few items | Noise looks like a real difference between models | Confidence intervals and a minimum detectable gap |
Two of these deserve a closer look.
Leakage is the error people least expect. A candidate agent with file access can open anything the benchmark’s account can read, and ours found the grading rubrics. Hiding only the authoring skill still leaked grading text on 8 of 50 questions. The fix and the audit are in Your AI Agent Can Read Its Own Answer Key.
Sample size is the error that looks like precision. With 50 questions per set, the smallest gap we can detect between two models is 0.27 to 0.45 points. Across eight models and 1,200 graded answers, only 2 of 7 gaps between neighbouring models were statistically significant. Good ground truth with too few items still cannot separate close models.
Ground truth in machine learning vs. deep learning contexts
The idea is the same everywhere: a trusted reference. What changes is the kind of reference, how much of it is needed, and how it is checked.
Ground truth machine learning: labels for training and testing
In classic supervised learning, ground truth machine learning data usually means labeled examples: a class for each row, a value to predict, a category for each document. The same kind of labels serves two roles. Training labels teach the model; test labels measure it. Keeping the two separate is what makes the measurement honest, and noisy training labels are often tolerated where noisy test labels are not.
Ground truth deep learning: scale, dense labels and noise
Ground truth deep learning datasets are usually far larger and denser: a mask for every pixel in segmentation, a box for every object in detection, a transcript for every second of audio. At that scale, labels come from many annotators or from models, so agreement between annotators and systematic review of samples replace checking every item by hand.
Ground truth for generative AI evaluation
Language models add a third context. There is rarely one correct wording, so machine learning ground truth becomes a reference answer plus a rubric, and the grader is often another model. That is where source anchoring and judge validation, described above, carry most of the weight. Our benchmark takes its name from that idea: Groundtruth is ground truth for reasoning over documents, where every correct answer has to trace back to the page that supports it.
| Context | Typical ground truth | Main risk | Main check |
|---|---|---|---|
| Classic machine learning | Class labels, target values | Label noise, leakage between train and test | Held-out test set, spot audits |
| Deep learning | Masks, boxes, transcripts at scale | Inconsistent annotators | Inter-annotator agreement, sampled review |
| Generative AI evaluation | Reference answers, rubrics, evidence anchors | Rubric defects, grader bias, corrupted sources | Fixture answers, verified anchors, validated judges |
What good ground truth data for AI will not tell you
Even well-built ground truth has limits. It describes one set of documents or examples, so a model that scores well on it has shown competence on that material, not in general. Across our three document sets the top tier of models held, but the order inside the lower half changed from one set to the next. The current results are on the Groundtruth leaderboard, and the reasoning behind them is in Which AI Model Is Best for Your Geology?
A reference also ages. For historical sources we treat the document as the authority for what it claims, and questions test the reasoning as the report gives it rather than asking a model to correct a 1920s author.
FAQs
What does ground truth mean in AI?
In AI, ground truth means the verified correct answer that a model’s output is compared against. It can be a label, a bounding box, a transcription or a reference answer with cited evidence. The term comes from remote sensing, where observations on the ground were used to check what satellite images seemed to show.
How is ground truth data created?
Ground truth data is created in four steps. Sourcing picks the material, extraction turns it into usable text or data, labeling defines the correct answers and how they are graded, and validation tests the result. Validation should run known test answers through the real grader rather than rely on another read-through.
Why does ground truth quality matter for evaluation?
Every evaluation score is a comparison against the ground truth, so its errors pass into every result. A wrong reference answer marks correct models wrong, and a rubric that rewards the same point twice lets weak answers score too high. No metric or larger model can correct a flawed reference set after the fact.
What is the difference between ground truth and labeled data?
Labeled data is any data with labels attached, including labels that are noisy, automatic or unchecked. Ground truth is labeled data that has been verified well enough to serve as the reference for training or measurement. All ground truth is labeled data, but only checked, trusted labels count as ground truth.