Five connected cards showing an AI evaluation framework: Dataset, Rubric, Judge, Metrics and Reporting
Each component can fail on its own, so each one needs its own check

TL;DR

  • An AI evaluation framework is the whole system for testing an AI. It works like setting, marking and reporting a school exam: the questions, the marking guide, the marker and the report card.
  • Each part can fail on its own. Good questions cannot rescue a marker who gives points for confident wording, and a correct score still misleads if the report hides how uncertain it is.
  • Our examples come from Groundtruth, Eigenform’s open source dynamic benchmark: 150 geology questions, eight AI models, and software that checks every score the AI marker gives.

Key Takeaways

  • A framework is a system, not a score. Our AI marker’s output must pass 11 automatic checks before any score counts.
  • The questions set the ceiling. In one old report we used, 18% of the pages were too damaged by scanning errors to write questions from.
  • The marking guide and the marker need testing too. Practice answers with known scores found 15 marking-guide mistakes that a careful read had missed.
  • A good report says where the test itself could be wrong. Across eight models, only 2 of 7 score gaps between neighbours were big enough to be real.

An AI evaluation framework is what stands between a model’s output and the claim that the model is good enough. Most discussion of evaluation is about choosing a metric. However, the metric is only the last link in a longer chain: where the questions come from, what counts as a correct answer, who grades it, and how you report the result.

We built that chain for Groundtruth, Eigenform’s open source dynamic benchmark, and ran it on real geology exploration reports. Most of what we learned came from one link failing while the others looked fine.

What is an AI evaluation framework?

An AI evaluation framework is the system that turns a model’s outputs into a judgment you can trust. It connects a dataset of tasks, a rubric that defines a correct answer, a judge that applies it, metrics that summarise the scores, and reporting that states how certain the result is. Each part needs its own design and its own test.

Libraries such as OpenAI Evals, DeepEval and promptfoo also call themselves evaluation frameworks, but they run tests and leave the data, the rubric and the judge to you. AI evaluation frameworks in this guide’s sense include those decisions.

Governance frameworks sit at the other end. The NIST AI Risk Management Framework organises risk work into four functions: Govern, Map, Measure and Manage. It says what an organisation should measure and why, but not how to score an answer.

The core components of an AI evaluation framework

Every AI evaluation framework has the same five components, and each one needs its own check.

Dataset

The dataset holds the tasks and their reference answers, the ground truth data for AI that every score depends on. It also bounds everything downstream: a wrong reference answer marks a correct model wrong, and no judge or metric can undo that later.

For a benchmark built from documents, the dataset also includes the text that the questions draw on. We ranked 190 candidate reports, just over eight million words, by measured legibility before choosing any. In one 1920s report, 75 of 419 text pages (18%) were too damaged by OCR errors to anchor a question to.

Every reference answer carries an evidence anchor, a pointer to the passage that supports it. A script then checks that each anchor is a verbatim substring of the extracted text. On one build, that check found 41 corrections to a ground truth dataset built from scanned PDFs, three of them anchors on the wrong page.

Rubric

The rubric defines what a correct answer must contain. A loose instruction such as “score this out of 10” rewards answers that sound like the domain, whether or not they are correct. Our rubrics split 10 points into non-overlapping components, one of them a gate worth 2 to 4 points. If an answer misses the concept the gate tests, the whole answer scores zero.

Once the gate passes, the other components only add points, and a do-not-credit list names claims that sound plausible but must earn nothing.

You have to test a rubric, not just re-read it. Before we grade any model, we run fixture answers (fabricated answers with a known expected score, four per question) through the production judge. On one benchmark, the author’s own review of the AI evaluation rubric found two double-credit defects, and the fixture run found fifteen more.

Judge

The judge applies the rubric. For open-ended answers it is usually another language model, an approach known as LLM-as-a-judge. That makes the judge part of the measuring instrument, so it needs the same scrutiny as the model under test.

Ours is one fixed model at temperature 0. It must end every response with a JSON block: the gate decision, each component’s award and reason, the total, and an adjudication flag. A validator checks eleven rules before accepting the score, for example that the awards add up to the total. A response that fails is re-graded up to three times. After that, the question is left unscored rather than given a guessed number.

A format check cannot catch a wrong judgment. One fixture reached the right conclusion by reasoning the rubric rejected. Even so, the judge gave it 10 out of 10 against an expected 0, in a clean, valid JSON block. The fixture run caught it; the validator could not.

In pairwise LLM evaluation, judges tend to favour one position regardless of content. We therefore ask twice with the order swapped and average the verdicts. As a result, a judge that always picks the first answer produces a tie, not a wrong winner.

Metrics

Metrics turn individual scores into a comparison, and choosing them is a large part of AI model evaluation. For rubric-graded answers the usual metric is the mean score per model. What matters is how small a difference that mean can resolve.

With 50 questions per set, the smallest gap we could distinguish from noise was 0.27 points on WAMEX, 0.33 on USGS and 0.45 on Technical. A 0.3-point improvement is a finding on one set and noise on another, from the same benchmark.

Across eight models and 1,200 graded answers, only 2 of 7 gaps between neighbouring models reached benchmark statistical significance. Bootstrap confidence intervals turned the ranking into three tiers: four models at the top, three in the middle and one below.

Reporting

Reporting decides what a reader can conclude. We publish tiers from confidence intervals rather than a plain ranking, and state which neighbouring gaps are significant. The Groundtruth leaderboard uses only pointwise scores, where the judge grades each answer against its rubric on its own.

A report should also say where the evaluation itself might be wrong. Ours discloses that the judge’s model family also appears on the leaderboard, where it places third. It also states that we scored each model once per set, and that costs are harness costs, not list prices. Answers the judge flags for adjudication keep their score; however, the harness lists them after every run for a person to check.

Cost belongs in the same report as quality. Among models of similar accuracy, cost per question ranged from about $0.04 to $0.37. So once scores are statistically tied, cost is what separates the best AI models for geology.

What changes in an AI agent evaluation framework

An agent plans, calls tools and decides when to search again. As a result, an AI agent evaluation framework has to measure the path as well as the answer. The five components stay the same, but each has more to cover, which is why AI agent evaluation scores reasoning, tool use, trajectory and outcome separately.

Agentic AI evaluation framework: hold one variable fixed

A score belongs to a model and its harness together, the harness being the software that gives the model its tools and instructions. So our leaderboard has two tracks. One holds the harness fixed and varies the model; the other pins one reference model and varies the harness.

The harness track is an agent evaluation framework that holds the question sets, rubrics, judge, reference model and result format constant. The reference model alone scores 6.81, the baseline any harness has to beat.

Holding the harness fixed does not make behaviour uniform. Eight models ran in one harness with the same three tools. Even so, Gemini 3.1 Pro Preview made 67% of its tool calls in the shell, while GPT 5.6 Sol made 0.4%. Tool mix did not predict the score either: two models with almost the same mix finished 0.77 points apart. So we score the path from saved transcripts, not against a fixed reference path.

Keeping the answer key out of reach

An agent with file and shell tools can reach anything its account can read, including the grading material. In our AI agent testing, an agent reached the rubric folder on 13 of 50 questions. Hiding or renaming things still leaked grading text on 8 of 50; only making the files absent during the run worked.

Detection backs up prevention. We audit every transcript for access to the grading key before its score counts. So far the audit has caught exactly one leaked answer, which we discarded and regenerated.

Building vs. buying an AI evaluation framework

Search for AI agent evaluation frameworks 2026 and most results are ranked lists from vendors, each placing its own product first. A more useful split is by component. Tools are good at running tests and storing results; however, they are weak at the parts that depend on your domain.

ComponentWhat off-the-shelf tools give youWhat you usually build
DatasetStorage, versioning, test sets drawn from tracesTasks and reference answers for your domain
RubricGeneric metrics such as relevance and faithfulnessGates, component points, do-not-credit lists
JudgeLLM-judge templates and scorersValidation of the judge on your own fixtures
MetricsAggregation, dashboards, pass/fail gates in CIResolution analysis for your sample size
ReportingExperiment tracking and comparison viewsA statement of what the results cannot tell you

Open-source libraries cover the running layer: Inspect, from the UK AI Security Institute, supports tool use, multi-turn tasks and model grading. Tracing platforms such as LangSmith and Langfuse add production traces and online scoring.

When to build your own

Build the parts no general metric captures. For us that was reasoning over specific documents, which a general benchmark cannot separate from what a model memorised. Our dynamic benchmarking pipeline generates questions, reference answers and rubrics from a set of documents and verifies each against its source. In addition, the generator, calibration process and scoring logic are MIT-licensed.

Building has a running cost in people, not only in compute. Our authoring flow for a custom LLM benchmark has one human approval and four scripted gates, and two passes still stay manual: re-scoring the fixtures and expert review.

An AI tool evaluation framework for buying

When you buy, evaluate the evaluation tool the way you would evaluate a model. Five questions cover most of it:

  1. Can you bring your own dataset and rubric, or only use its built-in metrics?
  2. Is the judge’s output checked against a schema, and what happens when the check fails: a retry, an empty score or a default number?
  3. Can you pin the judge model and its temperature, and does every result record which judge produced it?
  4. Does it report confidence intervals, or only a mean?
  5. For agents, can it keep grading material out of the agent’s reach, and does it keep full transcripts?

Built or bought, the framework still runs the three stages of how to benchmark AI models. First, construct the task set. Then run it with the model isolated. Finally, read the score against the resolution the sample supports.

Common failure modes across an AI evaluation framework

Each component fails in its own way, and each failure looks like a model result. The sections above cover the first three. The last three are easy to miss, because every routine check still passes.

FailureWhere it entersWhat catches it
Damaged or misread source textDatasetLegibility scores; verbatim anchor checks
Double credit across the gateRubricFixtures run through the production judge
Valid format, wrong judgmentJudgeFixtures, not the validator
The judge grades the wrong questionJudge inputReading a sample of judge transcripts
The metric measures the wrong thingMetricsChecking it against the outcome you care about
One document set hides reorderingsDataset coverageSeveral document sets

In the judge-input failure, the harness took the question text for the judge from the answer file rather than the rubric. When an answer file carried an empty or mismatched question field, the judge graded against the wrong question, and the harness only logged a warning.

The metric failure came from benchmarking a self-improving agent. As real task success climbed across generations, pass@1 fell, because lineages that wrote failing code on the first attempt got informative error messages to learn from.

Coverage fails quietly. Deepseek V4 Pro was the lowest of the middle tier on the Technical set (7.34) and the highest of those three models on WAMEX (8.46). With one document set, either ordering would have looked like the answer.

FAQs

What is the difference between an AI evaluation framework and an evaluation tool?

An evaluation tool runs tests and stores scores. An AI evaluation framework is the whole system around it: the dataset, the rubric, the judge, the metrics and the reporting rules, plus the checks on each. A tool can be one part of a framework, but it cannot decide what a correct answer is in your domain.

What are the components of an AI evaluation framework?

There are five. A dataset holds the tasks and their verified reference answers, a rubric defines a correct answer, and a judge applies the rubric. Then metrics turn scores into comparisons, and reporting states what the results can and cannot support. Each can fail independently, so each needs its own check.

Should you build or buy an AI evaluation framework?

Buy the running layer: test execution, logging, dashboards and tracing are solved problems. Build what depends on your domain: the tasks, the reference answers, the rubric and the validation of the judge against it. If no general metric captures what you need to measure, most of the framework’s value sits in the parts you build.

How do you know an AI evaluation framework is reliable?

Test it before trusting it. Run answers with known expected scores through the real judge, check that every reference answer traces to its source, and compute the smallest score gap your sample can detect. A reliable framework also reports its own limits, such as judge conflicts and single runs, next to the scores.