Agentic AI benchmarks: task cards pass through an AI agent and its tools to a checkmark
The task gets done, and the result gets checked

TL;DR

  • A benchmark is a standard test used to compare AI systems. Agentic AI benchmarks test AI agents, AIs that can use tools such as a web browser or a computer’s command line (the text window for typing instructions). They check whether the job actually got done: does the program now work, did the booking really appear in the system.
  • No single test covers every kind of work. The score depends on both the AI model, its “brain”, and the software around it that hands it tools, so pick the tests closest to your own tasks.

Key Takeaways

  • They test doing, not knowing. The agent gets a goal and some tools, works through many steps on its own, and passes only if the job is finished, like a driving test rather than a written exam.
  • Each test covers one kind of work. Fixing software, helping a customer, searching the web and using a desktop computer are tested separately, and doing well in one says little about the others.
  • The software around the AI changes the result. In our own test, eight AI models given exactly the same tools used them very differently: one did 67% of its work through the command line, another under 1%.
  • An agent can find the answers if you let it. Given access to the computer’s files, agents in our test found the marking notes, so a test must keep them out of reach.

Agentic AI benchmarks are now how labs show what their agents can do, in code repositories, terminals, websites and customer chats.

We run one ourselves. Groundtruth, Eigenform’s open source dynamic benchmark, gives agents real geology documents, a shell and search tools. A separate track holds the model fixed to test the software around it.

What Are Agentic AI Benchmarks?

Agentic AI benchmarks evaluate AI agents, systems that plan and act through tools over many steps, by giving them a goal inside a working environment and checking the outcome. Success is measured by the state the agent leaves behind, such as passing tests or a correct database record, not by how an answer is worded.

In practice, that design has three parts: an environment the agent can act in, a task with a clear goal, and an automatic check of the result. The check is what makes benchmarking AI agents different from grading chat answers. A patch either passes the repository’s tests or it does not, and a refund either appears in the database or it does not.

The Agentic AI Benchmarks Shaping Agent Evaluation

Most widely cited agent benchmarks fall into a few families, listed below with our own Groundtruth for comparison. Figures in this section come from each benchmark’s original paper, so they show difficulty at release, not today’s leaders.

BenchmarkWhat the agent doesHow success is checked
SWE-bench VerifiedFixes real GitHub issues in Python repositoriesThe repository’s tests pass after the patch
Terminal-Bench 2.0Completes hard tasks in a computer terminalTests written for each task
τ-benchServes a simulated customer under company policyThe final database state matches the goal
WebArenaCompletes tasks on working, self-hosted websitesChecks on the functional result
OSWorldUses real desktop apps on Ubuntu, Windows and macOSA custom evaluation script per task
GAIAAnswers real-world questions using browsing and toolsExact short answers
BrowseCompFinds hard-to-locate facts on the webShort answers checked against references
Groundtruth (ours)Reads real geology reports with a shell and search tools, then answers expert questionsAn independent LLM judge scores each answer 0–10 against a rubric tied to source passages

AI coding benchmarks: SWE-bench and Terminal-Bench

SWE-bench is the reference among AI coding benchmarks for agents. It holds 2,294 real GitHub issues from 12 popular Python repositories; the agent edits the codebase, and the project’s own tests decide whether the issue is fixed. At release, the best model solved 1.96% of them. SWE-bench Verified is a 500-task subset that people filtered by hand.

Terminal-Bench 2.0 moves the work into a command line. Its 89 tasks each come with their own environment, a human-written solution and tests, and frontier models and agents scored under 65% when it launched.

However, both differ from classic LLM coding benchmarks such as HumanEval, which score a single function written from a description. Here the agent has to find its way around a whole repository or system before it changes anything.

Tool use and customer service: τ-bench

τ-bench tests an agent talking to a customer, played by a language model, while it uses company tools and follows written policy in retail and airline settings. Scoring compares the database at the end of the conversation with the intended goal.

τ-bench also introduced pass^k, the chance an agent succeeds on all of k attempts at the same task. At release, GPT-4o solved under half the tasks, and its pass^8 in retail was below 25%: the same agent often succeeded once and failed on a repeat.

A follow-up, τ²-bench, adds a telecom support domain in which the customer also has to take actions, so the agent must guide a person as well as act itself.

Web and computer use: WebArena, OSWorld and BrowseComp

WebArena runs working websites for shopping, a discussion forum, software development and content management. At release its best GPT-4 agent completed 14.41% of tasks, against 78.24% for people.

OSWorld gives agents a real operating system and 369 tasks across desktop and web apps, each with its own evaluation script. People completed 72.36% of the tasks; the best model managed 12.24%, mostly failing to find the right place to click.

BrowseComp asks 1,266 questions whose answers are hard to find but short and easy to check, which makes persistence in searching the web measurable.

General assistants and agent traces: GAIA and TRAIL

GAIA poses 466 questions that are simple for people but need browsing, file handling and reasoning. At release, people scored 92% and GPT-4 with plugins 15%.

TRAIL tests the tester. It holds 148 agent traces from GAIA and SWE-bench runs, the full record of what an agent did, with 841 errors marked by people. The best model found and placed the errors correctly only 11% of the time, so asking a model to grade an agent’s trace is not yet reliable.

The AI Agent Harness Benchmark: Why the Wrapper Matters

An agent is a model plus a harness: the prompts, tools, memory and loop that decide what the model sees and does next. A score belongs to both, so an AI agent harness benchmark has to say which one it is measuring. The default view of the SWE-bench Verified leaderboard now runs every model in the same minimal agent, for exactly that reason.

We measured the effect directly. Eight models ran Groundtruth in the same opencode harness, with the same file-reading, search and shell tools. Gemini 3.1 Pro Preview made 67% of its tool calls in the shell; GPT 5.6 Sol made under 1%. GPT 5.6 Sol and Grok 4.6 used almost the same mix of tools and still finished 0.77 points apart, so tool habits alone did not predict the score.

For that reason, Groundtruth’s harness track pins the model to one reference model, currently GLM 4.7, and lets the harness vary. The method, holding the model fixed, turns any difference from the model’s own baseline into a measurement of the harness.

Both tracks report results on the Groundtruth leaderboard, in separate rows.

What Agentic AI Benchmarks Miss

Agentic AI benchmarks measure a lot, but three gaps come up repeatedly.

Agents look around

An agent with a shell can read anything its account can read. In our benchmark, agents found the grading rubrics on the same machine. Hiding only the authoring instructions, or renaming the rubric folder, still leaked grading text on 8 of 50 questions each. Moving the files out of the repository closed the leak, and reading its own answer key is now something every agent benchmark has to test for.

One run is not a measurement

Agents vary from run to run, which is what τ-bench’s pass^k exposes. In addition, small test sets add their own noise: on our 50-question sets, the smallest gap we can tell apart from chance is 0.27 to 0.45 points out of 10, as measured in the resolution of a 50-question benchmark.

Your tasks are not on the list

Finally, each benchmark covers one environment and one kind of task. A coding score says little about an agent reading drill reports or reconciling invoices. For that, the test has to come from your own material, which is what dynamic benchmarking generates from a set of documents.

How to Choose Benchmarks When Benchmarking AI Agents

  • Match the environment. Choose benchmarks whose tools and setting resemble your deployment: a repository, a terminal, a browser or a customer conversation.
  • Check what is held fixed. A leaderboard row is a model, a harness or both, and comparing across them mixes two effects.
  • Ask for repeated runs. Prefer results that report several attempts per task, or a reliability measure such as pass^k.
  • Add your own tasks. Public scores narrow the shortlist; a test built from your own work makes the decision.

Benchmarks are one part of AI agent evaluation, alongside checks on the agent’s steps, cost and failures in use.

FAQs

What do agentic AI benchmarks measure?

They measure whether an AI agent can complete a multi-step task with tools inside a working environment. Instead of grading the wording of an answer, they check the outcome: tests that pass, a database record that matches the goal, or a short answer that matches a reference.

What is the best benchmark for AI agents?

There is no single best one, because each covers a different kind of work. SWE-bench and Terminal-Bench suit coding agents, τ-bench suits customer-facing agents, WebArena, OSWorld and BrowseComp suit web and computer use, and GAIA suits general assistants. Choose the ones closest to your own tasks.

How do you evaluate AI agents for your own use?

Start with public agentic AI benchmarks to narrow the shortlist, then test on tasks taken from your own work. Fix the harness when comparing models, run each task several times, and check the agent’s steps as well as its final result. That is how to evaluate AI agents for a real decision.

Why do agent benchmark scores differ between leaderboards?

Leaderboards often run different harnesses, prompts, tool sets and numbers of attempts, and some report a model while others report a whole agent product. A model can therefore rank differently on two leaderboards for the same benchmark. Read what each one held fixed before comparing their numbers.