---
title: "Terminal-Bench agent tasks: measurement, limitations and evaluation guide · Undominated.ai"
canonical: https://undominated.ai/benchmarks/terminal-bench/
description: "Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result."
---

# Terminal-Bench agent tasks: measurement, limitations and evaluation guide · Undominated.ai

> Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result.

Evaluate agents operating in terminal environments

# Terminal-Bench agent tasks

Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result.

Source reviewed 2026-09-23 · [Official Terminal-Bench and Harbor announcement](https://www.tbench.ai/news/announcement-2-0)

 On this page [What it measures](#measurement)[Limitations](#limits)[Fair comparison](#controls)[Your evaluation](#experiment)

## What it measures

Terminal-Bench evaluates agent tasks in executable environments. Its maintainers describe Harbor as a harness for container-based evaluation and identify task verification and changing external dependencies as important sources of difficulty.

This is a method guide. Undominated does not publish or blend this benchmark’s scores into its model ranking.

## What the result does not establish

 - An environment that changes between runs can change outcomes without a model improvement.
- Different agent scaffolds, permissions or retry budgets are different experimental conditions.
- A verified task outcome does not establish that an agent can safely operate your production environment.

## Hold these conditions constant

 - Benchmark and task revision
- Container image, dependencies and network policy
- Agent version, tools, time limit and retry budget
- Verifier output, failure logs and reproducible run artifacts

## An evaluation for your team

Start with a disposable environment and synthetic credentials. Include recovery tasks and explicit permission boundaries. Count a task only after an independent verifier accepts its final state; keep the run logs alongside observed costs.

Write down the task version, attempt count, accepted outcomes, observed spend and latency. Preserve failures in the denominator. A new model version or tool configuration deserves a new record.

 [Review agent execution boundaries →](/agents/)[Inspect tool access →](/mcp-servers/)[Record a task-specific comparison →](/compare/)

Editorial guidance is not a new measurement. Follow the official source for its current benchmark definition and result history.

## Continue your investigation

 - [Inspect confidence intervals](/significance/)
- [Compare embedding evidence](/embeddings/)
- [Evaluate a shortlist](/compare/)
- [Read ranking rules](/methodology/)
