Evaluate agents operating in terminal environments
Terminal-Bench agent tasks
Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result.
Source reviewed 2026-09-23 · Official Terminal-Bench and Harbor announcement
What it measures
Terminal-Bench evaluates agent tasks in executable environments. Its maintainers describe Harbor as a harness for container-based evaluation and identify task verification and changing external dependencies as important sources of difficulty.
This is a method guide. Undominated does not publish or blend this benchmark’s scores into its model ranking.
What the result does not establish
- An environment that changes between runs can change outcomes without a model improvement.
- Different agent scaffolds, permissions or retry budgets are different experimental conditions.
- A verified task outcome does not establish that an agent can safely operate your production environment.
Hold these conditions constant
- Benchmark and task revision
- Container image, dependencies and network policy
- Agent version, tools, time limit and retry budget
- Verifier output, failure logs and reproducible run artifacts
An evaluation for your team
Start with a disposable environment and synthetic credentials. Include recovery tasks and explicit permission boundaries. Count a task only after an independent verifier accepts its final state; keep the run logs alongside observed costs.
Write down the task version, attempt count, accepted outcomes, observed spend and latency. Preserve failures in the denominator. A new model version or tool configuration deserves a new record.
Editorial guidance is not a new measurement. Follow the official source for its current benchmark definition and result history.