Evaluate agents operating in terminal environments

Terminal-Bench agent tasks

Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result.

Source reviewed 2026-09-23 · Official Terminal-Bench and Harbor announcement

What it measures

Terminal-Bench evaluates agent tasks in executable environments. Its maintainers describe Harbor as a harness for container-based evaluation and identify task verification and changing external dependencies as important sources of difficulty.

This is a method guide. Undominated does not publish or blend this benchmark’s scores into its model ranking.

What the result does not establish

  • An environment that changes between runs can change outcomes without a model improvement.
  • Different agent scaffolds, permissions or retry budgets are different experimental conditions.
  • A verified task outcome does not establish that an agent can safely operate your production environment.

Hold these conditions constant

  • Benchmark and task revision
  • Container image, dependencies and network policy
  • Agent version, tools, time limit and retry budget
  • Verifier output, failure logs and reproducible run artifacts

An evaluation for your team

Start with a disposable environment and synthetic credentials. Include recovery tasks and explicit permission boundaries. Count a task only after an independent verifier accepts its final state; keep the run logs alongside observed costs.

Write down the task version, attempt count, accepted outcomes, observed spend and latency. Preserve failures in the denominator. A new model version or tool configuration deserves a new record.

Editorial guidance is not a new measurement. Follow the official source for its current benchmark definition and result history.

Evidence & Ask