---
title: "AI benchmark guide: Arena, MTEB, coding and tool use · Undominated.ai"
canonical: https://undominated.ai/benchmarks/
description: "Choose evidence for your engineering task. Compare what AI benchmarks measure, inspect accepted coverage, and design a task-specific model evaluation."
---

# AI benchmark guide: Arena, MTEB, coding and tool use · Undominated.ai

> Choose evidence for your engineering task. Compare what AI benchmarks measure, inspect accepted coverage, and design a task-specific model evaluation.

Choose the evidence

# Which benchmark answers your question?

Start with the job your system must do. Preference, retrieval, issue resolution and tool use measure different outcomes. Keep their scores separate, then test your shortlist on your own acceptance criteria.

Compare conversational preference within a published task

## [Arena human preference](/benchmarks/arena/)

Use preference ratings to form a shortlist, then inspect uncertainty and the exact model variant before comparing cost.

**136 / 346** standard model entries have an accepted Arena score. Catalogue as of 2026-09-23.

 [Inspect measurement and limitations →](/benchmarks/arena/)

Choose embeddings for retrieval and related representation tasks

## [MTEB embedding evaluation](/benchmarks/mteb/)

An embedding average is useful only when its task collection, aggregation and model revision are clear. Retrieval quality still needs a corpus-specific check.

**20 / 57** embedding entries have a complete accepted benchmark score. Catalogue as of 2026-09-23.

 [Inspect measurement and limitations →](/benchmarks/mteb/)

Evaluate repository issue resolution

## [SWE-bench coding agents](/benchmarks/swe-bench/)

A coding result describes a model working through a particular agent and environment. Compare the full run configuration before treating a leaderboard entry as a model choice.

Method guide with links to the official results. Scores from this benchmark are not ingested into Undominated.

 [Inspect measurement and limitations →](/benchmarks/swe-bench/)

Evaluate agents operating in terminal environments

## [Terminal-Bench agent tasks](/benchmarks/terminal-bench/)

Terminal results depend on the task environment as well as the model. Pin the environment and verifier when comparing systems or repeating a result.

Method guide with links to the official results. Scores from this benchmark are not ingested into Undominated.

 [Inspect measurement and limitations →](/benchmarks/terminal-bench/)

Evaluate structured tool use and agentic interactions

## [Berkeley Function Calling Leaderboard](/benchmarks/bfcl/)

Tool-call correctness is more specific than general chat quality. Check the evaluation category and whether the model uses native function calling or a prompting workaround.

Method guide with links to the official results. Scores from this benchmark are not ingested into Undominated.

 [Inspect measurement and limitations →](/benchmarks/bfcl/)

## Turn a leaderboard into a decision

 - Choose the task and define an acceptable result.
- Match benchmark version, model variant, tools and budget.
- Keep cost, latency, correctness and failure recovery as separate observations.
- [Compare your shortlist](/compare/), attach private evaluation results, and save a record of the evidence used.

Guide sources reviewed 2026-09-23. Coverage counts come from the accepted catalogue; they are not a census of every public result.

## Continue your investigation

 - [Inspect confidence intervals](/significance/)
- [Compare embedding evidence](/embeddings/)
- [Evaluate a shortlist](/compare/)
- [Read ranking rules](/methodology/)
