Choose embeddings for retrieval and related representation tasks
MTEB embedding evaluation
An embedding average is useful only when its task collection, aggregation and model revision are clear. Retrieval quality still needs a corpus-specific check.
Source reviewed 2026-09-23 · Official MTEB project and task documentation
What it measures
MTEB evaluates embedding and retrieval systems across distinct tasks, languages and modalities. Benchmark definitions select the tasks; individual task metrics and aggregate scores answer different questions.
Undominated publishes a scoped set of accepted results; see the linked instrument for its exact task, date and coverage.
What the result does not establish
- Do not compare differently named benchmark collections as though their averages used the same denominator.
- A broad aggregate may hide weak performance on the retrieval task that matters to your corpus.
- On Undominated’s accepted English benchmark, incomplete task coverage is shown as partial. Missing task results are never filled with zero or averaged away.
Hold these conditions constant
- Exact benchmark definition and task list
- Model checkpoint, embedding dimensions and input instructions
- Task-level results and aggregation rule
- Corpus language, retrieval setup and reranking stage
An evaluation for your team
Create an evaluation set from the queries your users actually ask. Keep relevance labels and retrieval configuration fixed while measuring recall, ranking quality, observed latency and indexing cost separately.
Write down the task version, attempt count, accepted outcomes, observed spend and latency. Preserve failures in the denominator. A new model version or tool configuration deserves a new record.
Editorial guidance is not a new measurement. Follow the official source for its current benchmark definition and result history.