---
title: "MTEB embedding evaluation: measurement, limitations and evaluation guide · Undominated.ai"
canonical: https://undominated.ai/benchmarks/mteb/
description: "An embedding average is useful only when its task collection, aggregation and model revision are clear. Retrieval quality still needs a corpus-specific check."
---

# MTEB embedding evaluation: measurement, limitations and evaluation guide · Undominated.ai

> An embedding average is useful only when its task collection, aggregation and model revision are clear. Retrieval quality still needs a corpus-specific check.

Choose embeddings for retrieval and related representation tasks

# MTEB embedding evaluation

An embedding average is useful only when its task collection, aggregation and model revision are clear. Retrieval quality still needs a corpus-specific check.

Source reviewed 2026-09-23 · [Official MTEB project and task documentation](https://github.com/embeddings-benchmark/mteb)

 On this page [What it measures](#measurement)[Limitations](#limits)[Fair comparison](#controls)[Your evaluation](#experiment)

## What it measures

MTEB evaluates embedding and retrieval systems across distinct tasks, languages and modalities. Benchmark definitions select the tasks; individual task metrics and aggregate scores answer different questions.

Undominated publishes a scoped set of accepted results; see the linked instrument for its exact task, date and coverage.

## What the result does not establish

 - Do not compare differently named benchmark collections as though their averages used the same denominator.
- A broad aggregate may hide weak performance on the retrieval task that matters to your corpus.
- On Undominated’s accepted English benchmark, incomplete task coverage is shown as partial. Missing task results are never filled with zero or averaged away.

## Hold these conditions constant

 - Exact benchmark definition and task list
- Model checkpoint, embedding dimensions and input instructions
- Task-level results and aggregation rule
- Corpus language, retrieval setup and reranking stage

## An evaluation for your team

Create an evaluation set from the queries your users actually ask. Keep relevance labels and retrieval configuration fixed while measuring recall, ranking quality, observed latency and indexing cost separately.

Write down the task version, attempt count, accepted outcomes, observed spend and latency. Preserve failures in the denominator. A new model version or tool configuration deserves a new record.

 [Compare accepted embedding results →](/embeddings/)[Inspect retrieval tools →](/tools/)[Read publication rules →](/methodology/)

Editorial guidance is not a new measurement. Follow the official source for its current benchmark definition and result history.

## Continue your investigation

 - [Inspect confidence intervals](/significance/)
- [Compare embedding evidence](/embeddings/)
- [Evaluate a shortlist](/compare/)
- [Read ranking rules](/methodology/)
