# Retrieval evaluation worksheet

Blank working document. Replace placeholders with your own evidence; an empty field is unknown, not a passing result.

Owner: [fill in]
Source revision: [fill in]
Environment: [fill in]
Evidence date: [fill in]

## Frozen system

Corpus version/hash: ___
Chunker/embedding/reranker configuration: ___
Existing collection identity: ___
Read-only setting and denied provisioning permissions: ___
Runner version: ___

## Case labels

| Question ID | Expected passage IDs | Answerable? | Label source |
| --- | --- | --- | --- |
| ___ | ___ | unknown | ___ |

## Observed output

Retrieved IDs/ranks: ___
Answer and cited passages: ___
Unsupported statements: ___
Abstention behaviour: not run

## Experiment

Failure categories: ___
Case denominator/exclusions: ___
Single proposed change: ___
Held-out rerun evidence: not run
Data-sharing limits: ___

## Boundaries

- Qdrant MCP retrieves semantic memories; it is not a complete RAG evaluation harness. Read-only mode removes the storage tool, but does not by itself establish zero backend writes during setup. Use an existing collection and credentials that deny provisioning; FastEmbed can download models and use local compute.

- The reviewer supplies no dataset or evaluation dependencies. Its read-only Bash instruction is not a sandbox, and model calls or corpus uploads require separate authorisation.

Workflow: https://undominated.ai/workflows/#evaluate-retrieval-grounding

Original worksheet: MIT. Linked resources retain their own licences and setup requirements.
