# Evaluate retrieval grounding and abstention — acceptance checklist

Record evidence for every checked item. Unchecked or untested is not a pass. This checklist does not execute an integration or authorise external actions.

## Required inputs

- [ ] Versioned corpus, an existing populated collection, chunking/embedding configuration and retrieval code.

- [ ] Held-out questions with authorised expected evidence and an existing evaluation runner.

## Output checks

- [ ] Evaluation labels were set independently of the tested output.

- [ ] Unanswerable cases and missing-source cases are included.

- [ ] Retrieved source IDs trace to the frozen corpus, and inspection uses an existing collection without provisioning permissions.

- [ ] Any aggregate measure names its case denominator and excluded cases.

## Deliverables

- [ ] Corpus/configuration manifest

- [ ] Question and relevance set

- [ ] Per-case retrieval/answer evidence

- [ ] Failure analysis and next experiment

## Evidence and decision

Evidence links: [fill in]
Untested paths: [fill in]
Unresolved findings: [fill in]
Reviewer: [fill in]
Decision and scope: [fill in]

## Boundaries

- Qdrant MCP retrieves semantic memories; it is not a complete RAG evaluation harness. Read-only mode removes the storage tool, but does not by itself establish zero backend writes during setup. Use an existing collection and credentials that deny provisioning; FastEmbed can download models and use local compute.

- The reviewer supplies no dataset or evaluation dependencies. Its read-only Bash instruction is not a sandbox, and model calls or corpus uploads require separate authorisation.

Workflow: https://undominated.ai/workflows/#evaluate-retrieval-grounding
