Topic 2: Retrieval Metrics
4 min read·21 Sept 2026
Intuition: every metric answers a different question about the same ranked list. Picking the wrong one hides the problem you have.
| Metric | Answers | Rewards | Blind to |
|---|---|---|---|
| Recall@k | Did we find everything relevant? | Completeness | Where things ranked; how much junk came too |
| Precision@k | Was what we returned useful? | Cleanliness of the list | Relevant documents we missed |
| F1@k | Both at once | Balance | Ranking order |
| Hit rate@k | Did we find anything relevant? | The generator's ceiling | Everything else |
| MRR@k | How high was the first hit? | Getting one right answer to the top | Other relevant documents |
| MAP@k | How high were all the hits? | Ranking every relevant document well | Graded relevance |
| nDCG@k | How good is the ordering, with grades? | Perfect above partial, high above low | Needs graded labels |