Topic 3: Compression and Filtering
5 min read·21 Sept 2026
Why compress after reranking? A chunk that’s relevant overall may be mostly irrelevant for this question. Sending the whole thing costs tokens, adds distraction, and pushes other evidence out of the budget.
3.1 Contextual Compression: Extracting Only Relevant Sentences
| Method | Cost | Risk | Best for |
|---|---|---|---|
| Embedding-based (score each sentence) | Milliseconds | May cut a sentence the answer depends on | High volume, cost-sensitive |
| LLM extraction (copy relevant sentences) | One call per passage (or batch) | Invented text, dropped conditions | Small candidate sets, hard passages |
| Cross-encoder per sentence | Tens of milliseconds | Same as embedding, more accurate | When you already run a cross-encoder |
| No compression | Free | Wasted tokens | Short chunks (already compressed) |