CourseRAG · Module 8: Reranking and Post-Retrieval Processing · part 41 of 82
Part 41 · Module 8: Reranking and Post-Retrieval Processing

Topic 3: Compression and Filtering

5 min read·21 Sept 2026

Why compress after reranking? A chunk that’s relevant overall may be mostly irrelevant for this question. Sending the whole thing costs tokens, adds distraction, and pushes other evidence out of the budget.

3.1 Contextual Compression: Extracting Only Relevant Sentences

MethodCostRiskBest for
Embedding-based (score each sentence)MillisecondsMay cut a sentence the answer depends onHigh volume, cost-sensitive
LLM extraction (copy relevant sentences)One call per passage (or batch)Invented text, dropped conditionsSmall candidate sets, hard passages
Cross-encoder per sentenceTens of millisecondsSame as embedding, more accurateWhen you already run a cross-encoder
No compressionFreeWasted tokensShort chunks (already compressed)

The rest of this course is yours to keep

This course is bought on its own, once, and stays readable afterwards, including the parts added to it later.