Loading...🤓
March 8, 2025
Short questions and long documents live in different regions of embedding space — and that makes vector search brittle. HyDE bridges the gap by drafting a fictional answer first, then searching for real documents that resemble it.
HyDE — Hypothetical Document Embeddings — flips the usual order: instead of embedding the question directly, the LLM first writes a plausible answer to it. That fictional answer is embedded and used for the vector search. Because the synthetic document mirrors the style and vocabulary of the real ones in the database, retrieval lands far closer to the right material than a bare question ever could.
The pipeline begins when the language model drafts a plausible but unverified answer to the user's query. This hypothetical response is then converted into an embedding — a numerical representation — and used as the search vector against the vector database. Because the fabricated answer is structurally closer to actual documents than a question would be, the similarity search returns more relevant results. Those real documents then replace the hypothetical answer and serve as the grounding context for the final, verified response.

HyDE excels when users ask vague or conceptual questions. Consider someone asking "that one law about digital privacy in California." A direct keyword or embedding search on this phrasing might struggle. HyDE, however, generates a hypothetical summary of CCPA, uses that to locate the actual legal text, and then delivers a grounded answer citing the real source material.
Retrieval quality improves substantially for abstract, open-ended, or poorly articulated queries. The approach requires no specialised agent logic — it works within a straightforward pipeline, making it relatively easy to implement.
If the hypothetical answer is fundamentally wrong, the entire retrieval step will be misled, leading to incorrect or irrelevant results. This architecture also adds unnecessary overhead for simple factual lookups where direct matching would suffice.
HyDE is a strong fit for knowledge bases where user queries tend to be conceptual, exploratory, or loosely worded. For precise, factual retrieval tasks, simpler approaches will be faster and more reliable.
Gao et al. — Precise Zero-Shot Dense Retrieval without Relevance Labels (2022)