Loading...🤓
May 19, 2026
Comparative questions like "How does RAG variant X differ from Y?" break standard RAG. Agentic RAG splits them into sub-investigations — here is how our rag-service implementation handles that, step by step.
Agentic RAG is not one pattern but a spectrum. At one end sits multi-step retrieval, where an LLM decomposes a complex question into sub-questions, researches them one by one, and checks its own intermediate results. At the other end sit fully orchestrated multi-agent systems, where an orchestrator coordinates autonomous sub-agents with their own tools, data sources, and specialisations. The literature groups both under the label Agentic RAG. What matters in practice is where on that spectrum your implementation actually lives.
Our implementation sits at the first end: iterative, self-checking retrieval against a single knowledge base, with web search as a fallback. That is not a full multi-agent system — but it is already powerful enough to handle the comparative and analytical questions that a single retrieval pass cannot answer.
The flow starts with query decomposition. A local grading LLM — in our case a Mistral Small 3.1 (24B) running on our own inference host — is prompted in German to split the original question into two to four independent sub-questions. If the model fails or returns something unusable, a fail-open mechanism kicks in: the original question is treated as the only sub-question, and the pipeline moves on. A flaky grading model will never block a response.
Then the actual research begins. Each sub-question is embedded with intfloat-multilingual-e5-large and queried against Qdrant. If vector search returns nothing or its top score sits below the CRAG incorrect-threshold of 0.3, the system also queries SearXNG and merges those web results with whatever Qdrant returned for that sub-question — low-score vector hits are kept rather than thrown away, so partial matches still contribute to the synthesis. All collected documents are then deduplicated by URL across every sub-question, so the same source never shows up twice in the final context.
After each iteration, a completeness check runs. The grading LLM is asked again — in German, with a simple yes-or-no answer expected — whether the collected evidence is sufficient to answer the original question. If the model errors out, the check fails open so that an unstable grading model cannot keep the loop alive forever. If the answer is no, the LLM generates two or three fresh sub-questions that target the missing aspects, and the next iteration runs against those. A hard cap of three iterations prevents infinite loops even if the grading model stubbornly keeps saying no.
Once the loop ends, all deduplicated documents and web results are handed to the main answer LLM. A multi-source synthesis prompt instructs the model to connect information across sources, flag contradictions explicitly, and cite every claim as a Markdown link. Training knowledge is forbidden — the answer has to come from the supplied context or nowhere. The entire flow is narrated live through Server-Sent Events: our chat UI receives step events like query_decomposed, subquestion_started, subquestion_completed, completeness_checked, and web_search_completed, and renders them as a real-time animation of what the agent is doing. Users see the research happen rather than staring at a spinner.

Our knowledge base is focused on the German energy transition — renewable generation, grid stability, storage technologies, and the regulatory framework. A realistic question for this corpus is: "What is the better technology to stabilise the German power grid — gas-fired peaker plants or large-scale battery storage? Please weigh both with sources." A single vector search struggles with this kind of question regardless of what is in the index. The top hits tend to cluster around whichever technology dominates the corpus on any given day, and the comparative angle — when does each option actually win — rarely gets balanced retrieval on its own.
Our agentic mode decomposes the question into separate sub-questions — typically one per technology being compared, plus one for the comparison itself, for example: "What role do gas-fired power plants play in stabilising the German grid?", "What role do battery storage systems play in stabilising the German grid?", and "How do gas plants and batteries compare on cost, response time, and CO2 footprint in the German market?". Each sub-question runs its own Qdrant search against the renewable-energy corpus and pulls back whatever evidence is currently indexed for that specific angle — independent of what the other sub-questions retrieve. The completeness check then decides whether all angles are covered or whether another iteration with fresh gap-questions is needed. Once it terminates, the synthesis LLM merges the per-sub-question buckets into one structured comparative answer with clean citations — separate retrieval per angle, deduplication across angles, evidence-grounded synthesis at the end.
The biggest gain is the ability to answer comparative and analytical questions that a single retrieval step inherently cannot handle. Because every sub-question gets its own retrieval pass, the system covers queries that touch multiple facets of a knowledge base at the same time. The live SSE animation makes the process transparent to users — they see that real research is happening instead of waiting for a black box. The consistent fail-open philosophy ensures that a wobbly grading model never blocks an answer, it merely reduces the extra value on top. And where the knowledge base has gaps, the SearXNG fallback steps in automatically.
The price for all this is latency. Typical response times in agentic mode land at 8 to 15 seconds, while pure Corrective RAG in the same system finishes in 5 to 8. That is the cost of the extra LLM calls: query decomposition, completeness checks, and — when needed — gap-question generation all stack on top of retrieval. Worse, the quality of the whole response depends heavily on the quality of the decomposition. If the grading LLM splits the original question poorly, every subsequent retrieval looks in the wrong place, and the final synthesis cannot recover from that. For simple factual questions, the mode is simply overkill.
Our implementation is deliberately narrow: one vector store, multi-step retrieval, one synthesis at the end. The full picture of Agentic RAG that Singh et al. 2025 describe in their survey goes considerably further. The next step would be multiple domain-specialised sub-agents — one for product documentation, one for CRM data, one for pricing, one for technical specifications. Each of these agents would have its own retrieval strategy tailored to its data source.
On top of that, those sub-agents could wield tools far beyond vector search: direct API calls, SQL queries against structured databases, calculators for numerical work, even sandboxed code execution. A true orchestrator would then plan dynamically, based on the query, which agents to call in what order and which steps can run in parallel — instead of marching through the steps sequentially as we do today. At this level, tool selection happens via function calling: the orchestrator decides on the fly which tool fits which sub-question rather than relying on a fixed pipeline.
And finally, we are missing what truly defines a multi-agent system: consistency checks between agents. When two sub-agents return overlapping or conflicting information on a topic, the orchestrator should detect that and reconcile — perhaps by consulting a third source, or by flagging the contradictions explicitly in the final answer. Today, our system delegates that responsibility to the final LLM's synthesis prompt. That works, but it is nowhere near as structured as a dedicated reconciliation step would be. For a company stitching together heterogeneous data silos, this is the natural next step up.
Enable the agentic mode when your users ask comparative, multi-faceted, or research-style questions that touch several aspects of your knowledge base at once. For simple factual questions, Corrective RAG remains the better choice — faster, cheaper, and accurate enough. The extra effort only pays off when multiple sub-topics need to be weighed cleanly against each other and users expect a structured answer with each individual claim backed by a source.
Singh et al. — Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG (2025)