Dev News Daily ENDE

Spring AI shows a RAG pipeline that drops retrieved chunks which do not answer

A post on the Spring Blog by Christian Tzolov puts two pieces of Spring AI together: the Modular RAG pipeline in the spring-ai-rag module, and Jev, the small judging model from Spring AI TypeSafe that the previous post in the series introduced. The demo module, 05-1-modular-rag, sits next to the naive 05-rag version so the two can be compared.

The post names three problems with the classic approach of embedding the user's text and stuffing similar chunks into the prompt. Users do not type search queries, so embedding a chatty question verbatim drags its noise into the search. Similar is not the same as useful: a reference list that mentions the topic five times scores high and answers nothing. And retrieved text is untrusted input that goes straight into the prompt.

The RetrievalAugmentationAdvisor addresses these in phases, each a swappable interface. Before retrieval, one LLM call rewrites the question into a search-friendly query and another expands it into several queries; retrieval runs per query in parallel and the results are joined. After retrieval, JevDocumentFilter asks four narrow questions about each passage and drops those that fail, and JevDocumentReranker asks whether the passage could answer the query, sorts by score and keeps the top results. Filtering runs first so the reranker only pays for survivors, and both stages see the user's original question, not the rewritten one.

In the demo, a question about whether Hurricane Milton made landfall in Florida became four searches that returned six chunks. Four were reference-list entries and archive links; the filter removed them, and the reranker scored the remaining two at 0.97 and 0.88. With a limit of three, the model received two passages and answered correctly. The post sums up the distinction: vector similarity tells you what a text is about, Jev whether it answers the question.

Spring AI shows a RAG pipeline that drops retrieved chunks which do not answer
Spring AI shows a RAG pipeline that drops retrieved chunks which do not answer — Dev News Daily

Why it matters

Most RAG failures in practice are retrieval failures that look like model failures. Putting a relevance judge between the vector store and the prompt is cheap to try in Spring AI because it is one DocumentPostProcessor in an existing advisor.