Models (LLMs) by grounding their responses in external knowledge sources, thus mitigating hallucinations. However, selecting the most effective RAG architecture for specific Question Answering (QA) tasks remains a challenge due to the variety of available techniques. This paper presents a comparative study of six different RAG systems, ranging from basic vector search and large context models to advanced architectures such as Sentence Window Retrieval, Auto-Merging, Knowledge Graphs, and Agentic RAG. We designed and evaluated 15 system variants?selected from approximately 230 preliminary trials?on a specific corpus of scientific and technical documents, using the ?RAG triad? metrics (Context Relevance, Groundedness, Answer Relevance) provided by the Trulens framework with an LLM-as-a-judge approach. Our experimental results indicate that Knowledge Graph-based systems (RAG 4) generally offer the best balance of performance metrics. However, Sentence Window Retrieval (RAG 2) proves highly effective for broad, conceptual questions. We also discuss the potential of Agentic RAG systems and the trade-offs between accuracy, cost, and latency.
Evaluating RAG approaches for Question Answering
Maristella Agosti;
2026
Abstract
Models (LLMs) by grounding their responses in external knowledge sources, thus mitigating hallucinations. However, selecting the most effective RAG architecture for specific Question Answering (QA) tasks remains a challenge due to the variety of available techniques. This paper presents a comparative study of six different RAG systems, ranging from basic vector search and large context models to advanced architectures such as Sentence Window Retrieval, Auto-Merging, Knowledge Graphs, and Agentic RAG. We designed and evaluated 15 system variants?selected from approximately 230 preliminary trials?on a specific corpus of scientific and technical documents, using the ?RAG triad? metrics (Context Relevance, Groundedness, Answer Relevance) provided by the Trulens framework with an LLM-as-a-judge approach. Our experimental results indicate that Knowledge Graph-based systems (RAG 4) generally offer the best balance of performance metrics. However, Sentence Window Retrieval (RAG 2) proves highly effective for broad, conceptual questions. We also discuss the potential of Agentic RAG systems and the trade-offs between accuracy, cost, and latency.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




