Russian researchers from SberAI, MWS AI, ITMO and HSE University, have presented the new approach to evaluating RAG systems, which underpin modern AI assistants.
➡️ Why Most AI Assistant Tests Don't Work in Reality?
The key idea is to move away from static tests to a dynamic environment with constantly updated data. The work was accepted at the EACL 2026 international conference.
Classic benchmarks quickly become outdated and poorly reflect real-world conditions. In business, AI works with live knowledge bases, where the relevance and coherence of facts are important, not just accuracy on a fixed dataset. DRAGOn proposes to test AI systems on fresh news, automatically collecting a "knowledge map" from them.
Instead of simple questions like "who/where/when", the system creates multi-level logical tasks. To answer, AI must match several facts from different news, not just copy a piece of text, and the neural network-judge checks the answers.
What this brings in practice:
- Tasks become multi-step, not trivial;
- The ability to link facts, not copy answers is tested;
- The evaluation takes into account completeness and factual accuracy, not just word matching.
Methodology can be deployed within a company and test AI on its own data before implementation. This allows comparing solutions in real scenarios and reducing the risk of errors, especially in tasks of analytics, support, and working with documents.
Article
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
