TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 578 subscribers
Post #2104 172
DRAGOn methodology.

Russian researchers from SberAI, MWS AI, ITMO and HSE University, have presented the new approach to evaluating RAG systems, which underpin modern AI assistants.

➡️ Why Most AI Assistant Tests Don't Work in Reality?

The key idea is to move away from static tests to a dynamic environment with constantly updated data. The work was accepted at the EACL 2026 international conference.

Classic benchmarks quickly become outdated and poorly reflect real-world conditions. In business, AI works with live knowledge bases, where the relevance and coherence of facts are important, not just accuracy on a fixed dataset. DRAGOn proposes to test AI systems on fresh news, automatically collecting a "knowledge map" from them.

Instead of simple questions like "who/where/when", the system creates multi-level logical tasks. To answer, AI must match several facts from different news, not just copy a piece of text, and the neural network-judge checks the answers.

What this brings in practice:
- Tasks become multi-step, not trivial;
- The ability to link facts, not copy answers is tested;
- The evaluation takes into account completeness and factual accuracy, not just word matching.


Methodology can be deployed within a company and test AI on its own data before implementation. This allows comparing solutions in real scenarios and reducing the risk of errors, especially in tasks of analytics, support, and working with documents.

Article

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  2. Sep 14, 2026Post #2188
  3. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  4. Aug 22, 2026Post #2185
  5. Aug 21, 2026Deep systemic analysis of AI constraints from context to internal weight editing. 📂 PDF #…
  6. Aug 17, 2026Adaptive Gradient Thresholding Why Fixed Gradient Clipping Kills Deep RecSys When Feedback…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →