A typical situation: a company has an internal knowledge base, but employees still go to HR with questions. Because messaging a real person in chat is faster than searching for answers in the wiki.
At Evrone, we embedded an RAG-based AI assistant directly into the Evrone ERP personal account. The chat is on every page – employees ask questions and get answers without switching between windows.
On the technical side: the frontend is built with React, communication runs through WebSockets (asynchronous exchange and real-time response streaming). The backend uses Python with FastAPI – we chose a lightweight framework for performance.
The wiki is written in Markdown. We split the text into chunks based on headings (header-based splitting). Semantic search uses vector embeddings – employees can phrase questions naturally, and the model searches by meaning.
We use Qwen as the LLM. We had already used it before for a smart resume parser. On this project, we tested several models, including popular open-source and a number of Russian solutions. We evaluated response quality, inference speed, and token cost. Qwen showed the best results across all factors: it follows instructions accurately, understands Russian-language queries well, and is more cost-effective than the other options.
The system automatically tracks changes in the wiki: a Python script runs on a schedule or via a Webhook trigger, then rebuilds the vector index. A custom React-based admin panel allows adjusting the number of chunks, the LLM, the system prompt, and generation parameters on the fly.
Security: the infrastructure is in-house, so no data leaks occur. Topic filtering ensures the model only answers questions related to company documentation. Rate limiting is applied at the API level.
The architecture is universal: it can be adapted for customer support, a consultant on a website, or an AI agent inside your own ERP.
You can learn more about how we did this here.
Post #712
214

- 🔥 3
- 👍 1