🔴 Why old approach limited?
Cursor has long used a retrieval mechanism: the agent searches the codebase and adds the necessary pieces to the LLM context.
Previously, it was just a grep version, searching by string matching. This is fast but not always sufficiently relevant.
Now it has been replaced by a smarter semantic search. Essentially, RAG. That is, the relevance of code snippets is now evaluated by a special vector model that no longer just searches by keywords but matches meanings.
🟢 How was embedding model trained?
Cursor trained its own embedding model specifically tailored for code. Real agent work trajectories were used for this. Each session is a sequence: query -> search for relevant code snippets -> result. A separate LLM evaluated these trajectories to determine which found snippets were ultimately useful and which were noise.
Then we take our vector model and train it on triplets (query, relevant files, irrelevant files) so that its ranking matches the LLM's ranking, meaning more useful snippets are closer to the query in vector space.
By the way, grep search still remains somewhere: for example, it is indispensable when you need to quickly search by variable or function names. The results of the grep module and the vector model are combined.
What about the metrics in the end:
1️⃣ On offline evaluation on the specially collected "Cursor Context Bench" benchmark, the average accuracy improvement was about 12.5%.
2️⃣ In A/B tests, code retention increased on average by about 0.3%. This metric shows how much code generated by the agent remains in the user's project over time. On large codebases, there was even a +2.6% increase.
3️⃣ Also, the number of dissatisfied follow-up requests decreased by about 2.2% – when the user has to make corrections or additional requests because the agent failed on the first try.
The effect is not huge because not every query requires a search, but it exists and will be especially noticeable in large codebases.
🤖 Data Science, ML & Big Data with @DataXplore
