Standard LLMs hallucinate on large C++ codebases. I built an open-source MCP server using tree-sitter to fix that.
I’ve been experimenting heavily with AI agents for legacy modernization, but I kept hitting a massive wall when dealing with C++.
If you just dump a bunch of intertwined C++ files into Claude or an open-source model, it loses context. It tries to refactor a function without knowing what upstream dependencies rely on it, or it gets tangled in circular references and breaks the build. Regex-based parsers also constantly miss the nuances of C++ macros and templates.
So, I built **LegacyGraph-MCP**, an open-source Model Context Protocol (MCP) server that turns "spaghetti code" into a queryable Knowledge Graph for AI Agents.
**How it works under the hood:** Instead of forcing the LLM to read raw text and guess relationships, the agent queries the structure.
1. **AST Parsing:** It uses `tree-sitter-cpp` for 100% accurate parsing (no regex hacks).
2. **Graph RAG:** It loads the codebase into a NetworkX directed graph.
3. **Agent Interaction:** The LLM uses MCP tools to ask specific questions like:
* *“Which functions call* `process_client()`\*?”\*
* *“Are there any circular dependencies here?”*
* *“What are the orphan functions?”*
**Features:**
* **Omni-Ingestion:** It can clone a GitHub repo, apply a patch, scan a local directory, or accept raw file uploads.
* **Hybrid Deployment:** You can run it locally (`stdio`) directly tied to your IDE/Claude Desktop, or deploy it to the cloud (`HTTP`/`SSE`) with ephemeral `/tmp/` clones.
* **Token-Safe Visuals:** It dynamically generates bounded Mermaid.js graphs inline so the LLM can visually map out the dependencies it is working on.
**The Ask:** The core graph logic is working well (I’ve got an end-to-end verifier passing 100% on test data), but I’m looking for feedback from people who work with *really* messy, real-world legacy code.
Specifically:
1. Does the cycle detection logic hold up against your worst circular dependencies?
2. Are there specific query tools (besides `get_callers`, `get_callees`, `detect_cycles`, etc.) that would make your life easier when using an AI coding assistant?
**Repository:** [LegacyGraph-MCP](https://github.com/RohitYadav34980/LegacyGraph-MCP)
*(If you are using Smithery or Claude Desktop, I included the quick-start configs in the README).*
What else can we build with this?
I am open to all ideas! Could we use this for automated test scaffolding for isolated subgraphs? Dead-code pruning pipelines? Vulnerability blast-radius analysis? Drop your ideas in the comments.
Whether you are a backend engineer optimizing graph algorithms, or someone pushing the limits of AI agents, there is a place for you here. Let’s build the ultimate open-source legacy modernization engine together.
Would love to hear your thoughts or see if this breaks on your weirdest legacy code!
https://redd.it/1rzws51
@r_cpp
Post #24873
35