Without a map, a coding agent searches text and reads whole files: it spends tokens on code it doesn't need and answers questions about structure by guessing. leo-mcp gives it the real call graph of the repository, with file and line in every answer.
It started as an agent; I kept what added value
It began in May as my own coding agent on top of opencode, with a way of retrieving code I called KC-RAG (I explained it in this article). I kept the part that really added value, the code map, and turned it into an MCP server, the protocol Claude Code, Cursor, Codex and opencode already speak.
What already works
The graph tool answers questions about structure without using any model:
- where: where a symbol is defined.
- who_calls: who calls it.
- impact: everything that depends on it; in other words, what breaks if you change it.
- trace: the path from A to B, including the jump from a
fetch('/api/x')in the frontend to the backend route. - guard: like impact, but flagging the affected parts that have no test.
get_context takes a question in plain language and returns only the piece of the graph that answers it, compressed: between 80 and 97% less than reading the whole files, in my tests.
It includes a command line to index a repo and add it to a project in one step. The index lives in the user's cache, never inside your project, and only what changes is analysed again.
It warns before a risky edit
In Claude Code it hooks in before every edit: if the agent is about to touch a function that untested code depends on, it tells the agent and you, in one line. The rest of the time it stays quiet, on purpose: a warning that always fires ends up switched off.
It takes about 0.28 seconds and never blocks a change; that is covered by tests. No warning isn't a guarantee either: it's a static graph, and a dependency through a queue, a WebSocket or a name built from text doesn't show up.
Verified, not estimated
On every change, continuous integration rebuilds all the relationships with tools that aren't mine and compares them one by one. If the graph gets worse, the change doesn't go in.
- Python: compared with Python's own
aston about 100,000 symbols and 480,000 relationships, from this repo and three external projects. 100% precision and coverage. - TypeScript and JavaScript: compared with the
tsccompiler on the TypeScript repository itself, with 36,739 symbols. 99.95% precision and 99.93% coverage. I leave it unrounded. - Speed: with 271,493 symbols indexed, the slowest query takes under 2 milliseconds.
With real agents: 30% fewer tokens
With opencode, on 15 tasks repeated three times and counting subagents too: answers as good or better, 14% less time and 30% fewer tokens overall. The savings are in heavy tasks; on a normal one, spend stays about the same.
On someone else's project, with FastAPI and Next.js, it installed in 6 seconds, indexed in one and wrote nothing inside it. Claude Code answered the same questions with 26% fewer tokens, especially the structural ones. That test also uncovered several installation bugs, which I fixed.
What's missing
It isn't on npm or PyPI yet: it installs from the repository, which already has a template for collecting feedback. Python and TypeScript are parsed with their full syntax tree; Go, Java, Rust, PHP and others, approximately.
What I learned: measuring forces you to cut
I wanted to put "40% fewer tokens" in the README. I tried six configurations, none got there, and I took it out. What's left is smaller, but every figure can be reproduced with the commands in the repository.