Skip to content
←Projects
17 May 20263 min readIn beta

leo-mcp - the real code map for AI agents.

An MCP server that gives Claude Code or Cursor the real call graph of a repository: where each function is, who calls it and what breaks if you change it. Verified against the Python AST and the TypeScript compiler.

Pythontree-sitterMCPAST
What it is
An MCP server that gives coding agents the real call graph of a repository.
My role
Personal project, solo: design, development, tests and continuous integration.
Stack
Python · tree-sitter · MCP · AST
Result
100% precision in Python and 99.95% in TypeScript, verified on every change · with real agents, 30% fewer tokens and tasks 14% faster.
Status
In beta. Installs from the repository.
Links
View code ↗ (opens in a new tab)

Without a map, a coding agent searches text and reads whole files: it spends tokens on code it doesn't need and answers questions about structure by guessing. leo-mcp gives it the real call graph of the repository, with file and line in every answer.

It started as an agent; I kept what added value

It began in May as my own coding agent on top of opencode, with a way of retrieving code I called KC-RAG (I explained it in this article). I kept the part that really added value, the code map, and turned it into an MCP server, the protocol Claude Code, Cursor, Codex and opencode already speak.

What already works

The graph tool answers questions about structure without using any model:

  • where: where a symbol is defined.
  • who_calls: who calls it.
  • impact: everything that depends on it; in other words, what breaks if you change it.
  • trace: the path from A to B, including the jump from a fetch('/api/x') in the frontend to the backend route.
  • guard: like impact, but flagging the affected parts that have no test.

get_context takes a question in plain language and returns only the piece of the graph that answers it, compressed: between 80 and 97% less than reading the whole files, in my tests.

It includes a command line to index a repo and add it to a project in one step. The index lives in the user's cache, never inside your project, and only what changes is analysed again.

It warns before a risky edit

In Claude Code it hooks in before every edit: if the agent is about to touch a function that untested code depends on, it tells the agent and you, in one line. The rest of the time it stays quiet, on purpose: a warning that always fires ends up switched off.

It takes about 0.28 seconds and never blocks a change; that is covered by tests. No warning isn't a guarantee either: it's a static graph, and a dependency through a queue, a WebSocket or a name built from text doesn't show up.

Verified, not estimated

On every change, continuous integration rebuilds all the relationships with tools that aren't mine and compares them one by one. If the graph gets worse, the change doesn't go in.

  • Python: compared with Python's own ast on about 100,000 symbols and 480,000 relationships, from this repo and three external projects. 100% precision and coverage.
  • TypeScript and JavaScript: compared with the tsc compiler on the TypeScript repository itself, with 36,739 symbols. 99.95% precision and 99.93% coverage. I leave it unrounded.
  • Speed: with 271,493 symbols indexed, the slowest query takes under 2 milliseconds.

With real agents: 30% fewer tokens

With opencode, on 15 tasks repeated three times and counting subagents too: answers as good or better, 14% less time and 30% fewer tokens overall. The savings are in heavy tasks; on a normal one, spend stays about the same.

On someone else's project, with FastAPI and Next.js, it installed in 6 seconds, indexed in one and wrote nothing inside it. Claude Code answered the same questions with 26% fewer tokens, especially the structural ones. That test also uncovered several installation bugs, which I fixed.

What's missing

It isn't on npm or PyPI yet: it installs from the repository, which already has a template for collecting feedback. Python and TypeScript are parsed with their full syntax tree; Go, Java, Rust, PHP and others, approximately.

What I learned: measuring forces you to cut

I wanted to put "40% fewer tokens" in the README. I tried six configurations, none got there, and I took it out. What's left is smaller, but every figure can be reproduced with the commands in the repository.

Built byIsmael Manzano LeónFull Stack Developer · leo/ · leosoftware.dev
View on GitHub→ (opens in a new tab)

Does this fit what you are looking for? Write to me →