In Why chunking breaks your coding agent I described the problem: if you split code by size, the agent gets half functions, loses what they call inside and ends up bringing back code that only looks similar by name. Here's the part I still had to tell: what I did instead.
The underlying idea isn't new: code isn't plain text, it has structure. Functions have names and parameters, they call each other and modules import things from others. If you respect that when you store it, the agent gets pieces that make sense. I call it KC-RAG, for Knowledge Context, and it has four parts.
1. Cut the code where the code itself splits
A program, seen from inside, is a tree: the AST. Each function is a branch with its name, its parameters, its documentation and its body. If you walk that tree instead of counting tokens, you can pull out each function and each class whole. No more verifyUser() cut in half.
You don't have to invent anything for this: Python ships its own ast module, and tree-sitter covers a huge number of languages. I read the code, get the tree and keep each function and class as a complete piece, with its name, parameters, return value and documentation.
2. Know who calls whom
This is the part chunking doesn't have. With the pieces extracted, I go through the body of each one and note which other functions it calls. I store it both ways: verifyUser calls hashPassword, and hashPassword is called by verifyUser.
That changes the result a lot. If you ask "how do I validate users?", I no longer stop at verifyUser(): I also bring what it uses inside, which is exactly where the bug often is. Not because it looks like the question, but because it's genuinely connected.
3. Search by name and by meaning at once
Searching by meaning alone has a catch: verifyEmail and verifyUser "look" very similar to a model, but they're different functions. So I search two ways at once: by exact name, which doesn't get it wrong, and by meaning, for when you ask in your own words. Exact results come first and meaning-based ones fill in the rest.
4. Don't always give the same context
Understanding a function isn't the same as changing it or writing something new, so it makes no sense to always pass the same thing:
- To understand something: the whole function and only the names of what it calls.
- To change it: the function, the full code of what it uses inside and who calls it, because that's what can break.
- To write something new: almost no logic, just the structure of the project and which functions exist.
- To explore: a list of functions per file, to know what's there.
How it all fits together
- I walk the code and pull out each function and class whole.
- I build the map of who calls whom.
- I store everything, with its data, so it can be searched later.
- When a question arrives, I look at what kind of task it is.
- I search by name and by meaning at once.
- I follow the map as far as that task needs.
- I trim the result to what's needed and pass it to the agent.
One thing that mattered to me: storing the code doesn't depend on any model. It's a walk over the tree, it always gives the same result and it can't make anything up. The only thing that changes with the question is how much context I give.
An example: adding Google login to login()
Splitting the code by size: the whole auth.py is opened, "login oauth" is searched by similarity and loose pieces arrive: login() cut off and bits of hashPassword and the session, unrelated to each other. The model can't see where credentials are validated and proposes something that skips it. Counting the failed attempt, about 8,000 tokens.
With KC-RAG: it's detected as a change, login() is found by name and the map is followed to hashPassword, validateEmail and createSession. I pass the full code of those, and only the names of the rest. The model sees the whole logic and adds Google login without breaking validation. About 2,500 tokens in total.
The numbers measured for real, with a real agent and real tasks, are in the next article.
When it's worth it
It's worth it in large projects, with more than a hundred functions, because there the cost of building the map pays off quickly. Also when a mistake is expensive, as in security changes, and when token spend matters.
It isn't needed for small scripts, for new code that doesn't depend on anything existing, or when what you're looking for is documentation rather than code.
