Skip to content
←Blog
June 23, 20263 min read

Why chunking breaks your coding agent (and almost nobody notices)

With RAG over code, the problem is almost never the model but how we split the code. What I learned by breaking it.

RAGAgentsLLMAST
Why chunking breaks your coding agent (and almost nobody notices)

I've spent a while working on how to give context to coding agents, and there was one thing I couldn't explain: why the model sometimes answers as if half the file were missing. At first I thought it was the model, that it needed a bigger or more expensive one, but I was wrong. The problem is almost never the model: it's chunking, the way we split the code before passing it on.

The root mistake: treating code as if it were text

Classic RAG was born for documents, and the recipe is always the same: you split the text into fixed-size pieces - about 500 tokens -, compute an embedding for each piece and store them in a vector store. When you ask something, you retrieve the pieces most "similar" to your question.

The first time I set it up for code, it seemed logical. It took me a while to realise that it works with prose, but with code it fails silently - and that's the worst part, because it doesn't warn you.

The reason, once you see it, is simple: code isn't plain text. You can cut a novel almost anywhere and both halves still make sense. Not code. It has a syntactic structure - what's called an AST, the syntax tree - and a dependency graph between its pieces. If you cut by counting tokens, you destroy exactly that.

Symptom 1: functions cut in half

A 500-token chunk cuts where it happens to land, not where the function ends. So the model gets the first half of verifyUser() and nothing else. And here's the curious part: it doesn't know code is missing, so it makes up the rest. We call it "hallucination", but it isn't its fault - you gave it half a function.

Symptom 2: dependencies left out

Say verifyUser() calls hashPassword(), and the real bug is in hashPassword. The search brings back verifyUser because it matches your question, but hashPassword is in another piece that looks nothing like what you typed. Result: you never see it, and neither does the model.

This is what size-based chunking can't see: the "who calls whom" relationship. It treats each piece as an island. But to understand a function you almost always need its direct dependencies - what it calls and who calls it -, and that lives in the graph, not in how similar two strings of text are.

Symptom 3: semantic similarity plays tricks on you

You search for verify user and the vector store happily returns verifyEmail, verifyToken, validateUser. Similar names, close embeddings, code that's no use at all. Two names being alike doesn't mean the code is relevant - and in programming, similar names are everywhere.

What I took away

Size-based chunking destroys exactly what makes code mean something: its structure and its relationships. It optimises for "comfortably sized pieces" when it should optimise for meaningful pieces and their dependencies.

Retrieving "similar pieces" is simply the wrong question. The right one - what to retrieve instead, and how to use the AST and the graph to get it - is exactly what I'm working on now. But that's another article.

Written byIsmael Manzano LeónFull Stack Developer · leo/ · leosoftware.dev
Talk to Ismael→