Retrieval-augmented generation is simple to demo and hard to run. The demo takes a folder of documents, cuts them into pieces of a few hundred tokens, embeds the pieces and answers questions from whichever pieces look most similar to the question. On a handful of tidy pages it works. On the material businesses actually hold, it returns fragments that look right and miss the answer.
Code is the clearest case, and it is the one we solved first, for DevGuardian AI. The lessons carry to any material that has structure.
Why token-count chunking fails on code
A chunk of 500 tokens cut from a source file is not a unit of anything. It starts in the middle of one method and ends in the middle of another. It carries no sense of what calls it or what it calls. A question about why a change is risky needs exactly that: the dependency structure. A diff alone does not show it, and a similarity search over fragments shows it even less. The reviewer who can tell you that a small change to a shared interface breaks four services elsewhere is reasoning over a graph, not over text that looks alike.
What we built instead
DevGuardian AI is a code review platform with three role-based agents, an Architect, a QA reviewer and a Security reviewer. Underneath them is a retrieval layer designed for code, with four parts.
Code-aware chunking. Source is parsed, not split. Using Roslyn, each chunk is a class, a method, an interface or a SQL statement: a unit of code with meaning, with metadata that says what it is and where it lives.
A dependency graph. The codebase’s dependencies are modelled in Neo4j. When an agent looks at a change, it can ask what else breaks, not only what looks similar. This is the GraphRAG part, and it is the part that gives the Architect agent its judgment.
Hybrid retrieval. Semantic similarity, symbolic matching on names and signatures, and graph traversal are combined for each query. Vector similarity is one signal among three, which is where it belongs.
Incremental refresh. The index is keyed to git commit hashes and updates only what a commit changed. A review never re-embeds the whole repository, which is what makes it affordable to run on every pull request rather than nightly.
Around that sits a Project Memory Bank: a persistent record of technical debt, architectural violations and review history, so each review starts from everything the system has already learned about that codebase. Vectors are stored in SQLite and background jobs run on Hangfire; the stack was chosen so one engineer could run it.
The same design, beyond code
Every business has material with structure that token-count chunking destroys:
| Material | The natural unit | The relationships that matter |
|---|---|---|
| Contracts and policies | Clause, section, definition | Which clauses reference which definitions; which policy supersedes which |
| Product catalogues | SKU, variant, specification | Compatibility, substitutes, bundles |
| Support knowledge | Procedure, step, known issue | Which issue a procedure resolves; which product version it applies to |
| Financial and operational records | Transaction, account, period | Who owes whom; what rolls up into what |
| Code | Class, method, statement | Calls, implements, depends on |
The method is the same each time. Parse the material into its real units. Model the relationships, in a graph where they are many-to-many. Retrieve with more than one signal. Re-index only what changed, keyed to the version of the source. And measure: a gold set of real questions with known answers, scored before and after every change to chunking, embedding model or ranking. Without the gold set nobody knows whether a change helped, and most RAG systems we are asked to look at have never had one.
Three questions to ask about your retrieval
- What is a chunk? If the answer is a number of tokens, the system does not know what your material is.
- What happens when one document changes? If the answer is a nightly rebuild, the system cannot be trusted between rebuilds and will cost more as the corpus grows.
- How do you know it got better? If there is no question set with known answers, nobody does.
We run a two-week retrieval assessment that answers all three on a sample of your material, with measured results. The service is described at RAG and knowledge systems.