DocumentationAI agentsRAG

Documentation you can audit: observed, documented, inferred

AI-generated documentation reads beautifully, and nobody can tell which sentences came from the code and which the model dreamed up. Our founder compared seventeen tools that document a codebase, found the gap, and designed what he would build instead: facts with evidence and one of three labels.

Amir Pournasserian · October 8, 2026 · 4 min read

Based on published research

This note summarises research our founder published in full, with the comparison tables and method, at Observed, documented, inferred: a design for codebase documentation you can audit.

Most of what a team needs to know about a system is locked in its source code. Developers can read it, slowly. Managers, business stakeholders, support staff and coding agents cannot, so they rely on documents that are missing, out of date or written for someone else. A wave of AI tools now turns a repository into a wiki. On 1 October 2026 our founder compared seventeen of them, and then wrote down the design he would argue for instead. Both pieces are on his site; this note is what they mean for anyone who has to trust AI-written documentation.

The developer wiki is solved. Everyone else was left out.

Of the seventeen tools reviewed, hosted and open source, fourteen write a developer wiki with diagrams and chat, and eleven keep it up to date. Several run on your own machine with your own keys. That part of the market is a commodity, and the open-source bar is high.

The gaps are in who the wikis are for. One tool has a manager view, and it is an analytics product, not a documentation one. Two produce business documentation, and both are built for COBOL-era systems and sold through sales teams. And not one of the seventeen shows which statements were read from the code and which a model inferred. A reader gets fluent prose with no way to know what in it is true.

The design: one fact layer, many views

The alternative reads the code once into a fact base, where every fact carries its evidence and one of three labels, and then renders that fact base for each audience.

LabelMeaningExamples
ObservedDerived deterministically from code, configuration or historyA call edge, a complexity score, a churn count
DocumentedStated in human-written materialA doc comment, a README statement, an architecture decision record
InferredProduced by a modelA module summary, a capability description

Three rules keep the labels honest. The weakest label wins: a fact built from an observed edge and an inferred summary is inferred. No laundering: model output never becomes documented or observed, even on a later run, so the tool cannot promote its own guesses by re-reading them. One way in for human knowledge: human-written material in the source is the only documented input, and a docstring the tool proposes becomes documented only after a person merges it.

The labels appear in every output. No document, diagram or agent answer presents an inferred statement without saying so, and when evidence is missing the output says so instead of filling the gap.

Six principles behind it

Deterministic first, model last: indexing spends no model tokens, and a model is used only to summarise, group, name and narrate. One fact layer, many views: documents, diagrams and agent answers never read raw source, only facts. Evidence travels with every fact. Read-only: the source tree is never modified. Local first: the only thing that leaves the machine is text sent to the model the user configured. Language-neutral contracts, so the index and the tool surface can be implemented in any language.

Who gets what

ReaderWhat they needOutput
DevelopersHow the system is built and where to change itDeveloper docs, diagrams
ManagersSystem summary, health, risks, what changedManager brief
Business teamWhat the system does, in plain languageCapability catalogue, glossary
Operations and supportHow to run it and what its errors meanOperations runbook
Coding agentsStructured, cited facts on demandA knowledge base over MCP

When a file changes, only what depends on it is regenerated, and what went stale is reported. That is what makes a manager brief something that can stay current rather than a one-off summary.

Why this matters beyond documentation

The same problem sits inside every AI system that answers questions about a business: a fluent answer with no line between what was retrieved and what was composed. The labels are a general answer to it. We built the first version of this thinking into DevGuardian AI, whose Project Memory Bank keeps a persistent, evidenced record of a codebase’s technical debt and review history, and we apply the retrieval side of it in every knowledge system we build: the agent cites by identifier, the server renders from the evidence, and anything the model added is marked as the model’s.

The design, and the full comparison of the seventeen tools, are at the link above. It is a design, not a product, and every claim in it is written as a test someone can fail.

Working on this?

These are the patterns we use on client work. Tell us what you are building and we will say how they apply.

Book an AI consultation