AI services

RAG & Knowledge Systems

Retrieval built for your material, not generic chunking: hybrid search, GraphRAG, code-aware indexing and incremental refresh on pgvector, Azure AI Search or Neo4j.

At a glance

  • Hybrid retrieval: semantic, lexical and graph signals combined
  • Chunking that follows the structure of the material
  • Incremental re-indexing keyed to what changed
  • Evaluation sets, so retrieval quality is measured, not assumed

Retrieval-augmented generation is simple to demo and hard to run. Most RAG systems split documents by token count, embed everything, and hope vector similarity finds the right passage. On real material, such as contracts, code, product catalogues or policy manuals, that approach returns plausible fragments and misses the answer. We build retrieval the way the material is actually structured.

Who it's for

  • Companies with a body of documents, records or code that people and agents need accurate answers from
  • Product teams whose assistant answers confidently and wrongly
  • Engineering organisations that want AI to reason about a codebase, not just search it

What we build

  • Chunking that follows structure. Clauses, sections, tables, classes, methods, SQL statements: each chunk is a unit of meaning, with the metadata to filter on.
  • Hybrid retrieval. Semantic search, lexical search (Lucene, Azure AI Search, PostgreSQL full text) and graph traversal combined per query, so vector similarity is one signal among several.
  • GraphRAG. Dependency and relationship graphs on Neo4j, so a question about one thing can pull in what it depends on and what depends on it.
  • Incremental refresh. Indexes keyed to document versions or git commits that re-embed only what changed. A nightly full rebuild is a warning sign.
  • Cache-augmented generation where the corpus is small enough to keep in context and the cost model says so.
  • Evaluation. A question set with known answers, measured before and after every change to chunking, embedding model or ranking.

How it's done

We begin with your questions and a small gold set of answers, then build the index to serve them. Storage is pgvector, Azure AI Search or an in-memory store depending on scale; embeddings from Azure OpenAI or open models; ranking tuned against the gold set. Everything is observable: which chunks were retrieved, why, and what the model did with them.

Proof

DevGuardian AI: code-aware chunking, a Neo4j dependency graph, incremental embedding refresh keyed to commits, and hybrid retrieval, so a review agent can see what else breaks when a piece of code changes. Valco AI: a conversational layer over a modelled market database rather than a pile of documents.

Engagement shape

A two-week retrieval assessment on a sample of your material, with measured results, then the build.

Frequently asked questions

Do we need a vector database?

Often not a separate one. pgvector inside PostgreSQL or Azure AI Search covers most workloads and keeps retrieval next to the data it indexes. We add Neo4j when relationships matter, and a dedicated vector store only at a scale that needs it.

Why is our current RAG answering wrongly?

Usually one of three things: chunks that cut through the meaning, retrieval that returns similar-looking but wrong passages, or no measurement, so nobody knows which. An assessment with a gold question set finds out in a week.

Can this work on source code?

Yes, and it needs its own design: parse the code into classes, methods and statements, model the dependency graph, and re-index per commit. That is the retrieval layer we built for DevGuardian AI.

Talk to the people who would build it

Tell us what you are trying to do with rag & knowledge systems. We will come back with an honest take and a plan.

Book an AI consultation