AI agentsCase study notes.NET

Keeping AI agents honest: the server-side grounding pattern

An agent that puts real items, prices or actions in front of a user cannot be allowed to invent any of them. The pattern we built for MayAI makes that impossible by construction rather than by prompt. Here is how it works and where it applies.

Amir Pournasserian · October 8, 2026 · 4 min read

When we built MayAI, a consumer agent that turns one chat message into a food order, the hardest requirement was not finding restaurants. It was making sure the agent could never show a dish that was not on the menu or a price that was not real. A hallucinated paragraph in a chatbot is embarrassing. A hallucinated item in an order is an order nobody can fulfil.

Prompts help with this. They do not hold. The pattern that held is one we now use in every agent that puts real items, prices or actions in front of a person. We call it server-side grounding.

The rule: the model returns IDs, the server builds the screen

Every reply from the agent is a JSON object in a fixed schema. It contains the agent’s reasoning, a message for the user, up to five quick-reply options, and the identifiers of the items it wants to show. It does not contain the items.

The server then looks each identifier up in what the tools returned during that same request. A restaurant ID that came back from the search tool becomes a restaurant card, built from the tool’s data: the real name, the real menu, the real price. An ID that no tool returned in this request is dropped. It cannot appear on screen, because there is nothing to build the card from.

That is the whole trick. The model is free to reason, to choose and to phrase. It is not free to assert facts, because facts are rendered only from tool output.

Six supporting rules

The core rule is not enough on its own. These are the rules around it that made MayAI’s prototype behave.

Act first, then speak. The agent’s instructions tell it to call the tool, read the result and only then answer, and never to announce what it is about to do. “Let me look into that” followed by nothing is worse than no answer, and an agent left to its own devices produces a lot of it.

Ask when it is not sure. “Add the burger” is clear when one burger is on screen and a guess when there are three. The agent leaves the cart alone and asks which one. If a user asks for a change the agent cannot pass on, such as “no onions” to a menu that has no such option, it adds the dish and says so.

Widen an empty search and say what you did. An empty result is not a dead end. The agent tries a broader search and tells the user that it did.

Reason before answering, in a field nobody sees. Before the message, the agent fills in a reasoning field: what the user wants, what it already knows, what could go wrong, which tool gets there fastest. The app never shows it. It exists to make the next field better.

Small tool results. Restaurant and menu data arrives large, nested and inconsistent. A small service of our own sat between the data source and the agent and reduced each response to plain records, stripped of images, tracking fields and banners. Every field sent to the model costs time and money, and most of them were not needed to choose a dish.

One transaction per turn. The user’s message, the agent’s reply and the updated conversation are saved together. If the model call fails, none of it is saved, so a conversation can never be left half-written.

Why this beats prompting

A prompt is a request. A schema with server-side resolution is a constraint. The difference shows up at the edges: a model under a long conversation, a model swapped for a cheaper one, a model updated by its provider overnight. Prompts drift with all three. The server’s lookup does not care which model produced the ID.

It also makes the system testable. You can replay a conversation with a stub model that returns deliberately bad IDs and confirm that nothing invented renders. Try writing that test for a prompt.

Where it applies

Anywhere the agent’s output refers to something that exists elsewhere:

  • Commerce and booking. Products, prices, rooms, flights, slots.
  • Support and operations. Tickets, orders, customer records, the next step in a procedure.
  • Internal assistants. Documents, policies, people. The agent cites by ID; the server renders the citation from the real document.
  • Code review. Findings refer to real files and lines, resolved from the retrieval layer, not typed by the model.

The cost is a little more schema work up front and a reply format that is less free-form. In exchange, the question “did the agent make that up?” has a mechanical answer: it could not have.

The MayAI prototype was built in .NET on Microsoft Agent Framework by an architect and one developer. The wider lessons from that and from DevGuardian AI are in multi-agent systems in .NET.

Working on this?

These are the patterns we use on client work. Tell us what you are building and we will say how they apply.

Book an AI consultation