.NETAI agentsCase study notes

Multi-agent systems in .NET: what we learned building two of them

DevGuardian AI runs three role-based review agents; MayAI runs one agent with six tools. Both are built on Microsoft Agent Framework in .NET. Seven lessons about when to add agents, how to design tools, and what has to be in place before the first demo.

Amir Pournasserian · October 8, 2026 · 4 min read

Most of the writing about agents is in Python. Most of the enterprises we work with are on .NET, with Azure underneath. Between 2024 and 2025 we built two agent systems on Microsoft Agent Framework in .NET: DevGuardian AI, a code review platform with three role-based agents, and MayAI, a consumer agent with one agent and six tools. These are the lessons that survived both.

1. Start with one agent

MayAI does a lot with one agent: it searches restaurants and dishes, filters by cuisine and diet, reads menus, manages a cart and remembers a user’s preferences. That is one agent and six tools, with a fixed reply schema and a memory per user. It needed nothing more.

DevGuardian AI has three agents because the work genuinely has three mandates. An Architect, a QA reviewer and a Security reviewer each own a different question about the same change, and a single agent asked to hold all three perspectives at once held none of them well. That is the test: add a role when the mandates differ, not when the task is big.

2. Tools are the product

The agent is the easy part. The tools are where the quality lives. Few of them, each named and described so the model picks the right one, each returning compact, well-typed results. In MayAI, a small service sat between the live restaurant data and the agent purely to reduce each response to plain records, because every field sent to the model costs time and money and most of them were never needed to choose a dish.

In .NET this is a strength: tools are strongly typed methods with schemas generated from the types, so the contract between agent and system is enforced by the compiler before it is enforced by the model.

3. Structured replies, grounded by the server

Every reply is a JSON object in a fixed schema, and anything the reply refers to, a dish, a file, a finding, is an identifier that the server resolves against what the tools returned in that request. An identifier no tool returned is dropped. This is the single most important design decision in both systems, and it is described in full in keeping AI agents honest.

4. Memory is two different things

There is the conversation, which has to be saved on the server and restored with every message so a user can pick up where they left off, and there is the lasting knowledge: a user’s allergies and budget habits in MayAI, a repository’s technical debt and review history in DevGuardian AI’s Project Memory Bank. The first is a transaction per turn, saved together with the user’s message and the agent’s reply, so a failed model call saves nothing. The second is a tool the agent calls deliberately, so what is remembered is a decision, not a side effect.

5. Retrieval is its own system

DevGuardian AI’s agents are only as good as what they can see. Its retrieval layer, code parsed with Roslyn into class- and method-level chunks, a dependency graph in Neo4j, hybrid search and incremental re-indexing keyed to commits, was more work than the agents and worth more. The reasoning is in RAG for code and other structured material.

6. Observability before the second feature

An agent system has a new failure mode: it did the wrong reasonable thing, slowly, at a cost. You need a trace that follows one user request through every tool call and every model call, with the tokens and the latency of each. We put OpenTelemetry in from the first prototype and set budgets per tool call. Our founder’s current work on a hospitality brand’s guest-facing multi-agent platform runs the same way, with Application Insights behind it.

7. The .NET ecosystem is ready, with one gap

Microsoft Agent Framework gave us the agent loop, tool calling and structured outputs without leaving C#. Azure OpenAI and OpenAI-compatible endpoints through OpenRouter both worked behind the same abstraction, so MayAI could switch models without touching the agent. Blazor gave DevGuardian AI its interface and SvelteKit gave MayAI its mobile-first app. Hangfire ran the background jobs; SQLite held vectors at a scale where a vector database would have been ceremony.

The gap is testing. The C# SDK for the Model Context Protocol is a first-tier SDK, yet as of October 2026 no .NET-native tool inspects, tests or load tests an MCP server; teams wrap the Node-based official Inspector. The detail is in testing an MCP server. We live with the wrapper for now.

What we would tell a team starting today

Build one agent with three good tools and a fixed reply schema. Put the trace in before the demo. Decide what the agent may do without approval and what it may not. Then, and only then, ask whether the work has more than one mandate. The service this comes from is AI agents and multi-agent systems.

Working on this?

These are the patterns we use on client work. Tell us what you are building and we will say how they apply.

Book an AI consultation