A Model Context Protocol server is used by a language model, not by a person, so testing it means more than calling each tool by hand. Does it follow the specification? Can an agent pick the right tool from the descriptions? Will it pass review for the Claude connector directory or the ChatGPT app store, and work inside Cursor, VS Code and the command-line agents? How does it behave under load? On 1 October 2026 our founder reviewed about eighty tools that answer those questions and compared thirty-one of them capability by capability. The comparison tables and the market map are on his site. This note is the short version, and what we do about the gaps on client work.
Six jobs
A tool in this space does one or more of six jobs, with observability sitting beside them on the production side:
- Inspect a server by hand.
- Run functional tests and evaluations against it: does the model pick the right tool with the right arguments?
- Audit its design and protocol conformance.
- Check it against the rules of the hosts that will run it.
- Scan it for security problems.
- Load test it.
What the landscape looks like
Every capability exists somewhere, but no open, local tool combines them. Inspection, evaluations, audits, host checks and load testing are spread over about a dozen products.
MCPJam is the closest thing to a complete workbench. It ships saved evaluation suites, AI-generated test cases the user keeps or discards, a language-model judge, client emulation, directory-readiness checks and CI gates. Its AI features depend on its hosted backend and credits.
The official MCP Inspector is no longer a slow target. Version 2 has been generally available since July 2026 and ships about weekly. Its roadmap for November 2026 to January 2027 has saved collections, assertions, CI flows and a conformance runner. It has no language-model or evaluation work planned.
Three areas are saturated: manual inspection, security scanning, and the basic combination of YAML evaluations, a model judge and a GitHub Action. Protocol conformance is owned by the official suite.
The protocol changed fundamentally two months before the review. The 2026-07-28 specification is a stateless redesign, the fifth revision in about twenty months and a breaking one. Hosts lag it, so a tester has to speak both eras, and many small tools are stuck on older versions. Several went quiet within three to six months of launch.
.NET is a gap of its own. The C# SDK is a first-tier SDK and reached version 2.0 in July 2026, yet no .NET-native inspector, test or load tool was found. Teams on .NET wrap the Node-based official Inspector.
The gaps
The research names six: a fully local evaluation loop that runs on a model you choose; an open, versioned rule pack for many hosts backed by live behavioural checks; fix suggestions proven by re-scoring; load testing tied to functional scenarios rather than raw request rates; one combined report; and anything native to .NET.
What we do on client work
Until a single tool covers it, we assemble the six jobs ourselves and keep the results in one place.
- Conformance with the official suite, against both protocol eras, in CI.
- Host rules from our own dated rule pack, derived from the eight rulebooks, checked statically on every build and with live calls before a release.
- Evaluations from a gold set of real user requests, each with the tool and arguments a correct agent should choose, scored before and after any change to a description or a schema. The gold set is the most valuable artefact of the project, and we hand it over.
- Security scans plus our own injection cases: tool results that contain instructions, descriptions that try to redirect the model, inputs that smuggle a second request.
- Load tied to the scenarios in the gold set, so the numbers describe what agents will actually do, not an abstract request rate.
- Observability with OpenTelemetry from the first prototype, so a trace follows one user request through every tool call and model call.
The .NET gap matters to us because much of our work is in .NET. For now we wrap the official Inspector; we are watching the space closely, and the research above is how we keep that watch honest.
The service these checks belong to is described at MCP servers and AI integrations.