MCPAI agents

What each AI host requires of an MCP server

A server that follows the MCP specification can still be rejected by the Claude directory, turned away from the ChatGPT app store, or have its tool names silently cut short by Gemini CLI. The rules of eight hosts, where they conflict, and what that means for how we build.

Amir Pournasserian · October 8, 2026 · 4 min read

Based on published research

This note summarises research our founder published in full, with the comparison tables and method, at One MCP server, eight rulebooks.

The Model Context Protocol has one specification and many hosts. Each host that runs MCP servers, from the Claude connector directory to Gemini CLI, publishes rules of its own, and they do not agree with each other. While reviewing the tools that test MCP servers, our founder collected the rules as they stood on 1 October 2026. The full rulebook, host by host with sources, is on his site. This note is what we took from it for client work.

Eight rulebooks, one server

HostExamples of rules it adds
MCP specificationTool names of 1 to 128 characters; a valid input schema; input errors returned as results, not failures
Claude connector directoryNames of 64 characters or fewer; every tool has a title plus a read-only or destructive hint; read and write split into separate tools; no prompt-injection patterns; reasonably sized responses
Claude apps and Claude CodeAbout 150,000 characters per tool result in the apps; 25,000 tokens by default in Claude Code
ChatGPT appsRead-only, destructive and open-world hints on every tool, the absence of which is “a common cause of rejection”; non-promotional descriptions; minimal inputs
CodexA 10-second startup timeout and a 60-second tool timeout by default
VS Code with CopilotAt most 128 tools enabled per chat request
Windsurf100 tools in total
KiroNames of at most 64 characters including the server prefix, letters, digits and underscores only, starting with a letter
Gemini CLINames over 63 characters are truncated; some schema keywords are stripped

Where they conflict

Name length. The specification allows 128 characters. The Claude directory and Kiro stop at 64, and Kiro counts the server prefix. Gemini CLI does not reject a longer name; it truncates at 63, which is worse, because the server never finds out.

Characters. The specification permits hyphens and dots in names. Kiro permits neither. A name that is valid everywhere else fails there.

Annotations. The read-only, destructive and open-world hints are optional in the specification. Claude’s directory requires a title and a read-only or destructive hint on every tool. ChatGPT requires all three and names their absence as a common reason for rejection.

Budgets. Hosts cap different things: VS Code counts tools per chat request, Windsurf counts tools in total, Codex counts seconds, Claude counts the size of a result. A server built for one of these can exceed another without changing a line.

What a machine can check, and what it cannot

Roughly half of the rules can be checked statically from the tool list: name length and characters, the presence of annotations and titles, a valid input schema, the tool count. About a quarter need live calls: that every tool succeeds with valid input, that errors come back as results rather than generic failures, that responses stay within a host’s size limit, that the server starts within Codex’s ten seconds.

Most of what reviewers actually reject needs judgment: whether a description matches what the tool does, whether a read-only hint is honest, whether a description is promotional, whether anything in it reads as an injection. That is work for a language model with a rubric, with a person behind the verdict.

The rules also move. The host documentation was each read once on one day, and a validator that hard-codes them will rot. The research argues they have to be carried as versioned data, each rule with its source and the date it was read.

Descriptions are where servers fail

The rule that matters most is the one no regular expression can check. A study titled “Tool descriptions are smelly” found at least one defect in 97.1% of the 856 tool descriptions it examined, and poor descriptions reduce the chance that an agent selects the right tool. The most developed yardstick is Glama’s Tool Definition Quality Score, which grades six dimensions of a definition and is already run continuously across a registry of more than fifteen thousand servers, so maintainers are being graded whether they know it or not.

What validates against the rules today

The validators that exist are closed, paid or partial: Alpic Beacon, Manufact’s publishing checks and MCPJam’s hosted readiness runs. Nothing open covers the IDE and command-line hosts.

How this changes the way we build

  • We design tool names to the strictest rule, 64 characters with no hyphens or dots, so one server passes everywhere.
  • Every tool gets a title and all three hints from the first commit, with the hints checked against what the tool actually does.
  • Read and write are separate tools, always.
  • Result sizes are budgeted against Claude Code’s default, the smallest of the published limits.
  • Descriptions are written and reviewed as product copy for a model, then scored before submission.
  • The host rules we check against are kept as dated data in the project, not as code.

What this looks like on a real engagement is described at MCP servers and AI integrations.

Working on this?

These are the patterns we use on client work. Tell us what you are building and we will say how they apply.

Book an AI consultation