A practical guide to designing MCP tool schemas that agents can call reliably, covering enums, flat structures, parameter limits, usage examples, and the tradeoff between accuracy and context cost.
October 04, 2026

Anthropic’s own tool-use guidance reports a measurement worth sitting with: adding worked usage examples to tool definitions raised complex parameter-handling accuracy from 72 percent to 90 percent in internal testing. Nothing about the tool changed. The handler was identical, the parameters were identical, the model was identical. The only difference was how the tool described itself. That 18-point swing is the clearest available evidence that a tool schema is not documentation sitting next to the real interface. It is the interface, because the agent has nothing else to reason from, and a schema that is structurally correct but poorly described produces an agent that calls the right tool with the wrong arguments.
The schema is the agent’s entire view of the tool. Every piece of text in a tool definition, from the name to the parameter descriptions, is part of the model’s reasoning context, and it is all the model has when deciding whether to call, which to call, and what to pass.
A parameter typed as a bare string invites the agent to guess. Constraining finite-value fields with enums replaces guessing with selection, and the published guidance is consistent that agents guess wrong at a meaningful rate when the valid values are not in the schema.
GMI Cloud’s MCP server exposes a deliberately small tool set in v1, which reflects the same constraint from the other direction: selection accuracy degrades as the number of visible tools grows, so tool-set size is part of schema design rather than separate from it.
Nesting is the structural choice that costs the most accuracy. Complex nested objects confuse models and increase hallucination risk, which is why the consistent recommendation is to flatten request bodies into top-level parameters wherever the flatter design achieves the same thing.
Parameter count guidance converges around a small number. Google’s MCP Toolbox style guide targets fewer than five parameters per tool; AWS Prescriptive Guidance suggests around eight or fewer.
Every token in a tool definition is paid on every request, whether the tool is called or not. This is the tension that makes schema design a genuine tradeoff rather than a matter of adding more description until accuracy improves.
The useful mental model is that a tool definition is a prompt fragment that happens to be structured. The agent reads it the way it reads anything else in context, and its decisions follow from what that text says.
Three distinct decisions depend on it. Whether to call a tool at all, which requires the description to communicate purpose and scope. Which tool to call when several look similar, which requires the names and descriptions to differentiate. And what values to pass, which requires the parameter definitions to be unambiguous.
A schema can fail at each of these independently. A tool with a perfect input schema and a description reading “Gets deals” fails the first two: the description restates the function name and tells the agent nothing about scope or when the tool applies. A tool with an excellent description and a status parameter typed as a plain string fails the third: the agent knows it should call the tool and has to invent a value.
The design principle the guidance converges on. Design for agent experience, not API completeness. The instinct when building an MCP server on top of an existing API is to expose every endpoint, because that is what completeness looks like. The result is a tool surface shaped by the API’s internal structure rather than by the tasks an agent is trying to accomplish, and not every REST endpoint should become an MCP tool.
The tool definition has several text fields and they have distinct jobs. Mixing them is a common and avoidable error.
The tool name is a primary signal for selection. Agents use names to decide which tool to call, which means a name should be specific and verb-led rather than generic. Parameter names matter too, and the guidance here is pointed: rename parameters to match how a model would understand the domain, not how the database labels its columns.
The tool description carries what the agent needs to choose correctly: purpose, scope, units, which enums matter, how to disambiguate identifiers, and any guardrails on use. It should also say when not to use the tool and what the alternative is, which is the field most often left empty and the one that most directly reduces wrong-tool selection.
Parameter descriptions carry the meaning of each individual parameter. The explicit guidance is not to put input descriptions in the tool description: each parameter should describe itself where it is defined. A parameter description reading “The ID” is the canonical bad example, because it says nothing about which ID, in what format, or where the agent would obtain one.
Usage examples carry what the schema cannot express: when optional parameters should be included, how nested structures should actually look when populated, and what a realistic call looks like end to end. This is where the 72 to 90 percent accuracy gain was measured.
The single cheapest reliability improvement available is replacing permissive types with constrained ones.
The failure this fixes. A parameter declared as {"type": "string"} for a field that accepts four values gives the agent no information about which four. It will produce something plausible, and plausible is frequently wrong: "complete" where the API expects "completed", "urgent" where it expects "high".
What an enum does. When an agent sees "enum": ["pending", "shipped", "delivered", "cancelled"], the valid values are in its context at the moment it constructs the call. There is nothing to infer.
Where to apply it. Any field with a finite value set. Any field where a wrong value has consequences rather than just producing an error. Status fields, category fields, priority levels, environment names, and role identifiers are the usual candidates.
The pattern worth auditing for. Teams frequently define finite values as plain strings and validate them at runtime, because the validation already exists in the underlying service. The validation catches the error, so nothing appears broken. What it costs is a failed call, a returned error, and a retry, which for an agentic workflow is both latency and an opportunity to go off course. Promoting the runtime check into a schema enum moves the constraint to where the agent can see it.
Bounds belong in the schema too. minLength, maxLength, numeric minimums and maximums, and format constraints all narrow what the agent can produce. For safety-sensitive parameters this is a control rather than a convenience: a bound the schema enforces is a bound the agent cannot exceed regardless of what it decides to attempt.
Nesting is where schema structure most directly costs accuracy.
The guidance is consistent across sources. Complex nested objects confuse models and increase hallucination risk. Avoid nested dictionaries. Strive for the simplest schema that conveys the necessary information, and prefer a flatter design when it achieves the same purpose.
The practical translation. When mapping a REST endpoint whose request body is a simple object, flatten those fields into top-level tool parameters rather than preserving the nesting. An API expecting {"user": {"name": ..., "email": ..., "role": ...}} becomes a tool with name, email, and role as top-level parameters, with the server reassembling the structure before calling the API.
Where nesting is unavoidable. Genuinely repeating structures, such as an order with a variable number of line items each having a product identifier and a quantity. An array of objects is the correct representation and flattening it is not possible. When nesting is necessary, usage examples become more important rather than less, because showing a populated structure communicates what the schema describes abstractly.
The reassembly principle. The shape the agent sees and the shape the downstream API requires do not have to match. The MCP server sits between them and can translate. Designing the tool schema around what makes the agent reliable, then adapting to the API’s shape inside the handler, is the whole point of the server existing.
Two published targets bracket the recommended range: Google’s MCP Toolbox style guide suggests fewer than five parameters per tool, and AWS Prescriptive Guidance suggests around eight or fewer.
Why fewer is better. Each parameter is a decision the agent must make correctly. A tool with twelve parameters, eight of them optional and three with conditional requirements, asks the model to get twelve things right simultaneously, and the probability of a fully correct call falls with each one.
The tradeoff that is easy to get backwards. The obvious response to a complex tool is splitting it into several simpler ones. This is frequently correct and it is also where teams overcorrect, because tool-set size has its own accuracy cost.
The guidance here is specific: use parameters to handle variations rather than creating a tool per variation. Three tools named get_active_orders, get_pending_orders, and get_shipped_orders should be one get_orders tool with a status enum parameter. The enum communicates the same options with one entry in the tool list instead of three.
The distinction that resolves it. Split a tool when the variants do genuinely different things, particularly when one is safe and another is destructive. Combine with an enum parameter when the variants are the same operation over different values. A tool that sends a customer-visible message and a tool that saves a draft should be separate tools, because a wrong enum value in a combined tool sends a real message to a real customer. A tool that filters orders by status should be one tool.
This is the tension that makes schema design a tradeoff rather than a matter of writing more description.
Every token is paid on every request. Tool definitions sit in the model’s context for the entire conversation. A richly described tool with detailed parameter descriptions and worked examples consumes that context whether the tool is called once, many times, or never.
Context rot. Excessive or irrelevant tool definitions dilute the model’s attention with distractor tokens, and reasoning accuracy degrades. This is not a gradual linear decline: the published characterisation is that accuracy collapses once the tool surface grows past what the model can hold coherently.
The threshold that Anthropic acted on. Claude Code defers MCP tool descriptions when they exceed 10 percent of context, discovering them on demand through a search mechanism instead, with this behaviour enabled by default since version 2.1.7. That 10 percent figure is a useful planning number: a tool surface consuming more than a tenth of the context window is large enough that the platform itself treats it as a problem.
The measured progression from AWS’s tool design walkthrough. Moving from a bare schema to rich descriptions raised accuracy and made the definition noticeably larger. Moving from there to schema constraints with defaults raised accuracy further while shrinking the definition, because enums and defaults communicate the same information more compactly than prose does.
The ordering matters: prefer encoding information in the schema structure over encoding it in prose, because the schema is both more reliable for the agent and cheaper in context.
The retry calculus. Per-call context overhead is higher with a richer definition, but fewer failed calls and retries often make the total lower. A tool that is called correctly on the first attempt at the cost of 200 extra definition tokens is cheaper than one that is called wrongly, returns an error, and is retried.
The resolution to the tension between completeness and context cost is moving detail behind a second tool.
The pattern. Rather than embedding a large enum and extensive descriptions in a frequently loaded tool, the tool keeps short hint descriptions and a separate tool supplies the full detail when needed. A search tool might describe its subject_area parameter with a brief example rather than enumerating fifty valid values, with a get_taxonomy tool returning the full list.
Why it works. The agent pays the small definition cost on every request and the full enum cost only when it needs the enum. For straightforward queries where the hint is sufficient, it skips the taxonomy call entirely and searches directly.
When it is worth the extra round trip. When the detail is large, when it changes frequently enough that embedding it means redeploying to update it, and when a meaningful share of calls do not need it. When the enum has five values and never changes, embedding it is simpler and cheaper.
The platform-level version. Anthropic’s Tool Search Tool generalises this to the whole tool surface: rather than loading every tool definition into context, the agent discovers tools on demand. Programmatic tool calling, which lets the model orchestrate multiple tool calls in code rather than through individual round trips, was measured at roughly 37 percent token reduction for multi-step agent loops. Both became generally available in February 2026.
A tool whose work takes minutes rather than milliseconds needs the schema to say so, because the agent has no other way to know.
Two schema-level requirements. State explicitly in the tool description that the operation is long-running, and specify the polling workflow the agent should follow. An agent that does not know an operation is asynchronous will either wait for a response that is not coming or assume failure and retry, which for a generation operation means paying twice.
Provide a dedicated status tool. A get_operation or equivalent that takes the job identifier and returns a state the agent can poll until terminal. The initiating tool returns the identifier immediately; the status tool reports progress.
Make the terminal states explicit. The status tool’s output schema should enumerate the states, so the agent knows which values mean keep polling and which mean stop. A status field typed as a plain string leaves the agent inferring whether "processing" and "in_progress" are the same thing.
This pattern shows up directly in GMI Cloud’s MCP server, where video and image generation are inherently long-running and the tool surface is built around initiating work, polling for completion, and retrieving results as distinct steps rather than one blocking call.
Input schemas are mandatory in MCP. Output schemas are optional, and defining them is worth the effort for two reasons.
They tell the agent how to interpret what came back. A tool returning an untyped object leaves the agent inferring field meanings from names. An output schema that names every field with its meaning removes that inference, which matters most when the result feeds a subsequent tool call and a misread field propagates.
They enable client-side validation. Clients should validate results against the declared schema, which catches the case where the tool returns something structurally different from what it promised. Without a declared output schema there is nothing to validate against, and a malformed result reaches the agent’s reasoning as if it were correct.
The consistency obligation. Declaring an output schema is a commitment that the tool returns data conforming to it. A schema that is aspirational rather than enforced is worse than none, because the client validates against a contract the server does not honour.
A note on the standard. MCP defaults to JSON Schema 2020-12 when no $schema field is present, which is worth knowing when a schema behaves differently than expected across draft versions.
Schema quality is measurable, and the measurement is cheap relative to discovering problems in production.
Build a task set, not a parameter set. Twenty to thirty natural-language requests representing what users will actually ask, including ambiguous phrasings and requests near the boundary of what the tool covers. The question is whether the agent selects the right tool and constructs a valid call, not whether the handler works.
Measure three things separately. Tool selection accuracy: did the agent pick the right tool? Parameter validity: was the call structurally valid and within constraints? Parameter correctness: were the values right for the request, which is distinct from valid?
Run each task multiple times. Agent behaviour is non-deterministic, and a schema that works on one attempt may fail on the next for reasons a single run does not surface.
Test with the full tool surface loaded. A tool evaluated alone will show selection accuracy that does not survive contact with fifteen sibling tools, because selection accuracy depends on how many similar options are visible at once. Tool-set size is part of the same design problem, not a separate one.
Log every call in production. Every tool call an agent makes, with its parameters and outcome, is evidence about the schema. A parameter that is frequently wrong points at a description that needs work. A tool that is frequently called when a different one was appropriate points at descriptions that do not differentiate.
Three properties of the GMI Cloud MCP tool surface reflect the principles above.
A small tool set. Ten tools in v1 rather than an exhaustive mapping of the platform API. This is the selection-accuracy constraint applied at design time: a surface the agent can hold coherently produces better selection than one that covers every capability.
Non-destructive by construction. No deletes, no account changes, no irreversible operations. Schema constraints reduce the probability of a wrong call; absence of the capability removes it. For a tool surface exposed to autonomous agents, deciding what not to expose is the most reliable constraint available.
Estimation as a first-class tool. Cost estimates are a tool the agent can call rather than information buried in a description, which means the agent can quote before it spends. This is the same principle as enums: putting the information where the agent can act on it rather than where it has to infer it.
The tool surface is described in detail in GMI MCP Server: Turn your assistant into a production studio. For teams building agents that consume tools rather than exposing them, GMI Agentbox provides the deployment and runtime layer.
A tool schema is the agent’s complete view of a tool, and the measured effect of schema quality on call accuracy is large: 72 to 90 percent on complex parameter handling from adding worked examples alone.
The reliability levers, in rough order of effect per unit of effort: constrain finite values with enums so the agent selects rather than guesses, flatten nested structures wherever a flatter design works, keep parameters under roughly five to eight per tool, write descriptions that say when not to use the tool rather than restating its name, and add usage examples for anything a schema cannot express.
The counterweight is that every token in a definition is paid on every request, and excessive tool surface degrades reasoning accuracy through attention dilution. Prefer encoding constraints in schema structure over prose, because it is both more reliable and more compact, and move large or rarely needed detail behind a second tool when the context cost justifies the extra round trip.
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
