A production-focused guide to handling MCP operations that outlast client timeouts, using job handles, polling, idempotency, cancellation, and durable state to avoid duplicate work and lost results.
October 04, 2026

A tool call that generates a video takes minutes. A tool call in ChatGPT times out after roughly 60 seconds. That arithmetic is the whole problem, and the naive resolution, making the tool wait and hoping it finishes, fails in a specific and expensive way: the client gives up, the agent concludes the call failed, and it retries. The generation job that was running perfectly well is now running twice, and both will complete, and both will be billed. Long-running operations in MCP are not primarily a performance problem. They are a correctness problem, and the two mechanisms that solve it are a job handle returned immediately and an idempotency key that makes the retry harmless.
Client timeouts are shorter than many legitimate operations. ChatGPT tool calls time out after approximately 60 seconds, which is well inside the normal runtime of video generation, large batch processing, or a long simulation.
The protocol does not require idempotent task creation, but real networks do. A requestor that retries after a timeout should not spawn a duplicate background job, and the practical fix is accepting an idempotency key or hashing stable request inputs to dedupe to the same job.
GMI Cloud’s MCP server is built around this constraint, because image, video, and audio generation are inherently long-running and the tool surface separates initiating work from polling for completion and retrieving results.
Progress notifications are best-effort; polling is authoritative. Notifications are receiver-driven and may not arrive at all, and OpenAI does not send a progress token in tool calls, so a server that relies on notifications alone has clients that never learn the job finished.
Async promotion is the elegant variant. Rather than forcing the agent to choose between a synchronous and an asynchronous tool, the server runs the operation synchronously and promotes it to a background job when it exceeds a threshold, returning a job handle instead of a result.
Job state must survive a process restart. An in-memory job store works in a demo and loses every in-flight job on deployment, which for a paid generation operation means the user paid for work whose result is now unreachable.
Asynchronous job patterns are well established in API design. MCP adds three complications that change how the pattern must be built.
The caller is a language model, not application code. A client library handles a 202 response and a polling loop deterministically. An agent has to decide to poll, decide how long to wait between polls, and decide when to stop. All three decisions are made from text in the tool schema, which means the polling workflow has to be documented in the tool description rather than implemented in a client.
The retry behaviour is non-deterministic. Application code retries according to a policy you wrote. An agent retries because it inferred the call failed, which it may do after a timeout, after an ambiguous response, or because the user asked again. The retry is not governed by a policy you control, which is exactly why idempotency has to be enforced server-side.
The client timeout is not yours to set. A server cannot extend the host application’s tool timeout. ChatGPT’s roughly 60-second limit applies regardless of what the server would prefer, and other clients have their own. The server has to return within the shortest timeout it might encounter, which in practice means returning quickly and always.
The foundational pattern returns a tracking identifier immediately rather than blocking.
The shape. The initiating tool starts the work in the background and returns within a second with a job identifier and a status. The agent receives a response well inside any client timeout, tells the user the job has started, and polls separately.
A minimal initiating response carries the job identifier, the current status, and ideally an indication of expected duration so the agent can set a sensible polling interval rather than hammering the status tool every 200 milliseconds.
Three tools rather than one. The pattern needs an initiating tool, a status tool, and a result tool. Some implementations combine status and result, returning the output when the status is terminal, which works and makes the agent’s loop simpler at the cost of a response shape that varies by state.
The status values that matter. A conventional set covers pending, running, succeeded, failed, and cancelled. The important design property is that the terminal states are enumerated in the status tool’s output schema, so the agent knows which values mean stop polling rather than inferring it. A status typed as a plain string leaves the agent guessing whether processing and running are the same thing, which is the schema design problem applied to job lifecycle.
What the status response should carry beyond the state. A progress indicator where one is meaningful, the step currently executing, and a timestamp of the last update. The progress figure is what lets the agent report something useful to the user rather than repeating “still running” for four minutes.
A refinement worth considering when operation duration varies widely.
The problem with a fixed choice. Separate synchronous and asynchronous tools force the agent to predict how long the work will take before starting it. For a code execution tool, a one-line command and a two-hour simulation are the same tool call with different arguments, and the agent has no reliable basis for choosing.
The promotion pattern. The server runs the operation synchronously with a timeout, commonly 30 seconds. If it completes within that window, the result is returned directly and the interaction is a normal tool call. If it exceeds the threshold, the server promotes it to a background job and returns a job handle with a message indicating the promotion occurred.
The agent gets a fast result when the work is fast and a job handle when it is not, without having to predict which.
What the schema has to say. The tool description must state that the operation may return either a result or a job handle, and that a job handle means the agent should poll. Without that, an agent receiving a job identifier where it expected output will treat the response as an error.
Where it fits. Operations with high variance in duration: code execution, data processing over variable input sizes, queries against datasets of unknown scale. Operations with consistently long duration, such as video generation, are better served by always returning a handle, because the promotion branch never fires and the conditional response shape adds complexity for nothing.
MCP supports progress notifications, and the emerging MCP Tasks specification adds status notifications on lifecycle transitions. Neither replaces polling, and understanding why prevents a design that works in one client and silently fails in another.
The distinction that matters. Polling is requestor-driven and authoritative. Notifications are receiver-driven and best-effort. A client that receives a status notification can update its display immediately, and it must still poll until it observes a terminal state itself.
The practical gap today. MCP supports progress notifications, but OpenAI does not send a progress token in tool calls, which means a server relying on notifications to signal completion has ChatGPT clients that never learn the job finished. A design that works in one client and not another is worse than one that works adequately everywhere.
The design conclusion. Build the polling path first and completely. Treat notifications as an enhancement that improves the experience where supported, not as the completion mechanism. The status tool is the authoritative source of truth in every client.
What MCP Tasks adds. The specification formalises the pattern with tasks/get for polling and notifications/tasks/status for lifecycle transitions, with the notification payload being the full task object identical to what tasks/get returns. The framing worth noting is that a task is not simply one request that finishes later: it can become a workflow spanning multiple interactions, where the receiver may send additional requests or notifications while the task runs. Until client support is broad, custom polling tools remain the portable approach.
This is the failure that costs money, and the protocol does not prevent it for you.
How it happens. The agent calls the generating tool. The server starts the job. For any reason, the response does not reach the agent or the agent interprets it as a failure: a network blip, a client timeout, an ambiguous response. The agent retries. The server receives what looks like a new request and starts a second job.
Both jobs run to completion. Both are billed. The user sees one result and pays for two.
Why this is more likely with an agent than with application code. Retry behaviour is inferred rather than configured. An agent may retry after a timeout, after a response it did not parse cleanly, or because the user repeated the request. None of this is governed by a policy the server author wrote.
The mechanism. Accept an idempotency key on the initiating request. If a request arrives with a key that has been seen before, return the existing job rather than creating a new one. The second call receives the same job identifier, the agent polls the same job, and exactly one unit of work is performed and billed.
When the agent does not supply a key. Agents do not reliably generate and reuse idempotency keys, which means the server cannot depend on receiving one. The fallback is hashing the stable inputs of the request: the prompt, the model, the parameters that affect output. A request whose hash matches a recent job dedupes to that job.
The window matters. Deduping forever means a user who deliberately requests the same generation twice gets the first result back, which is wrong. A window of a few minutes covers the retry case without preventing legitimate repeats, and documenting the window in the tool description lets the agent reason about it.
What to return on a deduped request. The existing job identifier with its current status, not an error. An error invites another retry. Returning the job handle lets the agent proceed to polling, which is what it was going to do anyway.
Idempotency is a billing control. For a free read operation a duplicate is wasted compute. For a paid generation it is a charge the user did not intend. Treating it as a correctness detail rather than a billing control understates what is at stake, which is the reasoning behind GMI Cloud’s MCP server addressing duplicate submission explicitly alongside its cost estimation tooling.
An agent that can start a long job should be able to stop one, and the reasons are both cost and user experience.
When it is needed. The user changes their mind mid-generation. The agent determines the parameters were wrong and wants to restart with corrected ones. A supervising process detects the job has exceeded its budget.
The tool. A cancel_job taking the identifier and returning the resulting state. The server marks the job cancelled, signals the worker, and stops billing from that point where the underlying work supports partial cancellation.
The honesty requirement. Not all work can be cancelled cleanly. If a generation has already completed its expensive phase and is finalising, cancellation may not reduce cost at all. The tool should report what actually happened rather than always returning success, because an agent that believes cancellation prevented a charge will report that to the user.
Cancellation and idempotency interact. A cancelled job should not be returned to a subsequent deduped request as if it were still viable. A retry after a cancellation is a new request and should start new work, which means the dedupe lookup needs to exclude cancelled jobs.
The job store is where this pattern most commonly breaks in production.
The demo version. An in-memory dictionary mapping job identifiers to state. It works, it is three lines, and it loses every in-flight job when the process restarts.
Why that is worse than it sounds. A server deployment during a four-minute video generation means the job identifier the agent holds no longer resolves. The agent polls, receives a not-found response, and reports failure to the user. The underlying work may have completed on a worker that outlived the API process, and the result is now unreachable. The user paid for output they cannot retrieve.
The production requirement. A persistent store, with Redis and DynamoDB the usual choices. Job state must survive restarts, and in a horizontally scaled deployment it must be visible to every replica, since the status poll may reach a different instance than the one that created the job. The 2026-07-28 protocol revision made MCP remote servers stateless at the protocol level specifically so any replica can serve any request, which only works if job state lives outside the process.
Beyond the store itself. A job queue needs worker leases so a crashed worker’s job can be reclaimed rather than stalling forever, lease renewal during long tasks, persisted progress covering step, percentage, and last heartbeat, and dead-letter handling after a bounded retry count. The guidance worth repeating is not to treat the queue message as the job: the message is a pointer, the job state is the record.
What to observe. Queue lag, lease age, retry count, failure reason, and traces spanning the API and the workers. A job that silently stalls with no heartbeat is invisible without lease age monitoring, and it is the failure mode that produces a user waiting indefinitely for a result that is never coming.
The agent learns the entire workflow from the tool definitions, which puts specific obligations on the schema.
The initiating tool description must state that the operation is long-running, that it returns a job handle rather than a result, and what the polling workflow is. An agent without this will either wait for output that is not coming or treat the handle as a failure.
The status tool output schema must enumerate the terminal states so the agent knows when to stop polling, and should include progress so the agent can report something meaningful rather than repeating the same message.
Expected duration in the initiating response lets the agent choose a polling interval proportionate to the work. A job expected to take four minutes polled every two seconds generates 120 unnecessary calls.
The idempotency window, if one exists, should be documented so the agent can reason about whether a repeated request will produce new work.
These are instances of the general principle that a tool schema is the agent’s complete view of a tool, covered in depth in MCP Tool Schema Design. The long-running case makes the point sharply, because an undocumented asynchronous workflow does not degrade gracefully: the agent either hangs or reports a failure that did not occur.
Image, video, and audio generation are long-running by nature, which makes this pattern structural rather than an edge case in the GMI Cloud MCP server’s design.
Work is initiated and tracked separately from retrieval. Generation tools start the work and return promptly; the agent polls and retrieves when the job completes. This is why a video generation request does not hit the client timeout that a blocking implementation would.
Duplicate submission is handled rather than left to the client. Because generation is billed, a retry that spawns a second job is a charge the user did not intend, and the server treats this as a billing control rather than a networking detail.
Estimation precedes execution. Cost estimates are a separate tool the agent can call before committing to a generation, which is the same design instinct applied earlier in the sequence: give the agent the information to avoid the expensive mistake rather than only handling it afterwards.
Everything lands in console history and billing. Jobs are visible outside the agent conversation, which means a job whose handle is lost in a conversation is still retrievable through the console rather than being orphaned.
The full tool surface is described in GMI MCP Server: Turn your assistant into a production studio. For teams building agents that orchestrate long-running work rather than assistants that trigger it, GMI Agentbox provides the runtime and deployment layer.
The constraint is simple and the consequences are not. Client tool calls time out in roughly a minute, many legitimate operations take longer, and an agent that experiences a timeout retries. Without a job handle the call fails; without idempotency the retry duplicates the work and the charge.
The pattern that holds has four parts. Return a job handle immediately rather than blocking. Expose a status tool with enumerated terminal states so the agent knows when to stop polling. Dedupe retries through an idempotency key or a hash of stable inputs, returning the existing job rather than an error. And persist job state outside the process so a deployment does not orphan in-flight work.
Two design choices separate implementations that work from those that work in one client. Build the polling path as the authoritative mechanism and treat progress notifications as an enhancement, because notifications are best-effort and at least one major client does not send the progress token at all. And document the entire asynchronous workflow in the tool schema, because the agent has nothing else to learn it from.
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
