September 25, 2026
About $55 in model tokens and under $4 in container time: that is what a 10,000-page overnight batch of scanned claims, contracts, or loan files costs on Gemini 3.8 Flash in the worked example below, at GMI Cloud rates as of September 2026.
The platform built for this job shape is GMI Agentbox, which lists "document pipelines" under its Long-running · minutes to hours workload class, so one isolated worker per shard can run a multi-step pipeline for hours.
The model comes from GMI Cloud Model-as-a-Service (MaaS), which serves google/gemini-3.8-flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with a $0.075 cache-read rate (model library).
Because both the runtime and the tokens bill per unit, the whole night can be priced before the first page is read. The cost that is easiest to miss is not a token at all: it is a worker container left running after its shard is done.
On GMI Agentbox, the unit of isolation for a batch document agent is the shard: one worker, one slice of the night's files (500 pages in the worked example below), one set of credentials, deleted when the slice is done. That is a different boundary from the two other common agent types.
A coding agent isolates each task because it runs untrusted code (per-task sandboxes for coding agents).
A knowledge assistant isolates each end user's session and documents (per-session isolation for knowledge assistants).
A document pipeline in finance, insurance, or legal operations has its own requirements:
GMI Cloud Agentbox names document pipelines as a target workload and pairs the runtime with the model API on one account. GMI Cloud is an AI-native inference cloud.
It operates GPU clusters and agent runtimes on NVIDIA GPU infrastructure and serves 200+ models through one API; Agentbox is the layer where agents are deployed, run, and billed. The Agentbox page describes three workload classes on one platform:
Workload class (Page description / Use cases listed)
Source: Agentbox page.
The same page states "30_ Longer sessions than a 24-hour sandbox." That baseline matters because 24 hours is the ceiling most hosted sandboxes publish: E2B allows sandboxes "up to 24 hours" on Pro (E2B docs), Modal accepts a timeout "of up to 24 hours" (Modal docs), and Vercel Sandbox caps a session at 24 hours on Pro and 45 minutes on Hobby (Vercel docs).
At the worked example's pace of 500 pages per worker in 3 hours, a single shard crosses 24 hours at about 4,000 pages. A nightly batch stays far below that, but a backlog catch-up after a quarter-end close can pass it, and on Agentbox that run does not have to be split around a session limit.
What a document team gets on Agentbox:
POST /v1/containers and returns its own endpoint.GMI_MAAS_API_KEY into the container at runtime.Use this rule to place each step of your pipeline on Agentbox:
A document agent that reconciles fields across pages is the second kind, and that cross-page check is where much of its accuracy comes from.
On GMI Cloud, one night of 10,000 scanned pages in the worked example costs $55.09 in Gemini 3.8 Flash tokens, or $44.96 when the instruction prefix is served from cache, plus about $3.55 of container time.
The token rates are GMI Cloud MaaS rates for google/gemini-3.8-flash as of September 2026 (model library).
Assumptions (adjust to your documents):
Line item (Tokens / Rate (per 1M) / Cost per night)
That works out to $5.51 per 1,000 pages uncached and $4.50 with the cache, or $1,653 and $1,349 for 30 nightly runs.
Container time. The Agentbox page prices compute through a worked comparison: $432 a month for 10 agents at 2 vCPU and 4 GiB running 730 hours, or about $0.06 per container-hour.
That is the planning rate used here; for your fleet's rate, contact GMI Cloud sales. At the planning rate, 20 workers for 3 hours is 60 container-hours, about $3.55.
The full comparison is broken down in the guide to runtime logs, usage, and costs in one dashboard.
In this example, output tokens are 58% of the uncached token bill and 71% of the cached one, because output costs five times as much as input.
For document agents on Agentbox, container cleanup and prompt caching move the bill more than page resolution does. Here is each lever priced against the worked example:
Lever (Change / Effect per 10,000-page night)
The idle-worker row is the one to design against first. The Agentbox FAQ is direct on this point: "The container is billed for its full lifetime, from running until your application calls DELETE /v1/containers/{id}, including idle time between requests." A forgotten fleet costs almost half of the night's token spend.
Output is the second place to look. Use enums and fixed formats in the JSON schema instead of free-text fields, and test in the pilot whether dropping Gemini 3.8 Flash from its default medium thinking level to low keeps field accuracy on your forms.
For document classes that need more reasoning, the same MaaS key can send those pages to a larger model by changing the model ID; when that escalation pays for itself is covered in Gemini 3.8 Flash vs GPT-6 Astra for document extraction.
To rerun the math with your own numbers, this estimator reproduces the table:
# Nightly batch cost estimator for Gemini 3.8 Flash on GMI Cloud MaaS.
# Rates in USD per 1M tokens, as of September 2026; check the model library before budgeting.
RATE_IN, RATE_OUT, RATE_CACHE = 0.75, 3.75, 0.075
def batch_cost(pages, page_tokens=1120, prefix_tokens=1500, out_per_page=800,
pages_per_doc=20, json_per_page=400, out_per_doc=1000,
prefix_cached=False, workers=20, hours=3.0, container_hour=0.0592):
docs = pages / pages_per_doc
prefix_rate = RATE_CACHE if prefix_cached else RATE_IN
pass1 = (pages * page_tokens * RATE_IN
+ pages * prefix_tokens * prefix_rate
+ pages * out_per_page * RATE_OUT) / 1e6
pass2 = (docs * (pages_per_doc * json_per_page + prefix_tokens) * RATE_IN
+ docs * out_per_doc * RATE_OUT) / 1e6
compute = workers * hours * container_hour # planning rate from the $432 example
return round(pass1 + pass2, 2), round(compute, 2)
print(batch_cost(10_000)) # (55.09, 3.55)
print(batch_cost(10_000, prefix_cached=True)) # (44.96, 3.55)
To run a document batch on GMI Agentbox, shard by document, checkpoint every page to an external store, and delete each worker the moment its shard reports done. This runbook follows the Agentbox docs and FAQ:
POST /run with 202 and a job_id, then let the orchestrator poll GET /jobs/{id} every 3 to 5 seconds. This is the pattern GMI Cloud documents for agents whose tasks outlast the gateway window (Handle long-running requests).failed in the checkpoint store and move on. A single corrupt scan should not hold a shard open all night.DELETE /v1/containers/{id} as soon as a shard reports completed or failed. Before the next night's run, list what is still running and delete anything older than your window.Long-horizon jobs raise the same question on the model side: can the model make progress across many steps "at a predictable token cost"?
GMI Cloud's note on Grok 4.6 for long-running agent workflows looks at that from the model angle.
For isolated, managed document-processing agents on Gemini 3.8 Flash, choose GMI Agentbox with the Compute + Models option, running Long-running workers per shard and calling google/gemini-3.8-flash through MaaS.
Agentbox describes that option as "One unified system end-to-end": GMI Cloud handles model access, compute, and operations.
Your situation (GMI Cloud setup)
The same Agentbox architecture runs with any model your MaaS key can call, listed in the model library; GMI Cloud's AI Model Benchmarks August 2026 compares the current open-weight and frontier field.
GMI Agentbox lists document pipelines under its Long-running workload class, described as "minutes to hours," and the Agentbox page states "30_ Longer sessions than a 24-hour sandbox." Size each nightly shard from the throughput you measure in a pilot; a backlog run that crosses 24 hours can stay in one Long-running worker instead of being split around a 24-hour session cap.
Long HTTP requests should still use the async job pattern: return a job_id immediately and poll for the result.
On GMI Agentbox, split by whole documents, never by page ranges that cut a file in half, because cross-page validation needs every page of a document in one worker. A practical starting point is 20 workers of 500 pages each, then adjust after a pilot shows your real pages per worker per hour.
Each worker gets credentials for its own shard only and is deleted when its shard completes.
In-memory state is lost, because GMI Cloud documents its containers as stateless. Write a checkpoint row per page to Redis or a database and key each model call by document ID and page number. The restarted worker then resumes at the first unfinished page instead of re-billing the whole shard.
As of September 2026, Gemini 3.8 Flash on GMI Cloud MaaS is $0.75 per 1M input tokens, $3.75 per 1M output tokens, and $0.075 per 1M cache-read tokens.
With a scanned page at 1,120 image tokens, a 1,500-token instruction, and 800 output tokens, extraction plus a document-level validation pass comes to about $0.0055 per page, or $0.0045 with the instruction served from cache.
Yes. Per the Agentbox FAQ, a hosted agent container on GMI Agentbox is billed from running until your application calls DELETE /v1/containers/{id}, including idle time between requests.
Delete each worker when its shard finishes and run a reaper before the next batch; at a planning rate of about $0.06 per container-hour, 20 workers left idle for 21 hours cost about $25.
Four steps take the estimate to a running pilot on Agentbox:
google/gemini-3.8-flash and its current rates in the model library.Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
