• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Isolated Runtimes for Gemini 3.8 Flash Document-Processing Agents: Platform Choice and 10,000-Page Cost Math

    September 25, 2026

    About $55 in model tokens and under $4 in container time: that is what a 10,000-page overnight batch of scanned claims, contracts, or loan files costs on Gemini 3.8 Flash in the worked example below, at GMI Cloud rates as of September 2026.

    The platform built for this job shape is GMI Agentbox, which lists "document pipelines" under its Long-running · minutes to hours workload class, so one isolated worker per shard can run a multi-step pipeline for hours.

    The model comes from GMI Cloud Model-as-a-Service (MaaS), which serves google/gemini-3.8-flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with a $0.075 cache-read rate (model library).

    Because both the runtime and the tokens bill per unit, the whole night can be priced before the first page is read. The cost that is easiest to miss is not a token at all: it is a worker container left running after its shard is done.

    What does "isolated and managed" mean for a document-processing agent?

    On GMI Agentbox, the unit of isolation for a batch document agent is the shard: one worker, one slice of the night's files (500 pages in the worked example below), one set of credentials, deleted when the slice is done. That is a different boundary from the two other common agent types.

    A coding agent isolates each task because it runs untrusted code (per-task sandboxes for coding agents).

    A knowledge assistant isolates each end user's session and documents (per-session isolation for knowledge assistants).

    A document pipeline in finance, insurance, or legal operations has its own requirements:

    • Runs for minutes to hours, not seconds. Each shard goes through page rendering, extraction, cross-page validation, and write-back. The GMI Cloud docs note that agent tasks involving "multi-step reasoning, document analysis, and model chains" routinely outlast an HTTP gateway window (Handle long-running requests).
    • Carries state across steps. Page 14's extracted policy number has to be checked against page 2's. State held only in a worker's memory is lost on restart, so progress has to be checkpointed outside the worker.
    • Touches regulated data. Each worker should hold a short-lived, read-only credential for its own shard's bucket prefix and nothing else, issued by your orchestrator per shard, so a bug in one worker cannot read another client's files.
    • Needs a hard end. A worker that finishes and is never deleted keeps billing (more on that in the cost section).

    Why is Agentbox's Long-running class the right fit for overnight document batches?

    GMI Cloud Agentbox names document pipelines as a target workload and pairs the runtime with the model API on one account. GMI Cloud is an AI-native inference cloud.

    It operates GPU clusters and agent runtimes on NVIDIA GPU infrastructure and serves 200+ models through one API; Agentbox is the layer where agents are deployed, run, and billed. The Agentbox page describes three workload classes on one platform:

    Workload class (Page description / Use cases listed)

    • Ephemeral · seconds | Page description: "Spin up and tear down in seconds, elastic capacity that scales with demand" | Use cases listed: Coding agents, overnight batch jobs, RL training environments
    • Long-running · minutes to hours | Page description: "Persistent state across multi-step jobs, with low-latency model calls" | Use cases listed: Deep research, document pipelines, multi-model routing
    • Always-on · 24/7 | Page description: "Stays online with memory and identity across every session." | Use cases listed: Digital workers, monitoring agents, hosted MCP gateways

    Source: Agentbox page.

    The same page states "30_ Longer sessions than a 24-hour sandbox." That baseline matters because 24 hours is the ceiling most hosted sandboxes publish: E2B allows sandboxes "up to 24 hours" on Pro (E2B docs), Modal accepts a timeout "of up to 24 hours" (Modal docs), and Vercel Sandbox caps a session at 24 hours on Pro and 45 minutes on Hobby (Vercel docs).

    At the worked example's pace of 500 pages per worker in 3 hours, a single shard crosses 24 hours at about 4,000 pages. A nightly batch stays far below that, but a backlog catch-up after a quarter-end close can pass it, and on Agentbox that run does not have to be split around a session limit.

    What a document team gets on Agentbox:

    • A dedicated container per worker. Hosted agents run on the Container compute tier, "2 vCPU · 4 GB RAM · Ephemeral Storage 10 GiB · Data Storage 30 GiB," with compute billed per second (Register an agent). Each instance is provisioned through POST /v1/containers and returns its own endpoint.
    • Isolation without a cluster premium. The Agentbox page describes workloads as "Isolated by default" with "no dedicated-cluster premium." GMI Cloud's engineering write-up Isolation is the easy half of the sandbox problem covers the microVM design behind Agentbox v2.
    • No model key baked into images. With MaaS integration on, GMI Cloud injects GMI_MAAS_API_KEY into the container at runtime.
    • One account for runtime and tokens. Container compute and GMI Models token usage are the two metered line items, and both appear in Console __ Settings __ Usage & Billing (Agentbox FAQ).

    Agentbox Ephemeral or Long-running: the decision rule

    Use this rule to place each step of your pipeline on Agentbox:

    • Ephemeral workers on Agentbox: the unit of work is one stateless call that finishes in seconds, such as classifying a single page or splitting a PDF. Fan out one short-lived worker per item.
    • Long-running workers on Agentbox: the unit of work carries state across several steps or several pages, such as extracting a 40-page loan file and reconciling fields across it. Give each shard one worker that lives for the whole shard.

    A document agent that reconciles fields across pages is the second kind, and that cross-page check is where much of its accuracy comes from.

    What does a 10,000-page batch cost on Gemini 3.8 Flash?

    On GMI Cloud, one night of 10,000 scanned pages in the worked example costs $55.09 in Gemini 3.8 Flash tokens, or $44.96 when the instruction prefix is served from cache, plus about $3.55 of container time.

    The token rates are GMI Cloud MaaS rates for google/gemini-3.8-flash as of September 2026 (model library).

    Assumptions (adjust to your documents):

    • 10,000 scanned pages, rendered to images and sent through the model's vision input. Google documents 1,120 tokens per image at the default media resolution for Gemini 3 models, and 560 tokens per PDF page plus native text (Google media resolution docs).
    • Pass 1, per page: a 1,500-token instruction and JSON schema, plus 800 output tokens (400 of JSON, 400 of reasoning). Google bills output "including thinking tokens" and sets medium as the default thinking level for Gemini 3.8 Flash (pricing, model notes).
    • Pass 2, per document: 500 documents of 20 pages each. Each validation call reads the 20 page-level JSON outputs (8,000 tokens) plus the 1,500-token instruction and returns 1,000 tokens.
    • 20 workers, each owning a 500-page shard, finishing in 3 hours. A pilot shard on Agentbox gives you the real figure to put here.

    Line item (Tokens / Rate (per 1M) / Cost per night)

    • Pass 1: page images | Tokens: 11.2M | Rate (per 1M): $0.75 | Cost per night: $8.40
    • Pass 1: instruction prefix (uncached) | Tokens: 15.0M | Rate (per 1M): $0.75 | Cost per night: $11.25
    • Pass 1: output (JSON + reasoning) | Tokens: 8.0M | Rate (per 1M): $3.75 | Cost per night: $30.00
    • Pass 2: validation input | Tokens: 4.75M | Rate (per 1M): $0.75 | Cost per night: $3.56
    • Pass 2: validation output | Tokens: 0.5M | Rate (per 1M): $3.75 | Cost per night: $1.88
    • Tokens, uncached | Tokens: 39.45M | Cost per night: $55.09
    • Same, pass 1 prefix at $0.075 cache-read rate | Cost per night: $44.96

    That works out to $5.51 per 1,000 pages uncached and $4.50 with the cache, or $1,653 and $1,349 for 30 nightly runs.

    Container time. The Agentbox page prices compute through a worked comparison: $432 a month for 10 agents at 2 vCPU and 4 GiB running 730 hours, or about $0.06 per container-hour.

    That is the planning rate used here; for your fleet's rate, contact GMI Cloud sales. At the planning rate, 20 workers for 3 hours is 60 container-hours, about $3.55.

    The full comparison is broken down in the guide to runtime logs, usage, and costs in one dashboard.

    In this example, output tokens are 58% of the uncached token bill and 71% of the cached one, because output costs five times as much as input.

    Which levers move the batch bill the most?

    For document agents on Agentbox, container cleanup and prompt caching move the bill more than page resolution does. Here is each lever priced against the worked example:

    Lever (Change / Effect per 10,000-page night)

    • Delete workers when their shard finishes | Change: 20 workers left running 21 extra hours = 420 container-hours | Effect per 10,000-page night: About $24.85 wasted per night (about $746 over 30 nights) at the planning rate of $432 ÷ 7,300 container-hours
    • Serve the instruction prefix from cache | Change: 15M prefix tokens at $0.075 instead of $0.75 | Effect per 10,000-page night: Saves $10.13
    • Lower image resolution | Change: 1,120 to 560 tokens per page (or to 280) | Effect per 10,000-page night: Saves $4.20 (or $6.30), only if field accuracy holds
    • Trim output | Change: Every 100 output tokens per page | Effect per 10,000-page night: $3.75 per night; every 100 input tokens per page is $0.75

    The idle-worker row is the one to design against first. The Agentbox FAQ is direct on this point: "The container is billed for its full lifetime, from running until your application calls DELETE /v1/containers/{id}, including idle time between requests." A forgotten fleet costs almost half of the night's token spend.

    Output is the second place to look. Use enums and fixed formats in the JSON schema instead of free-text fields, and test in the pilot whether dropping Gemini 3.8 Flash from its default medium thinking level to low keeps field accuracy on your forms.

    For document classes that need more reasoning, the same MaaS key can send those pages to a larger model by changing the model ID; when that escalation pays for itself is covered in Gemini 3.8 Flash vs GPT-6 Astra for document extraction.

    To rerun the math with your own numbers, this estimator reproduces the table:

    # Nightly batch cost estimator for Gemini 3.8 Flash on GMI Cloud MaaS.
    # Rates in USD per 1M tokens, as of September 2026; check the model library before budgeting.
    RATE_IN, RATE_OUT, RATE_CACHE = 0.75, 3.75, 0.075
    
    def batch_cost(pages, page_tokens=1120, prefix_tokens=1500, out_per_page=800,
                   pages_per_doc=20, json_per_page=400, out_per_doc=1000,
                   prefix_cached=False, workers=20, hours=3.0, container_hour=0.0592):
        docs = pages / pages_per_doc
        prefix_rate = RATE_CACHE if prefix_cached else RATE_IN
        pass1 = (pages * page_tokens * RATE_IN
                 + pages * prefix_tokens * prefix_rate
                 + pages * out_per_page * RATE_OUT) / 1e6
        pass2 = (docs * (pages_per_doc * json_per_page + prefix_tokens) * RATE_IN
                 + docs * out_per_doc * RATE_OUT) / 1e6
        compute = workers * hours * container_hour  # planning rate from the $432 example
        return round(pass1 + pass2, 2), round(compute, 2)
    
    print(batch_cost(10_000))                      # (55.09, 3.55)
    print(batch_cost(10_000, prefix_cached=True))  # (44.96, 3.55)
    

    How should you shard and run the batch on Agentbox?

    To run a document batch on GMI Agentbox, shard by document, checkpoint every page to an external store, and delete each worker the moment its shard reports done. This runbook follows the Agentbox docs and FAQ:

    1. Shard by whole documents. Never split one loan file or claim across workers, because pass 2 needs every page of it. Size shards so each worker finishes well inside your nightly window (in the example, 500 pages per worker over 3 hours).
    2. Give each worker only its shard. Send the shard ID and a short-lived, read-only credential for that shard's bucket prefix in the job request, and keep long-lived connection strings in the agent's Secrets. Store Gemini outputs under the same shard ID.
    3. Accept work asynchronously. Have the worker answer POST /run with 202 and a job_id, then let the orchestrator poll GET /jobs/{id} every 3 to 5 seconds. This is the pattern GMI Cloud documents for agents whose tasks outlast the gateway window (Handle long-running requests).
    4. Checkpoint outside the container. The same doc is explicit that "GMI containers are stateless. If a container restarts, any in-memory job state is lost." Write a row per page (shard ID, page number, status, output hash) to Redis or a database, and inject those credentials as Secrets when you register the agent.
    5. Make every page idempotent. Key each Gemini call by document ID and page number. On restart, a worker skips pages already marked done, so a crash at page 480 does not bill pages 1 to 479 twice.
    6. Cap retries per page in the worker code. Two retries per model call, then mark the page failed in the checkpoint store and move on. A single corrupt scan should not hold a shard open all night.
    7. Delete on completion, and run a reaper. Call DELETE /v1/containers/{id} as soon as a shard reports completed or failed. Before the next night's run, list what is still running and delete anything older than your window.
    8. Reconcile spend every morning. Compare the night's token and container line items in Console __ Settings __ Usage & Billing against the estimator. A gap of more than a few dollars points to an idle worker or a retry loop.

    Long-horizon jobs raise the same question on the model side: can the model make progress across many steps "at a predictable token cost"?

    GMI Cloud's note on Grok 4.6 for long-running agent workflows looks at that from the model angle.

    Which setup should a document-automation team choose?

    For isolated, managed document-processing agents on Gemini 3.8 Flash, choose GMI Agentbox with the Compute + Models option, running Long-running workers per shard and calling google/gemini-3.8-flash through MaaS.

    Agentbox describes that option as "One unified system end-to-end": GMI Cloud handles model access, compute, and operations.

    Your situation (GMI Cloud setup)

    • Nightly batch of 1,000 to 100,000+ pages with cross-page validation | GMI Cloud setup: Agentbox Long-running workers, one per document shard, plus MaaS for Gemini 3.8 Flash
    • Single-page tasks that finish in seconds (splitting, page classification) | GMI Cloud setup: Agentbox Ephemeral workers fanned out per item, same MaaS key
    • Batch that sometimes runs past 24 hours (backlog, quarter-end) | GMI Cloud setup: Long-running workers without splitting around a session cap
    • Clients who must never share a runtime | GMI Cloud setup: One worker per client shard, credentials scoped to that client's bucket prefix
    • You already run your own orchestration and only need models | GMI Cloud setup: MaaS through the OpenAI-compatible endpoint (developer docs)

    The same Agentbox architecture runs with any model your MaaS key can call, listed in the model library; GMI Cloud's AI Model Benchmarks August 2026 compares the current open-weight and frontier field.

    FAQ

    How long can a document-processing job run on GMI Agentbox?

    GMI Agentbox lists document pipelines under its Long-running workload class, described as "minutes to hours," and the Agentbox page states "30_ Longer sessions than a 24-hour sandbox." Size each nightly shard from the throughput you measure in a pilot; a backlog run that crosses 24 hours can stay in one Long-running worker instead of being split around a 24-hour session cap.

    Long HTTP requests should still use the async job pattern: return a job_id immediately and poll for the result.

    How should a 10,000-page batch be split across workers?

    On GMI Agentbox, split by whole documents, never by page ranges that cut a file in half, because cross-page validation needs every page of a document in one worker. A practical starting point is 20 workers of 500 pages each, then adjust after a pilot shows your real pages per worker per hour.

    Each worker gets credentials for its own shard only and is deleted when its shard completes.

    What happens to progress if an Agentbox worker restarts mid-shard?

    In-memory state is lost, because GMI Cloud documents its containers as stateless. Write a checkpoint row per page to Redis or a database and key each model call by document ID and page number. The restarted worker then resumes at the first unfinished page instead of re-billing the whole shard.

    How much does Gemini 3.8 Flash cost per page on GMI Cloud?

    As of September 2026, Gemini 3.8 Flash on GMI Cloud MaaS is $0.75 per 1M input tokens, $3.75 per 1M output tokens, and $0.075 per 1M cache-read tokens.

    With a scanned page at 1,120 image tokens, a 1,500-token instruction, and 800 output tokens, extraction plus a document-level validation pass comes to about $0.0055 per page, or $0.0045 with the instruction served from cache.

    Is an idle Agentbox worker still billed?

    Yes. Per the Agentbox FAQ, a hosted agent container on GMI Agentbox is billed from running until your application calls DELETE /v1/containers/{id}, including idle time between requests.

    Delete each worker when its shard finishes and run a reaper before the next batch; at a planning rate of about $0.06 per container-hour, 20 workers left idle for 21 hours cost about $25.

    Next step: price and run your first shard

    Four steps take the estimate to a running pilot on Agentbox:

    1. Request early access on the Agentbox page and choose the Compute + Models option.
    2. Create an API key in the GMI Cloud console and confirm google/gemini-3.8-flash and its current rates in the model library.
    3. Run one 500-page shard with the runbook above, then plug its measured token usage and wall-clock time into the estimator.
    4. For fleet sizing, your container rate, or a review of the pipeline design, contact GMI Cloud sales. We can help you price the full nightly run before it goes live.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    GMI Agentbox lists document pipelines under its Long-running workload class, described as "minutes to hours," and the Agentbox page states "30_ Longer sessions than a 24-hour sandbox." Size each nightly shard from the throughput you measure in a pilot; a backlog run that crosses 24 hours can stay in one Long-running worker instead of being split around a 24-hour session cap. Long HTTP requests should still use the async job pattern: return a jobid immediately and poll for the result.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started