• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Isolated, Managed Environments for GPT-6 Astra Knowledge Assistants: Which Platform to Choose

    September 25, 2026

    For a knowledge assistant, "isolated" means three separate things: every end-user session runs in its own runtime, every tenant's documents sit behind their own retrieval boundary, and the model layer does not keep the prompts that carry those documents.

    GMI Cloud puts the runtime and the model on one platform: GMI Agentbox gives each end user a dedicated, isolated instance, and GMI Cloud Model-as-a-Service (MaaS) serves GPT-6 Astra at a list price of $10 / $50 per 1M input / output tokens, on a platform that offers zero-retention configurations for sensitive workloads, with the configuration for your Astra traffic confirmed at onboarding (model rates are listed in the GMI Cloud model library, GPU rates on the pricing page).

    The retrieval boundary is a design rule you enforce inside that per-session runtime, and this guide spells it out.

    When one assistant serves several departments or several client companies from confidential documents, the failure that ends the project is one tenant's contract text surfacing in another tenant's answer. Each layer below closes a different path to that failure.

    What does "isolated" actually mean for a knowledge assistant?

    Isolating a GPT-6 Astra knowledge assistant means isolating data, not just compute.

    A coding agent needs a sandbox because it executes untrusted code (that case is covered in GMI Cloud's guide to isolated environments for coding agents).

    A knowledge assistant mostly reads, retrieves, and summarizes, so the leak paths are about where documents and conversation state can travel.

    Isolation layer (Leak it prevents / Who enforces it / GMI Cloud component)

    • 1. Session runtime | Leak it prevents: User B's conversation state, uploaded files, or tool credentials still sitting in memory when user A's request arrives | Who enforces it: The hosting platform | GMI Cloud component: Agentbox: "Each end user gets their own isolated instance, with state and tool access scoped to that session"
    • 2. Tenant retrieval | Leak it prevents: A query from Company A matching chunks indexed for Company B | Who enforces it: Your application design, running inside the isolated runtime | GMI Cloud component: Agentbox instance receives only that tenant's retrieval credentials at creation
    • 3. Model-side retention | Leak it prevents: Prompts containing confidential chunks being stored by the model layer | Who enforces it: The model provider's configuration | GMI Cloud component: MaaS: "Zero-retention configurations for sensitive workloads"; confirm in writing which configuration covers your Astra traffic

    Most security reviews fail an assistant on layer 2, not layer 1. A perfectly isolated container still leaks if it queries a shared index with a post-retrieval filter.

    GMI Cloud is an AI-native inference cloud that offers GPU clusters, model APIs, and an agent runtime on one platform, which is why the runtime and the model can share one account and one invoice here.

    Which platforms should you shortlist for a GPT-6 Astra knowledge assistant?

    Start with GMI Cloud Agentbox plus MaaS: per-session runtime isolation and GPT-6 Astra access come from one GMI Cloud account and one invoice.

    GMI Cloud Agentbox + MaaS

    • Session runtime isolation: Dedicated isolated instance per end user; "Isolated by default, no dedicated-cluster premium"
    • GPT-6 Astra access: Through MaaS on the same account as the runtime, openai/gpt-6-astra, $10 / $50 per 1M tokens list
    • Model-side retention control: MaaS offers zero-retention configurations for sensitive workloads and per-client customization of pricing, policies, and deployment; confirm the configuration for your Astra traffic at onboarding
    • Billing: Runtime and models on one invoice

    Amazon Bedrock AgentCore Runtime

    • Session runtime isolation: Dedicated microVM per user session, terminated and memory sanitized when the session ends; sessions up to 8 hours on microVMs
    • GPT-6 Astra access: Available on Amazon Bedrock since September 8, 2026
    • Model-side retention control: Governed by your AWS account and Bedrock settings
    • Billing: AWS bill

    OpenAI API + your own runtime

    • Session runtime isolation: You build it (containers, VMs, or a sandbox vendor)
    • GPT-6 Astra access: Direct, $10 / $50 per 1M tokens standard
    • Model-side retention control: Abuse-monitoring logs kept up to 30 days by default; Zero Data Retention requires prior approval
    • Billing: Separate bills for runtime and model

    Self-hosted on Kubernetes

    • Session runtime isolation: You build and patch namespaces, network policies, and pod lifecycle
    • GPT-6 Astra access: Through an external API
    • Model-side retention control: Whatever you negotiate with the model vendor
    • Billing: Your cloud bill plus model bill

    Sources: Agentbox, MaaS, AgentCore session isolation, OpenAI data controls, GPT-6 Astra on Amazon Bedrock.

    The recommendation for a multi-tenant knowledge assistant on Astra is GMI Cloud: Agentbox for the per-session runtime, MaaS for the model, and your tenant-scoped retrieval running inside the Agentbox instance.

    Agentbox is in early access, so the first step is requesting access on the Agentbox page.

    How does Agentbox keep each end-user session isolated?

    Agentbox gives every end user a dedicated container instance instead of a shared worker pool.

    When an enterprise customer onboards, it calls POST /v1/containers with the agent's template_id and receives a dedicated container endpoint; per the Agentbox FAQ, "Each end user gets their own isolated instance, with state and tool access scoped to that session." The product FAQ describes the same thing as a "dedicated isolated runtime for every end-user session."

    Three properties matter for a knowledge assistant specifically:

    • No shared memory between users. Conversation history, uploaded PDFs, and any per-user cache live inside that instance. There is no process-level state that a second user's request can hit.
    • One system for runtime and model access. With the Compute + Models option, Agentbox runs the instance and routes its GPT-6 Astra calls through GMI Cloud MaaS as "one unified system end-to-end," with one account and one invoice for both.
    • Isolation at the machine level. GMI Cloud's engineering write-up Isolation is the easy half of the sandbox problem explains that each Agentbox v2 workload runs in its own microVM with a guest kernel.

    The production evidence for multi-tenant use comes from Morphic, an agentic workspace for project management, wikis, and workflow automation built by SocratesLabs and featured as a customer story on the Agentbox page.

    Running on Agentbox, the team reports "Workload segregation out of the box for secure multi-tenant deployment" and "~5x lower cloud cost than the previous AWS-based setup." In the words of Joshua Sum, CEO of SocratesLabs: "We have segregation and infrastructure support done out of the box, and our cloud setup was almost five times cheaper than AWS." A wiki-and-workspace product is close to the knowledge-assistant pattern: many customers, each with private documents, one agent codebase.

    How do you keep one tenant's documents out of another tenant's answers?

    On GMI Cloud Agentbox, bind retrieval to the tenant before the query runs, never after. The runtime boundary from Agentbox only helps if the instance itself can reach nothing but its own tenant's data. The rules for teams building on Agentbox:

    1. One index or namespace per tenant. Separate vector collections (or at minimum hard namespaces) per client company or department. A metadata filter on a shared collection is a post-retrieval control; one missing filter clause and the leak is live.
    2. Inject tenant-scoped credentials at instance creation. Pass the tenant's retrieval token into the container when your application provisions it, so the instance physically cannot authenticate against another tenant's index.
    3. Scope tools the same way. If the assistant can open files in SharePoint, Google Drive, or a ticketing system, the connector credentials belong to that session's instance, which is what Agentbox's session-scoped tool access is for.
    4. No cross-tenant answer cache. Semantic caches keyed only on question text will return Company A's answer to Company B. Key caches by tenant ID, or keep them inside the instance.
    5. Log the tenant ID on every retrieval and tool call. GPT-6 Astra hides more of its reasoning than earlier models, so action logs matter more than traces; GMI Cloud's note Before you scale Astra traffic lists what to capture. Per-agent usage then shows up in each agent's Analytics tab; the full dashboard setup is covered in GMI Cloud's guide to runtime logs, usage, and costs in one dashboard.

    With per-session instances from Agentbox and per-tenant indexes from your design, isolation stops depending on every developer remembering a filter.

    What does retrieval context cost on GPT-6 Astra?

    Retrieval size is the biggest cost lever for a GPT-6 Astra knowledge assistant on GMI Cloud MaaS, because every retrieved chunk is billed as input on every query.

    The figures below use the openai/gpt-6-astra rates in the Console model library on September 25, 2026, at list price: $10 per 1M input tokens, $50 per 1M output tokens, $1.00 per 1M cached input tokens, and $20 / $75 once a request's input passes 272K tokens.

    Astra was also listed at $7.50 / $37.50 per 1M input / output tokens on September 25, 2026 under a limited-time discount; budget on the list price used below.

    Worked example: one assistant, 2,000 queries a day. Assumptions (adjust to your traffic): 4,000 tokens of shared instructions served from cache, 8 retrieved chunks of 800 tokens, 2,000 tokens of conversation history, a 100-token question, and a 600-token answer.

    Cost item (Tokens per query / List rate (per 1M) / Cost per query)

    • Shared instructions (cached) | Tokens per query: 4,000 | List rate (per 1M): $1.00 | Cost per query: $0.0040
    • Retrieved chunks | Tokens per query: 6,400 | List rate (per 1M): $10.00 | Cost per query: $0.0640
    • History + question | Tokens per query: 2,100 | List rate (per 1M): $10.00 | Cost per query: $0.0210
    • Answer | Tokens per query: 600 | List rate (per 1M): $50.00 | Cost per query: $0.0300
    • Total | Cost per query: $0.1190

    At 2,000 queries a day, that is $238 a day or $7,140 for a 30-day month at list price. Three numbers fall out of this table:

    • Each extra 800-token chunk costs $0.008 per query, or $480 a month at this volume. Going from 8 to 12 chunks lifts the monthly bill from $7,140 to $9,060.
    • Caching the shared instructions saves $2,160 a month here. Without the cache, those 4,000 tokens bill at the full input rate and the month comes to $9,300.
    • Stuffing a whole document set into the prompt crosses the 272K tier. A 300,000-token request costs $6.00 in input alone at the $20 per 1M list rate, plus about $0.045 for the answer: roughly $6.05 per query, about 51 times the retrieval design above.

    A tenant-scoped index queried from each Agentbox instance returns a small, relevant set of chunks from one company's documents, while "just send everything the tenant owns" is both a wider exposure surface and a 51x cost multiplier, so the isolation design and the cost design point the same way.

    Which setup should you pick for your deployment?

    On GMI Cloud the setup is Agentbox plus MaaS in every case; what changes is the isolation unit, which you pick by who sits on the other side of the boundary. These thresholds apply when scoping knowledge assistants on GMI Cloud:

    If your situation is... (Isolation unit / GMI Cloud setup)

    • External client companies share one assistant product | Isolation unit: One Agentbox instance per end-user session, one index per client company | GMI Cloud setup: Agentbox (Compute + Models option) + MaaS, with the retention configuration for Astra confirmed in writing
    • Internal departments with different clearance levels (HR, legal, finance) | Isolation unit: One Agentbox instance per session, one index per department | GMI Cloud setup: Same, with department-scoped retrieval tokens
    • One team, one document set, no clearance differences | Isolation unit: Per-session instance still recommended; a single index is acceptable | GMI Cloud setup: Agentbox + MaaS
    • Average retrieved context above ~50K tokens per query | Isolation unit: Revisit chunking before scaling; you are paying $0.50 or more per query in input alone at list price | GMI Cloud setup: Track GMI Models token usage in Console __ Settings __ Usage & Billing, and move routine questions to a lower-cost MaaS model
    • Any request path that can exceed 272K input tokens | Isolation unit: Cap context in code; the second tier doubles the input rate | GMI Cloud setup: Enforce a hard input token limit in the Agentbox instance before the MaaS request is sent

    Agentbox offers three access options: Compute only, Models only, or Compute + Models as "one unified system end-to-end." For a knowledge assistant on Astra, Compute + Models is the option to choose, because it puts the per-session runtime and Astra access in one system under one invoice.

    If a later workload needs a cheaper model for routine questions, MaaS lets you change the model ID without changing the integration; comparing Astra against a lower-cost model on extraction accuracy is covered in Gemini 3.8 Flash vs GPT-6 Astra for document extraction.

    How do you get started on GMI Cloud?

    A GPT-6 Astra knowledge assistant pilot on GMI Cloud takes four steps, from security review to running traffic:

    1. Request Agentbox early access on the Agentbox page and choose the Compute + Models option.
    2. Create a GMI Cloud API key in the Console and confirm openai/gpt-6-astra appears in the model library. The developer docs cover the OpenAI-compatible endpoint.
    3. Confirm model-side retention terms. MaaS offers zero-retention configurations for sensitive workloads; have GMI Cloud confirm in writing which configuration covers your Astra traffic, and attach it to your security review alongside the per-session isolation design.
    4. Load one tenant, then two. Run a red-team pass where a user from tenant B asks for tenant A's document titles. With per-tenant indexes and session-scoped credentials, the correct answer is "not found."

    To size the pilot, take the worked example above, swap in your query volume and chunk count, and check current Astra rates in the model library (GPU-hour rates for dedicated capacity are on the pricing page).

    For a scoped quote or help with the tenant design, contact GMI Cloud sales.

    FAQ

    Does each end user get a separate runtime on Agentbox?

    Yes, each end user gets a separate runtime on GMI Cloud Agentbox. When an enterprise provisions a hosted agent through POST /v1/containers, it receives a dedicated container endpoint, and each end user gets their own isolated instance with state and tool access scoped to that session.

    Conversation history, uploaded files, and tool credentials do not carry over between users.

    Does GMI Cloud keep the prompts my assistant sends to GPT-6 Astra?

    GMI Cloud MaaS offers zero-retention configurations for sensitive workloads, listed among its production features next to per-client customization of policies. During onboarding, confirm with GMI Cloud which retention configuration covers your Astra traffic, and keep that written confirmation in your security review.

    For reference, OpenAI's own API keeps abuse-monitoring logs for up to 30 days by default unless a customer is approved for Zero Data Retention.

    How much does GPT-6 Astra cost on GMI Cloud?

    GPT-6 Astra on GMI Cloud MaaS has a list price of $10 per 1M input tokens and $50 per 1M output tokens, with cached input at $1.00 per 1M; requests with more than 272K input tokens bill at $20 / $75.

    Check the model library for the current rate before sizing a budget. A retrieval-based assistant should cap its context well below 272K input tokens.

    Is a metadata filter on a shared vector index enough for tenant isolation?

    A metadata filter on a shared vector index is not recommended for confidential documents. A filter is applied at query time, so one missing clause exposes every tenant in the collection.

    Separate indexes or hard namespaces per tenant, with tenant-scoped credentials injected into each Agentbox instance, remove that failure mode.

    Can the same assistant switch from GPT-6 Astra to a cheaper model later?

    Yes. A GPT-6 Astra assistant on GMI Cloud can switch models later, because GMI Cloud serves 200+ models through one OpenAI-compatible endpoint (developer docs), so switching is a model ID change rather than a new integration.

    Keep Astra for complex, multi-document questions and move routine lookups to a lower-cost model once you have accuracy data.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Yes, each end user gets a separate runtime on GMI Cloud Agentbox. When an enterprise provisions a hosted agent through POST /v1/containers, it receives a dedicated container endpoint, and each end user gets their own isolated instance with state and tool access scoped to that session. Conversation history, uploaded files, and tool credentials do not carry over between users.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started