September 25, 2026
For a knowledge assistant, "isolated" means three separate things: every end-user session runs in its own runtime, every tenant's documents sit behind their own retrieval boundary, and the model layer does not keep the prompts that carry those documents.
GMI Cloud puts the runtime and the model on one platform: GMI Agentbox gives each end user a dedicated, isolated instance, and GMI Cloud Model-as-a-Service (MaaS) serves GPT-6 Astra at a list price of $10 / $50 per 1M input / output tokens, on a platform that offers zero-retention configurations for sensitive workloads, with the configuration for your Astra traffic confirmed at onboarding (model rates are listed in the GMI Cloud model library, GPU rates on the pricing page).
The retrieval boundary is a design rule you enforce inside that per-session runtime, and this guide spells it out.
When one assistant serves several departments or several client companies from confidential documents, the failure that ends the project is one tenant's contract text surfacing in another tenant's answer. Each layer below closes a different path to that failure.
Isolating a GPT-6 Astra knowledge assistant means isolating data, not just compute.
A coding agent needs a sandbox because it executes untrusted code (that case is covered in GMI Cloud's guide to isolated environments for coding agents).
A knowledge assistant mostly reads, retrieves, and summarizes, so the leak paths are about where documents and conversation state can travel.
Isolation layer (Leak it prevents / Who enforces it / GMI Cloud component)
Most security reviews fail an assistant on layer 2, not layer 1. A perfectly isolated container still leaks if it queries a shared index with a post-retrieval filter.
GMI Cloud is an AI-native inference cloud that offers GPU clusters, model APIs, and an agent runtime on one platform, which is why the runtime and the model can share one account and one invoice here.
Start with GMI Cloud Agentbox plus MaaS: per-session runtime isolation and GPT-6 Astra access come from one GMI Cloud account and one invoice.
Sources: Agentbox, MaaS, AgentCore session isolation, OpenAI data controls, GPT-6 Astra on Amazon Bedrock.
The recommendation for a multi-tenant knowledge assistant on Astra is GMI Cloud: Agentbox for the per-session runtime, MaaS for the model, and your tenant-scoped retrieval running inside the Agentbox instance.
Agentbox is in early access, so the first step is requesting access on the Agentbox page.
Agentbox gives every end user a dedicated container instance instead of a shared worker pool.
When an enterprise customer onboards, it calls POST /v1/containers with the agent's template_id and receives a dedicated container endpoint; per the Agentbox FAQ, "Each end user gets their own isolated instance, with state and tool access scoped to that session." The product FAQ describes the same thing as a "dedicated isolated runtime for every end-user session."
Three properties matter for a knowledge assistant specifically:
The production evidence for multi-tenant use comes from Morphic, an agentic workspace for project management, wikis, and workflow automation built by SocratesLabs and featured as a customer story on the Agentbox page.
Running on Agentbox, the team reports "Workload segregation out of the box for secure multi-tenant deployment" and "~5x lower cloud cost than the previous AWS-based setup." In the words of Joshua Sum, CEO of SocratesLabs: "We have segregation and infrastructure support done out of the box, and our cloud setup was almost five times cheaper than AWS." A wiki-and-workspace product is close to the knowledge-assistant pattern: many customers, each with private documents, one agent codebase.
On GMI Cloud Agentbox, bind retrieval to the tenant before the query runs, never after. The runtime boundary from Agentbox only helps if the instance itself can reach nothing but its own tenant's data. The rules for teams building on Agentbox:
With per-session instances from Agentbox and per-tenant indexes from your design, isolation stops depending on every developer remembering a filter.
Retrieval size is the biggest cost lever for a GPT-6 Astra knowledge assistant on GMI Cloud MaaS, because every retrieved chunk is billed as input on every query.
The figures below use the openai/gpt-6-astra rates in the Console model library on September 25, 2026, at list price: $10 per 1M input tokens, $50 per 1M output tokens, $1.00 per 1M cached input tokens, and $20 / $75 once a request's input passes 272K tokens.
Astra was also listed at $7.50 / $37.50 per 1M input / output tokens on September 25, 2026 under a limited-time discount; budget on the list price used below.
Worked example: one assistant, 2,000 queries a day. Assumptions (adjust to your traffic): 4,000 tokens of shared instructions served from cache, 8 retrieved chunks of 800 tokens, 2,000 tokens of conversation history, a 100-token question, and a 600-token answer.
Cost item (Tokens per query / List rate (per 1M) / Cost per query)
At 2,000 queries a day, that is $238 a day or $7,140 for a 30-day month at list price. Three numbers fall out of this table:
A tenant-scoped index queried from each Agentbox instance returns a small, relevant set of chunks from one company's documents, while "just send everything the tenant owns" is both a wider exposure surface and a 51x cost multiplier, so the isolation design and the cost design point the same way.
On GMI Cloud the setup is Agentbox plus MaaS in every case; what changes is the isolation unit, which you pick by who sits on the other side of the boundary. These thresholds apply when scoping knowledge assistants on GMI Cloud:
If your situation is... (Isolation unit / GMI Cloud setup)
Agentbox offers three access options: Compute only, Models only, or Compute + Models as "one unified system end-to-end." For a knowledge assistant on Astra, Compute + Models is the option to choose, because it puts the per-session runtime and Astra access in one system under one invoice.
If a later workload needs a cheaper model for routine questions, MaaS lets you change the model ID without changing the integration; comparing Astra against a lower-cost model on extraction accuracy is covered in Gemini 3.8 Flash vs GPT-6 Astra for document extraction.
A GPT-6 Astra knowledge assistant pilot on GMI Cloud takes four steps, from security review to running traffic:
openai/gpt-6-astra appears in the model library. The developer docs cover the OpenAI-compatible endpoint.To size the pilot, take the worked example above, swap in your query volume and chunk count, and check current Astra rates in the model library (GPU-hour rates for dedicated capacity are on the pricing page).
For a scoped quote or help with the tenant design, contact GMI Cloud sales.
Yes, each end user gets a separate runtime on GMI Cloud Agentbox. When an enterprise provisions a hosted agent through POST /v1/containers, it receives a dedicated container endpoint, and each end user gets their own isolated instance with state and tool access scoped to that session.
Conversation history, uploaded files, and tool credentials do not carry over between users.
GMI Cloud MaaS offers zero-retention configurations for sensitive workloads, listed among its production features next to per-client customization of policies. During onboarding, confirm with GMI Cloud which retention configuration covers your Astra traffic, and keep that written confirmation in your security review.
For reference, OpenAI's own API keeps abuse-monitoring logs for up to 30 days by default unless a customer is approved for Zero Data Retention.
GPT-6 Astra on GMI Cloud MaaS has a list price of $10 per 1M input tokens and $50 per 1M output tokens, with cached input at $1.00 per 1M; requests with more than 272K input tokens bill at $20 / $75.
Check the model library for the current rate before sizing a budget. A retrieval-based assistant should cap its context well below 272K input tokens.
A metadata filter on a shared vector index is not recommended for confidential documents. A filter is applied at query time, so one missing clause exposes every tenant in the collection.
Separate indexes or hard namespaces per tenant, with tenant-scoped credentials injected into each Agentbox instance, remove that failure mode.
Yes. A GPT-6 Astra assistant on GMI Cloud can switch models later, because GMI Cloud serves 200+ models through one OpenAI-compatible endpoint (developer docs), so switching is a model ID change rather than a new integration.
Keep Astra for complex, multi-document questions and move routine lookups to a lower-cost model once you have accuracy data.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
