Kimi K3 is coming to GMI on Day 0. It's also the reason we're opening early access to something we've been building alongside it: the GMI Coding Plan. It is a flat monthly subscription that includes K3 and other open sourced coding models we host1.
K3 in GMI Coding Plan
Moonshot's K3 is a 2.8-trillion-parameter open-weight model with a 1M-token context window, the largest open-weight release to date. On coding it placed in the top three across six benchmarks, leading SWE Marathon and Program Bench, and trailing GPT-5.6 Sol on Terminal Bench 2.1 by half a point.
Per-token, K3 is $3.00 / 1M input and $15.00 / 1M output2, the highest-priced model on our hub. On the Coding Plan you pay a fixed monthly amount instead, and K3 draws from your credit pool like any other premium model.
How the Coding Plan works
A monthly pool of Premium Credits: Every premium model on GMI: K3, Kimi K2.7-Code, GLM-5.2, Qwen3.7 Max, DeepSeek-V4-Pro, and the rest, draws from that same pool at its own published rate.
3 models, near-unlimited. Every tier also includes near-unlimited use of our standard lane: DeepSeek-V4-Flash, GLM-4.7-Flash, GPT-OSS-120B, for the 80% of work that doesn't need a frontier model. Subject to a fair use policy: heavy legitimate coding is fine, resale and batch pipelines aren't.
One key, OpenAI-compatible. Works with Claude Code, Cline, Roo, Continue, Zed, Aider, or your own agent.
New models included on arrival. No new subscription when the next release lands.
Standard lane models (near-unlimited, fair use):
Model | Input / output per 1M |
|---|---|
DeepSeek-V4-Flash | $0.10 / $0.20 |
GPT-OSS-120B | $0.05 / $0.25 |
GLM-4.7-FlashGPT-OSS-120B | $0.07 / $0.40 |
Premium lane models (credit pool):
Model | Credits / 1M input | Credits / 1M output | Blended credits / 1M (80/20) |
Kimi K3 | 300 | 1,500 | 540 |
Qwen3.7 Max | 250 | 750 | 350 |
Kimi K2.7-Code-Highspeed | 190 | 800 | 312 |
Qwen3.6 Max Preview | 130 | 780 | 260 |
GLM-5.2-FP8 | 140 | 440 | 200 |
Qwen3 Coder 480B A35B | 90 | 450 | 162 |
Kimi K2.7-Code | 95 | 400 | 156 |
GLM-5.1 | 98 | 308 | 140 |
DeepSeek-V4-Pro | 113 | 226 | 136 |
GLM-5 (H200) | 60 | 192 | 86 |
Qwen3.6 Plus | 50 | 300 | 100 |
MiniMax-M2.5 (H200) | 30 | 120 | 48 |
DeepSeek-V3.2-Exp | 27 | 41 | 30 |
Plans
Lite | Standard | Pro | Max | |
Monthly | $9.99 | $19.99 | $59.99 | $199.99 |
Annual (2 months free) | $99.90 | $199.90 | $599.90 | $1,999.90 |
Premium lane Credits / month | 1,000 | 2,200 | 7,200 | 26,000 |
Standard lane models | Included | Included | Included | Included |
Throughput priority | — | — | Higher | Highest |
If K3 is your main model, look at Pro or Max. Lite and Standard are built for a mixed diet: standard lane for the routine work, premium credits for the hard problems. Running K3 for everything on an entry tier will burn through credits quickly.
Launch promo: 20% off your first month on Standard and Pro: $15.99 and $47.99.
Why early access
We'd want to ship to a few hundred developers who'll tell us what's broken than to everyone at once.
Access before general availability, in waves
Your feedback goes straight to the team building it
We'll tell you what's changing before it changes
In exchange, we want to know what you're actually building with, so tier sizing and routing defaults match real coding workloads instead of our assumptions.
→ Join Coding Plan early access
Questions? Find us on X or Discord.
Model included is based on real time availability and local regulations based on your country/territories
As of July 24, 2026
Build AI Without Limits
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

