September 25, 2026
Run the comparison on GMI Cloud's Model-as-a-Service (MaaS) API.
One OpenAI-compatible endpoint and one API key serve both dated builds, deepseek-ai/DeepSeek-V4-Pro-0813 and Qwen/Qwen3.8-Max-0902, so the SDK, base URL, request body, and billing are identical and the comparison isolates the model instead of two different integrations.
A full 100-task evaluation round on your own repository costs roughly $9.50 at GMI Cloud catalog list prices as of September 2026, and the harness below fits in about 60 lines of Python.
If you are choosing the default model for a coding product, the answer should come from your own codebase, not from a leaderboard. The steps below get you there in an afternoon.
A model A/B is only valid when the model is the single variable that changes. When the two models sit behind different providers, at least five other things change with them:
Variable (Two separate provider APIs / One GMI Cloud MaaS endpoint)
Each of those differences can shift latency, error rates, or output length for reasons that have nothing to do with the model, which muddies exactly the gap you are trying to measure.
Aggregators close some of the gap: OpenRouter, for example, lists both models, but it can route one model slug to different upstream hosts, so a careful test there means pinning a single provider per model.
On GMI Cloud MaaS both models answer on the same base URL with the same key, so both arms of the test share one integration: the same client, key, and request format.
GMI Cloud MaaS gives you both models, current pricing, and zero-retention configurations for private code, on one account.
GMI Cloud is an AI-native inference cloud, and MaaS is its serverless API layer: "A Model-as-a-Service platform for LLM, image, video, and audio models, with unified APIs, discounted pricing, and enterprise-grade guarantees" (MaaS page).
The two builds, as listed in the GMI Cloud model catalog as of September 2026 (model library, MaaS):
(DeepSeek V4 Pro 0813 / Qwen3.8 Max 0902)
On September 25, 2026 the model library also listed DeepSeek V4 Pro 0813 at a 20% promotional rate ($1.056 input, $3.168 output, $0.0352 cache read per 1M). The cost math in this guide uses list prices, so the budget holds after the promotion ends.
For background on each release, see DeepSeek V4 Pro Steps Out of Preview: The 0813 Build Is Live and Qwen3.8-Max: 2.4T Parameters Available Now.
Four GMI Cloud MaaS capabilities matter most for this evaluation:
Build the task set from your own merged pull requests, freeze everything except the model ID, and grade by running your tests. Public scores tell you how a model does on someone else's repository; your product ships against yours.
test_patch file, and the test command. Apply only the test_patch at the parent commit once and confirm the command fails; drop any task that already passes without a fix.temperature: 0 and the same max_tokens for both models. Hosted models can still return slightly different outputs across identical calls at temperature 0, so budget for a repeat run if the first result is close. Set max_tokens high enough for reasoning output; Qwen3.8 Max 0902 runs hybrid thinking by default, so a tight cap penalizes it unfairly.-0813 and -0902 IDs. The catalog also lists undated DeepSeek-V4-Pro and Qwen3.8-Max entries at different prices, and mixing them corrupts both your quality and cost numbers.git apply, then apply the task's test_patch, then run the test command. A patch that fails to apply counts as a failure.prompt_tokens, completion_tokens, and wall-clock latency per call from the API response.The Python script below sends every task to both models through one GMI Cloud MaaS client, grades each patch with your tests, and writes a CSV.
Create an API key at console.gmicloud.ai, export it as GMI_API_KEY, and put your tasks in tasks.jsonl with the fields id, repo (the path to a local mirror of the repository, so the 200 per-run clones are fast local copies instead of network fetches), commit, prompt, test_patch (absolute path to the PR's test diff), and test_cmd.
import csv, json, os, random, re, subprocess, tempfile, time
from openai import OpenAI
client = OpenAI(base_url="https://api.gmi-serving.com/v1",
api_key=os.environ["GMI_API_KEY"])
# USD per 1M tokens, GMI Cloud MaaS catalog list prices as of September 2026.
# Re-check the console model page before each run.
MODELS = {
"deepseek-ai/DeepSeek-V4-Pro-0813": {"in": 1.32, "out": 3.96},
"Qwen/Qwen3.8-Max-0902": {"in": 1.65, "out": 4.951},
}
SYSTEM = ("You are fixing a bug in a Git repository. Reply with one unified "
"diff inside a single ```diff block and nothing else.")
PARAMS = {"temperature": 0, "max_tokens": 16000}
DIFF = re.compile(r"```diff\n(.*?)```", re.S)
def run_task(task, model):
work = tempfile.mkdtemp()
subprocess.run(["git", "clone", "-q", task["repo"], work], check=True)
subprocess.run(["git", "checkout", "-q", task["commit"]], cwd=work, check=True)
t0 = time.perf_counter()
resp = client.chat.completions.create(
model=model,
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": task["prompt"]}],
**PARAMS)
latency = time.perf_counter() - t0
match = DIFF.search(resp.choices[0].message.content or "")
passed = False
if match:
patch = os.path.join(work, "model.patch")
with open(patch, "w") as f:
f.write(match.group(1))
applied = subprocess.run(["git", "apply", patch], cwd=work).returncode == 0
if applied: # add the PR's regression tests on top of the model's fix
applied = subprocess.run(["git", "apply", task["test_patch"]],
cwd=work).returncode == 0
if applied:
try:
passed = subprocess.run(task["test_cmd"], shell=True, cwd=work,
timeout=900).returncode == 0
except subprocess.TimeoutExpired:
passed = False
u, p = resp.usage, MODELS[model]
cost = (u.prompt_tokens * p["in"] + u.completion_tokens * p["out"]) / 1e6
return {"task": task["id"], "model": model, "passed": passed,
"latency_s": round(latency, 2), "prompt_tokens": u.prompt_tokens,
"completion_tokens": u.completion_tokens, "cost_usd": round(cost, 5)}
tasks = [json.loads(line) for line in open("tasks.jsonl")]
rows = []
for task in tasks:
order = list(MODELS)
random.shuffle(order) # randomize which model runs first
for model in order:
rows.append(run_task(task, model))
with open("results.csv", "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=rows[0].keys())
writer.writeheader()
writer.writerows(rows)
The cost column prices every prompt token at the standard uncached input rate. Cache reads would lower the real bill, and billable cache writes (listed for Qwen3.8 Max 0902) would raise it, so reconcile the column against your console usage after the run.
Run the test commands inside a container or throwaway VM, since they execute model-written code.
A 100-task comparison of DeepSeek V4 Pro 0813 and Qwen3.8 Max 0902 on GMI Cloud MaaS costs about $9.50 for both models together, under a typical bug-fix prompt shape. The assumptions: each task sends 20,000 input tokens (issue text plus the touched files) and receives 4,000 output tokens (reasoning plus the diff).
Prices are GMI Cloud MaaS catalog list prices as of September 2026 (model library).
Scenario (100 tasks) (DeepSeek V4 Pro 0813 / Qwen3.8 Max 0902 / Both models)
How the first row is computed: DeepSeek is 20,000 _ $1.32/1M + 4,000 _ $3.96/1M = $0.0264 + $0.0158 = $0.0422. Qwen is 20,000 _ $1.65/1M + 4,000 _ $4.951/1M = $0.0330 + $0.0198 = $0.0528.
The cache row uses the cache read rates from the table above; the GMI Cloud catalog also lists a cache write price of $2.063 per 1M tokens for Qwen3.8 Max 0902, so check whether writes apply to your prompts before counting on that row.
A statistically meaningful comparison on your own code costs less than a single engineer-hour, so there is no reason to pick a default coding model from public benchmarks alone.
Decide on paired results and cost per solved task, not on raw pass rate. The two models ran the same 100 tasks, so the informative tasks are the ones where exactly one model passed.
Rule 1: count discordant tasks. Tasks both models solved or both failed carry no signal about which is better. On the remaining tasks, one model needs a clear majority before the difference is real.
An exact two-sided sign test at the 5% level gives these thresholds (computed from the binomial distribution; reproduce any row with SciPy's binomtest, where the listed win count is the smallest with p < 0.05):
Tasks where only one model passed (Wins needed by the leader)
If the leader falls short, the difference is not significant at this sample size. Treat the models as tied for the decision: add another 100 tasks or let cost decide.
Rule 2: compare cost per solved task. Divide each model's total cost by its number of passes. As a hypothetical: if DeepSeek V4 Pro 0813 solves 58 tasks for $4.22, that is $0.073 per fix; if Qwen3.8 Max 0902 solves 64 tasks for $5.28, that is $0.083 per fix.
When Rule 1 says tied, the lower cost per fix wins.
Rule 3: check p95 latency for interactive features. For inline edits and chat inside an IDE, compare p95 wall-clock time as well as quality. For background agents that open PRs, pass rate and cost per fix matter more.
This short script turns results.csv into those three numbers:
import csv, statistics
from collections import defaultdict
rows = list(csv.DictReader(open("results.csv")))
by_task = defaultdict(dict)
for r in rows:
by_task[r["task"]][r["model"]] = r["passed"] == "True"
a, b = "deepseek-ai/DeepSeek-V4-Pro-0813", "Qwen/Qwen3.8-Max-0902"
for m in (a, b):
rs = [r for r in rows if r["model"] == m]
solved = sum(r["passed"] == "True" for r in rs)
cost = sum(float(r["cost_usd"]) for r in rs)
lat = sorted(float(r["latency_s"]) for r in rs)
per_fix = f"${cost / solved:.3f}" if solved else "n/a"
print(f"{m}: {solved}/{len(rs)} solved, ${cost:.2f} total, "
f"{per_fix} per fix, "
f"p50 {statistics.median(lat):.1f}s, p95 {lat[int(0.95 * (len(lat) - 1))]:.1f}s")
only_a = sum(t[a] and not t[b] for t in by_task.values())
only_b = sum(t[b] and not t[a] for t in by_task.values())
print(f"Only DeepSeek solved: {only_a} | Only Qwen solved: {only_b}")
Ship the winner as the default through the same MaaS endpoint, since the model ID you tested is the model ID you deploy. Three follow-ups keep the decision current:
tasks.jsonl in the repo and re-run the harness when a new build of either model appears in the model library.How much does it cost to compare DeepSeek V4 Pro 0813 and Qwen3.8 Max 0902 on 100 coding tasks? On GMI Cloud MaaS, about $9.50 for one round at 20,000 input and 4,000 output tokens per task: $4.22 for DeepSeek V4 Pro 0813 and $5.28 for Qwen3.8 Max 0902, at catalog list prices as of September 2026.
Three repeats cost about $29. Current per-token rates are listed in the GMI Cloud model library.
How do I keep a coding model comparison fair?
Change only the model ID. Use one endpoint and one SDK client, the same prompt template, temperature: 0, the same max_tokens, and interleaved run order, then grade by applying the diff and running your tests.
GMI Cloud MaaS serves both models behind one OpenAI-compatible API, which removes the integration differences between two separate provider accounts.
Should I use the dated model IDs or the undated ones?
Use the dated IDs, deepseek-ai/DeepSeek-V4-Pro-0813 and Qwen/Qwen3.8-Max-0902, so the eval is reproducible.
The GMI Cloud catalog lists undated DeepSeek-V4-Pro and Qwen3.8-Max entries as separate models with their own prices, so mixing the two gives you results you cannot trace back to a single build.
How many tasks do I need before trusting the result? For a DeepSeek V4 Pro 0813 versus Qwen3.8 Max 0902 coding comparison, 100 paired tasks is a practical starting point, but the number that matters is how many tasks only one model solved.
With 20 such tasks, the leader needs 15 wins to clear a two-sided sign test at 5%; with 40, it needs 27. Below that, the difference is not significant, so decide on cost per fix.
Can I send proprietary code to the models during the evaluation? GMI Cloud MaaS offers zero-retention configurations for sensitive workloads, which is the setting to request before sending private repository content.
Keep test execution in a sandbox either way, because the tests run code the model wrote.
Create a MaaS API key in the GMI Cloud console, copy the two dated model IDs from the model library, and run the harness against 100 tasks from your own history.
For endpoint details, see the Developers page; for current per-token rates, see the model library, and for dedicated GPU-hour rates, see pricing; for zero-retention or volume terms, contact the GMI Cloud team.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
