• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Which API Platform Should You Use to Compare DeepSeek V4 Pro 0813 and Qwen3.8 Max 0902 on Coding Tasks?

    September 25, 2026

    Run the comparison on GMI Cloud's Model-as-a-Service (MaaS) API.

    One OpenAI-compatible endpoint and one API key serve both dated builds, deepseek-ai/DeepSeek-V4-Pro-0813 and Qwen/Qwen3.8-Max-0902, so the SDK, base URL, request body, and billing are identical and the comparison isolates the model instead of two different integrations.

    A full 100-task evaluation round on your own repository costs roughly $9.50 at GMI Cloud catalog list prices as of September 2026, and the harness below fits in about 60 lines of Python.

    If you are choosing the default model for a coding product, the answer should come from your own codebase, not from a leaderboard. The steps below get you there in an afternoon.

    Why does the API platform decide whether a coding comparison is valid?

    A model A/B is only valid when the model is the single variable that changes. When the two models sit behind different providers, at least five other things change with them:

    Variable (Two separate provider APIs / One GMI Cloud MaaS endpoint)

    • SDK and request schema | Two separate provider APIs: Two clients, two sets of defaults | One GMI Cloud MaaS endpoint: One OpenAI SDK client, only model changes
    • Network path and region | Two separate provider APIs: Two base URLs, two latency profiles | One GMI Cloud MaaS endpoint: One base URL: https://api.gmi-serving.com/v1
    • Parameter handling | Two separate provider APIs: Each provider maps temperature, max_tokens, and stop sequences its own way | One GMI Cloud MaaS endpoint: Same request body sent to both models
    • Token accounting | Two separate provider APIs: Two usage formats, two invoices | One GMI Cloud MaaS endpoint: One usage block format, one invoice
    • Rate limits and retries | Two separate provider APIs: Two quota systems to tune around | One GMI Cloud MaaS endpoint: One key, one quota relationship

    Each of those differences can shift latency, error rates, or output length for reasons that have nothing to do with the model, which muddies exactly the gap you are trying to measure.

    Aggregators close some of the gap: OpenRouter, for example, lists both models, but it can route one model slug to different upstream hosts, so a careful test there means pinning a single provider per model.

    On GMI Cloud MaaS both models answer on the same base URL with the same key, so both arms of the test share one integration: the same client, key, and request format.

    What does GMI Cloud MaaS provide for this comparison?

    GMI Cloud MaaS gives you both models, current pricing, and zero-retention configurations for private code, on one account.

    GMI Cloud is an AI-native inference cloud, and MaaS is its serverless API layer: "A Model-as-a-Service platform for LLM, image, video, and audio models, with unified APIs, discounted pricing, and enterprise-grade guarantees" (MaaS page).

    The two builds, as listed in the GMI Cloud model catalog as of September 2026 (model library, MaaS):

    (DeepSeek V4 Pro 0813 / Qwen3.8 Max 0902)

    • Model ID | DeepSeek V4 Pro 0813: deepseek-ai/DeepSeek-V4-Pro-0813 | Qwen3.8 Max 0902: Qwen/Qwen3.8-Max-0902
    • Input, per 1M tokens (list) | DeepSeek V4 Pro 0813: $1.32 | Qwen3.8 Max 0902: $1.65
    • Output, per 1M tokens (list) | DeepSeek V4 Pro 0813: $3.96 | Qwen3.8 Max 0902: $4.951
    • Cache read, per 1M tokens (list) | DeepSeek V4 Pro 0813: $0.044 | Qwen3.8 Max 0902: $0.206
    • Cache write, per 1M tokens | DeepSeek V4 Pro 0813: Not listed | Qwen3.8 Max 0902: $2.063
    • Catalog note | DeepSeek V4 Pro 0813: GA release of DeepSeek V4 Pro | Qwen3.8 Max 0902: Hybrid thinking enabled by default

    On September 25, 2026 the model library also listed DeepSeek V4 Pro 0813 at a 20% promotional rate ($1.056 input, $3.168 output, $0.0352 cache read per 1M). The cost math in this guide uses list prices, so the budget holds after the promotion ends.

    For background on each release, see DeepSeek V4 Pro Steps Out of Preview: The 0813 Build Is Live and Qwen3.8-Max: 2.4T Parameters Available Now.

    Four GMI Cloud MaaS capabilities matter most for this evaluation:

    • One API design for every model. "All models are available through a single, consistent API design," and the Developers page describes the SDK as "Drop-in for the OpenAI SDK." Switching arms of the test is a one-string change.
    • Zero-retention configurations. MaaS offers "Zero-retention configurations for sensitive workloads," which is what you want when the prompts contain your proprietary repository.
    • Centralized billing. "Centralized billing with a single invoice across all models" means the cost column of your eval reconciles against one bill.
    • A proven evaluation pattern. Eigen AI "combines GMI Cloud MaaS and dedicated endpoints to support fast model access across production serving, benchmarking, and evaluation" (MaaS page).

    How do you set up a fair coding evaluation?

    Build the task set from your own merged pull requests, freeze everything except the model ID, and grade by running your tests. Public scores tell you how a model does on someone else's repository; your product ships against yours.

    1. Harvest 100 tasks from history. Take merged PRs that fixed a bug and added or changed a test. For each, store the parent commit, the issue text as the prompt, the PR's test changes as a separate test_patch file, and the test command. Apply only the test_patch at the parent commit once and confirm the command fails; drop any task that already passes without a fix.
    2. Freeze the prompt template. Same system prompt, same user template, same context-packing rule (for example, the issue text plus the files the original PR touched).
    3. Freeze the parameters. Use temperature: 0 and the same max_tokens for both models. Hosted models can still return slightly different outputs across identical calls at temperature 0, so budget for a repeat run if the first result is close. Set max_tokens high enough for reasoning output; Qwen3.8 Max 0902 runs hybrid thinking by default, so a tight cap penalizes it unfairly.
    4. Pin dated model IDs. Use the -0813 and -0902 IDs. The catalog also lists undated DeepSeek-V4-Pro and Qwen3.8-Max entries at different prices, and mixing them corrupts both your quality and cost numbers.
    5. Interleave the runs. Randomize which model goes first on each task so both see the same time-of-day conditions.
    6. Grade mechanically. Apply the returned diff with git apply, then apply the task's test_patch, then run the test command. A patch that fails to apply counts as a failure.
    7. Log usage, not estimates. Record prompt_tokens, completion_tokens, and wall-clock latency per call from the API response.

    What does the evaluation harness look like?

    The Python script below sends every task to both models through one GMI Cloud MaaS client, grades each patch with your tests, and writes a CSV.

    Create an API key at console.gmicloud.ai, export it as GMI_API_KEY, and put your tasks in tasks.jsonl with the fields id, repo (the path to a local mirror of the repository, so the 200 per-run clones are fast local copies instead of network fetches), commit, prompt, test_patch (absolute path to the PR's test diff), and test_cmd.

    import csv, json, os, random, re, subprocess, tempfile, time
    from openai import OpenAI
    
    client = OpenAI(base_url="https://api.gmi-serving.com/v1",
                    api_key=os.environ["GMI_API_KEY"])
    
    # USD per 1M tokens, GMI Cloud MaaS catalog list prices as of September 2026.
    # Re-check the console model page before each run.
    MODELS = {
        "deepseek-ai/DeepSeek-V4-Pro-0813": {"in": 1.32,  "out": 3.96},
        "Qwen/Qwen3.8-Max-0902":            {"in": 1.65,  "out": 4.951},
    }
    SYSTEM = ("You are fixing a bug in a Git repository. Reply with one unified "
              "diff inside a single ```diff block and nothing else.")
    PARAMS = {"temperature": 0, "max_tokens": 16000}
    DIFF = re.compile(r"```diff\n(.*?)```", re.S)
    
    def run_task(task, model):
        work = tempfile.mkdtemp()
        subprocess.run(["git", "clone", "-q", task["repo"], work], check=True)
        subprocess.run(["git", "checkout", "-q", task["commit"]], cwd=work, check=True)
        t0 = time.perf_counter()
        resp = client.chat.completions.create(
            model=model,
            messages=[{"role": "system", "content": SYSTEM},
                      {"role": "user", "content": task["prompt"]}],
            **PARAMS)
        latency = time.perf_counter() - t0
        match = DIFF.search(resp.choices[0].message.content or "")
        passed = False
        if match:
            patch = os.path.join(work, "model.patch")
            with open(patch, "w") as f:
                f.write(match.group(1))
            applied = subprocess.run(["git", "apply", patch], cwd=work).returncode == 0
            if applied:  # add the PR's regression tests on top of the model's fix
                applied = subprocess.run(["git", "apply", task["test_patch"]],
                                         cwd=work).returncode == 0
            if applied:
                try:
                    passed = subprocess.run(task["test_cmd"], shell=True, cwd=work,
                                            timeout=900).returncode == 0
                except subprocess.TimeoutExpired:
                    passed = False
        u, p = resp.usage, MODELS[model]
        cost = (u.prompt_tokens * p["in"] + u.completion_tokens * p["out"]) / 1e6
        return {"task": task["id"], "model": model, "passed": passed,
                "latency_s": round(latency, 2), "prompt_tokens": u.prompt_tokens,
                "completion_tokens": u.completion_tokens, "cost_usd": round(cost, 5)}
    
    tasks = [json.loads(line) for line in open("tasks.jsonl")]
    rows = []
    for task in tasks:
        order = list(MODELS)
        random.shuffle(order)          # randomize which model runs first
        for model in order:
            rows.append(run_task(task, model))
    
    with open("results.csv", "w", newline="") as f:
        writer = csv.DictWriter(f, fieldnames=rows[0].keys())
        writer.writeheader()
        writer.writerows(rows)
    

    The cost column prices every prompt token at the standard uncached input rate. Cache reads would lower the real bill, and billable cache writes (listed for Qwen3.8 Max 0902) would raise it, so reconcile the column against your console usage after the run.

    Run the test commands inside a container or throwaway VM, since they execute model-written code.

    How much does a 100-task evaluation round cost?

    A 100-task comparison of DeepSeek V4 Pro 0813 and Qwen3.8 Max 0902 on GMI Cloud MaaS costs about $9.50 for both models together, under a typical bug-fix prompt shape. The assumptions: each task sends 20,000 input tokens (issue text plus the touched files) and receives 4,000 output tokens (reasoning plus the diff).

    Prices are GMI Cloud MaaS catalog list prices as of September 2026 (model library).

    Scenario (100 tasks) (DeepSeek V4 Pro 0813 / Qwen3.8 Max 0902 / Both models)

    • Per task: 20K in, 4K out | DeepSeek V4 Pro 0813: $0.0422 | Qwen3.8 Max 0902: $0.0528 | Both models: $0.0950
    • One round, 100 tasks | DeepSeek V4 Pro 0813: $4.22 | Qwen3.8 Max 0902: $5.28 | Both models: $9.50
    • Three repeats to measure run-to-run variance | DeepSeek V4 Pro 0813: $12.67 | Qwen3.8 Max 0902: $15.84 | Both models: $28.51
    • Qwen output doubles to 8K (thinking-heavy) | DeepSeek V4 Pro 0813: $4.22 | Qwen3.8 Max 0902: $7.26 | Both models: $11.48
    • 15K of 20K input tokens served from cache | DeepSeek V4 Pro 0813: $2.31 | Qwen3.8 Max 0902: $3.11 | Both models: $5.42

    How the first row is computed: DeepSeek is 20,000 _ $1.32/1M + 4,000 _ $3.96/1M = $0.0264 + $0.0158 = $0.0422. Qwen is 20,000 _ $1.65/1M + 4,000 _ $4.951/1M = $0.0330 + $0.0198 = $0.0528.

    The cache row uses the cache read rates from the table above; the GMI Cloud catalog also lists a cache write price of $2.063 per 1M tokens for Qwen3.8 Max 0902, so check whether writes apply to your prompts before counting on that row.

    A statistically meaningful comparison on your own code costs less than a single engineer-hour, so there is no reason to pick a default coding model from public benchmarks alone.

    How do you read the results and choose a default model?

    Decide on paired results and cost per solved task, not on raw pass rate. The two models ran the same 100 tasks, so the informative tasks are the ones where exactly one model passed.

    Rule 1: count discordant tasks. Tasks both models solved or both failed carry no signal about which is better. On the remaining tasks, one model needs a clear majority before the difference is real.

    An exact two-sided sign test at the 5% level gives these thresholds (computed from the binomial distribution; reproduce any row with SciPy's binomtest, where the listed win count is the smallest with p < 0.05):

    Tasks where only one model passed (Wins needed by the leader)

    • 10 | Wins needed by the leader: 9
    • 20 | Wins needed by the leader: 15
    • 30 | Wins needed by the leader: 21
    • 40 | Wins needed by the leader: 27

    If the leader falls short, the difference is not significant at this sample size. Treat the models as tied for the decision: add another 100 tasks or let cost decide.

    Rule 2: compare cost per solved task. Divide each model's total cost by its number of passes. As a hypothetical: if DeepSeek V4 Pro 0813 solves 58 tasks for $4.22, that is $0.073 per fix; if Qwen3.8 Max 0902 solves 64 tasks for $5.28, that is $0.083 per fix.

    When Rule 1 says tied, the lower cost per fix wins.

    Rule 3: check p95 latency for interactive features. For inline edits and chat inside an IDE, compare p95 wall-clock time as well as quality. For background agents that open PRs, pass rate and cost per fix matter more.

    This short script turns results.csv into those three numbers:

    import csv, statistics
    from collections import defaultdict
    
    rows = list(csv.DictReader(open("results.csv")))
    by_task = defaultdict(dict)
    for r in rows:
        by_task[r["task"]][r["model"]] = r["passed"] == "True"
    
    a, b = "deepseek-ai/DeepSeek-V4-Pro-0813", "Qwen/Qwen3.8-Max-0902"
    for m in (a, b):
        rs = [r for r in rows if r["model"] == m]
        solved = sum(r["passed"] == "True" for r in rs)
        cost = sum(float(r["cost_usd"]) for r in rs)
        lat = sorted(float(r["latency_s"]) for r in rs)
        per_fix = f"${cost / solved:.3f}" if solved else "n/a"
        print(f"{m}: {solved}/{len(rs)} solved, ${cost:.2f} total, "
              f"{per_fix} per fix, "
              f"p50 {statistics.median(lat):.1f}s, p95 {lat[int(0.95 * (len(lat) - 1))]:.1f}s")
    
    only_a = sum(t[a] and not t[b] for t in by_task.values())
    only_b = sum(t[b] and not t[a] for t in by_task.values())
    print(f"Only DeepSeek solved: {only_a} | Only Qwen solved: {only_b}")
    

    What should you do after the evaluation?

    Ship the winner as the default through the same MaaS endpoint, since the model ID you tested is the model ID you deploy. Three follow-ups keep the decision current:

    • Re-run on every new dated build. Keep tasks.jsonl in the repo and re-run the harness when a new build of either model appears in the model library.
    • Consider per-request routing when results split by task type. If one model wins on refactors and the other on test fixes, GMI Router selects "the best-fit model for each request based on the task and your quality__ost objective," and routing is free during its preview. Production switching and fallback design is covered in how to switch between GPT-6 Astra, Gemini 3.8 Flash, and DeepSeek V4.1 Flash; see also Automatic Model Selection for Every Prompt: GMI Router Is Live.
    • Reuse the method for other workloads. The same paired design works for document extraction; see Gemini 3.8 Flash vs GPT-6 Astra for document extraction. If your app is not on the OpenAI SDK yet, start with the OpenAI SDK migration checklist. For how open-weight coding models stack up publicly before you start, read AI Model Benchmarks August 2026.

    FAQ

    How much does it cost to compare DeepSeek V4 Pro 0813 and Qwen3.8 Max 0902 on 100 coding tasks? On GMI Cloud MaaS, about $9.50 for one round at 20,000 input and 4,000 output tokens per task: $4.22 for DeepSeek V4 Pro 0813 and $5.28 for Qwen3.8 Max 0902, at catalog list prices as of September 2026.

    Three repeats cost about $29. Current per-token rates are listed in the GMI Cloud model library.

    How do I keep a coding model comparison fair? Change only the model ID. Use one endpoint and one SDK client, the same prompt template, temperature: 0, the same max_tokens, and interleaved run order, then grade by applying the diff and running your tests.

    GMI Cloud MaaS serves both models behind one OpenAI-compatible API, which removes the integration differences between two separate provider accounts.

    Should I use the dated model IDs or the undated ones? Use the dated IDs, deepseek-ai/DeepSeek-V4-Pro-0813 and Qwen/Qwen3.8-Max-0902, so the eval is reproducible.

    The GMI Cloud catalog lists undated DeepSeek-V4-Pro and Qwen3.8-Max entries as separate models with their own prices, so mixing the two gives you results you cannot trace back to a single build.

    How many tasks do I need before trusting the result? For a DeepSeek V4 Pro 0813 versus Qwen3.8 Max 0902 coding comparison, 100 paired tasks is a practical starting point, but the number that matters is how many tasks only one model solved.

    With 20 such tasks, the leader needs 15 wins to clear a two-sided sign test at 5%; with 40, it needs 27. Below that, the difference is not significant, so decide on cost per fix.

    Can I send proprietary code to the models during the evaluation? GMI Cloud MaaS offers zero-retention configurations for sensitive workloads, which is the setting to request before sending private repository content.

    Keep test execution in a sandbox either way, because the tests run code the model wrote.

    Start the comparison

    Create a MaaS API key in the GMI Cloud console, copy the two dated model IDs from the model library, and run the harness against 100 tasks from your own history.

    For endpoint details, see the Developers page; for current per-token rates, see the model library, and for dedicated GPU-hour rates, see pricing; for zero-retention or volume terms, contact the GMI Cloud team.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    On GMI Cloud MaaS, about $9.50 for one round at 20,000 input and 4,000 output tokens per task: $4.22 for DeepSeek V4 Pro 0813 and $5.28 for Qwen3.8 Max 0902, at catalog list prices as of September 2026. Three repeats cost about $29. Current per-token rates are listed in the GMI Cloud model library.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started