• 運算
  • 客戶
  • 價格
登入
More Blog Posts
XDiscordLinkedInYouTube

產品

  • GPU
  • MaaS
  • Studio

開發者

  • 模型總覽
  • 技術文件
  • 詞彙表

公司

  • 關於我們
  • 部落格
  • 活動
  • 合作夥伴
  • 新創計劃
  • 職涯
  • 大使計畫
  • 使命與願景

熱門模型

    掌握 AI 最新動態

    提交即表示您瞭解我們會收集並使用您提交的資訊,其中可能包含個人資訊。

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    隱私政策使用條款法律文件
    More Blog Posts
    Research

    Flash Used to Mean Compromise. Four Models Just Proved It Doesn't.

    DeepSeek V4.1 Flash, Gemini 3.8 Flash, GLM 5.3 Flash, and Step 3.7 Flash compared on agentic coding benchmarks, latency, speed, and cost per turn. Which is the best flash model for your workflow?

    2026年10月06日

    Two Flash models are now within a point of, or ahead of, Claude Opus 5 on agentic coding while costing 5% and 15% as much per output token. I tested DeepSeek V4.1 Flash, GLM 5.3 Flash, Step 3.7 Flash, and Gemini 3.8 Flash using the same four prompts in GMI’s Model Arena, then compared the labs’ benchmark claims with independent reruns wherever those were available. The video captures the runs. This post covers the results that held up.

    Not long ago, “Flash” in a model name signaled a tradeoff. You chose it when speed and price mattered more than top-end capability, knowing there would be a meaningful performance gap.

    That gap has narrowed sharply. On September 10, DeepSeek released V4.1 Flash and said it outperformed V4 Pro, its own flagship, on performance, cost, speed, and total time. The company intended to retire V4 Pro four days later, keeping it available only after user feedback. When a lab’s lower-cost model can make its premium model feel unnecessary, the old tier labels no longer carry the same meaning.

    The four at a glance

    DeepSeek V4.1 Flash

    Gemini 3.8 Flash

    GLM 5.3 Flash

    Step 3.7 Flash

    Released, 2026

    Sep 10

    Sep 2

    Aug 26

    May 29

    Active parameters

    8B read, 16B generate

    not disclosed

    18B

    11B

    Context

    1M

    1M

    1M

    256K

    Inputs

    text, image

    text, image, audio, video

    text, image, video

    text, image

    Weights

    MIT

    closed

    MIT

    Apache 2.0

    Output price per 1M tokens

    $1.20

    $3.75, rising to $7.50 on Jan 1

    $0.50

    $1.15

    The three open-weight models use mixture-of-experts architectures, activating between 8 billion and 18 billion parameters per token. Google has not published Gemini’s architecture. All four are available through GMI Cloud with just one model ID changed.

    The models

    DeepSeek V4.1 Flash is DeepSeek’s third Flash release in five months and its first built on a new architecture that DeepSeek calls a causal encoder-decoder. Its most consequential figure is KV-cache usage: roughly 890 bytes per token, about one-quarter of V4 Flash. That makes it particularly economical on long agent histories, where cached context often becomes the largest part of the bill. It supports vision natively, and the weights are released under MIT.

    Gemini 3.8 Flash arrived 20 days after 3.7 Flash and builds directly on that foundation. Google’s core argument is that the model can expend more effort on difficult tasks: additional reasoning steps, tool calls, and output tokens, with an effort setting that lets users reduce that behavior. It is the only model in this group with native audio input, and its $0.75 input and $3.75 output prices are scheduled to double on January 1.

    GLM 5.3 Flash spent six days atop OpenRouter’s coding charts as an anonymous free model called Ox Alpha before Z.ai identified it on August 26. It is Z.ai’s first natively multimodal GLM, processing text, images, and video within the same model. Its hybrid sparse-plus-linear attention stack is designed to reduce KV-cache requirements by more than four times at long context, according to Z.ai. It is also the lowest-priced model in this group.

    Step 3.7 Flash reflects StepFun’s emphasis on agent workflows rather than leaderboard placement. It combines a 196B mixture-of-experts backbone with a 1.8B vision encoder, and its benchmark table foregrounds BrowseComp and DeepSearchQA instead of Terminal Bench. Its Advisor Mode allows the Flash model to execute a task end to end while consulting a larger model only at difficult decision points. StepFun reports that this approach reaches 97% of Claude Opus 4.6’s coding performance at $0.19 per task, compared with $1.76.

    What the labs claim, and what holds up

    Terminal Bench 2.1 is the agentic-coding benchmark reported by all four labs. Vals.ai reruns it independently using its own harness.

    Model

    Lab-reported

    Vals.ai independent, Sep 27

    Gap

    DeepSWE 1.1, lab-reported

    DeepSeek V4.1 Flash

    90.6

    74.5

    −16.1

    74.2

    Gemini 3.8 Flash

    89.4

    81.3

    −8.1

    73.7

    GLM 5.3 Flash

    84.3

    62.9

    −21.4

    63.4

    Step 3.7 Flash

    59.6

    not run

    not reported

    Claude Opus 5, reference

    89.1

    84.6

    −4.5

    74.0

    StepFun’s blog lists Step 3.7 Flash at 59.6; its own model card says 59.5.

    Three conclusions stand out.

    The headline is supported by the labs’ own numbers. Two Flash models are within a point of, or ahead of, Claude Opus 5 on both coding benchmarks, while priced at $1.20 and $3.75 per million output tokens versus $25 for Opus 5.

    Flash-model claims contract more in independent testing than flagship claims. Under Vals’ rerun, the Flash models fell by 8 to 21 points, while Opus 5 dropped by 4.5 points. The ranking changes as well: Gemini leads the group in the independent run by nearly seven points. Flash benchmark tables deserve a wider confidence margin than flagship tables.

    The underlying improvement is still evident even when absolute scores are lower. Vals measured a 3.7-point increase from Gemini 3.7 to 3.8 Flash and a 7.5-point increase from DeepSeek V4 Flash (the 0731 build) to V4.1 Flash. The labs reported gains of 3.6 and 7.9 points, respectively.

    One other independent figure is useful context: Gemini scored 3.5 points above DeepSeek on the Vals Index, but cost 17 times more to run across the suite, $5.73 versus $0.33.

    Four prompts in Model Arena

    All four models answering the same two-sentence MoE prompt in Model Arena, with the latency breakdown and speed-versus-cost chart below.

    I gave all four models the same prompts at their default settings: a two-sentence explanation, a five-item list, a canvas particle system, and a CSS-only aurora. The per-prompt results are in the video. These totals come from one run per model, so they are directional rather than benchmark results.

    DeepSeek V4.1 Flash

    Gemini 3.8 Flash

    Step 3.7 Flash

    GLM 5.3 Flash

    Completed as asked

    4 of 4

    4 of 4

    3 of 4

    2 of 4

    Time to first token

    1.3 to 1.9 s

    3.7 to 76.6 s

    2.1 to 11.2 s

    1.4 to 4.4 s

    Total time

    73 s

    201 s

    177 s

    403 s, fourth unfinished

    Total cost

    $0.016

    $0.049

    $0.033

    $0.006, three prompts

    Answer share of output tokens

    41%

    64%

    100%

    2%

    Median generation speed

    203 tok/s

    88 tok/s

    160 tok/s

    46 tok/s

    DeepSeek completed all four prompts in 36% of Gemini’s time and at 32% of Gemini’s cost. Its first token also arrived in no more than 1.9 seconds in any run.

    For Gemini, the delay before the first token largely reflects thinking time. It rose from 3.7 seconds on the easiest prompt to 77 seconds on the most difficult. In a chat interface, use a lower effort setting or stream the reasoning process.

    GLM produced 18,510 output tokens across three prompts but only 385 answer tokens. On both coding prompts, it reasoned until reaching the token cap and returned no final answer. The lowest token price does not necessarily produce the lowest cost per completed response. Put a cap on reasoning and verify that a final output is actually returned.

    Step neither reasoned nor stopped generating: it used 10,768 tokens to produce a five-item list. Set max_tokens.

    The CSS-only aurora prompt rendered by DeepSeek, Step, and Gemini. GLM hit its reasoning cap and returned no output.

    Speed vs. cost

    Artificial Analysis continuously measures speed and latency through each lab’s own API. Its figures below were read on October 1, 2026, with DeepSeek at max effort and Gemini at high effort, and they drift from week to week. My Model Arena medians come from four prompts routed through GMI endpoints. The ordering is consistent across both sets of measurements.

    DeepSeek V4.1 Flash

    Gemini 3.8 Flash

    GLM 5.3 Flash

    Step 3.7 Flash

    Output speed, Artificial Analysis

    209 tok/s

    245 tok/s

    47 tok/s

    191 tok/s

    Time to first token, Artificial Analysis

    0.99 s

    17.5 s

    3.3 s

    2.7 s

    List price, input / output per 1M tokens

    $0.30 / $1.20

    $0.75 / $3.75

    $0.15 / $0.50

    $0.20 / $1.15

    One agent turn, 100K in + 20K out

    $0.054

    $0.150

    $0.025

    $0.043

    Same turn from January 1

    $0.054

    $0.300

    $0.025

    $0.043

    Gemini generates fastest once it begins, but it also has the longest startup delay. At high effort, Artificial Analysis measured 17 seconds to its first token; my median was 39 seconds. That is the tradeoff Google selected, and the effort control is the mechanism for reducing it.

    The agent-turn row is the relevant budgeting figure. At current pricing, one Gemini turn costs roughly six GLM turns. Starting January 1, it costs twelve, and more than five DeepSeek turns.

    Where I would route each one

    • Coding agents and latency-sensitive work: DeepSeek V4.1 Flash.

    • Low-cost multimodal and knowledge work: GLM 5.3 Flash, with a reasoning cap. Its 1773 on GDPval-AA exceeds Gemini’s 1545 on Google’s own card.

    • Audio input and the strongest independent coding result: Gemini 3.8 Flash. Budget for $7.50 output pricing from January 1.

    • Search and tool-use agents: Step 3.7 Flash, with 75.8 on BrowseComp and roughly 190 tokens per second. Pair it with a larger advisor model; StepFun reports 97% of Claude Opus 4.6’s coding performance at about one-ninth of the cost using that approach.

    Try it on GMI Cloud

    Every model here is available through one base URL and one API key. Switching among them requires changing only a single line in the request.

    curl https://api.gmi-serving.com/v1/chat/completions \
      -H "Authorization: Bearer $GMI_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "google/gemini-3.8-flash",
        "messages": [{"role": "user", "content": "Explain MoE architecture in 2 sentences"}]
      }'

    Replace the model field with deepseek-ai/DeepSeek-V4.1-Flash, zai-org/GLM-5.3-Flash, or stepfun-ai/Step-3.7-Flash to run the same prompt against the other models.

    Model Arena: Select the four Flash models, paste in one of your prompts, and inspect the latency and cost charts for your own workload rather than relying on mine.

    The question to answer

    Go beyond asking which Flash model is best. Ask which of your current flagship calls a Flash model can now absorb.

    Choose one route in your product, run your real prompts through the arena, and study the speed-versus-cost chart. If a Flash model lands in the bottom-right quadrant while maintaining acceptable output quality, that route no longer needs a flagship model.

    Ready to build? Get Started

    Grace Deng

    Developer Relations

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started