• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Announcements

    DeepSeek V4 Pro Steps Out of Preview: The 0813 Build Is Live

    DeepSeek V4 Pro leaves preview with the official 0813 build. Explore specs, pricing, architecture, and how it compares to Claude Opus 4.8.

    August 12, 2026

    DeepSeek's flagship model has left preview status. As of August 12, 2026, the deepseek-v4-pro endpoint on DeepSeek's official API now points to a new build: DeepSeek-V4-Pro-0813. The change is confirmed directly on DeepSeek's own API documentation and pricing pages, closing out a preview window that stretched back to the model's original debut on April 24, 2026.

    The Architecture Behind It

    V4 Pro is a Mixture of Experts system built at massive scale: 1.6 trillion total parameters, with roughly 49 billion active per token. The model uses hybrid attention mechanisms aimed at cutting inference costs at long context lengths. The model was pre-trained on more than 32 trillion tokens.

    Specs at a glance:

    • Total parameters: 1.6 trillion (MoE), about 49B active per token

    • Context window: 1 million tokens

    • Max output: 384,000 tokens

    • Reasoning modes: non-thinking, high effort, and max effort ("V4 Pro Max")

    • Extras: tool calling, JSON output, agent developer support

    Pricing Breakdown

    Pro pricing carries over from the preview period: approximately $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 to $0.93 per million output tokens depending on provider. Concurrency is capped at 500 requests for Pro, versus 2,500 for Flash, positioning Pro for heavier reasoning and long-context coding work rather than high-throughput traffic.

    How This Fits the V4 Rollout

    DeepSeek staged this release deliberately. The V4 series first appeared as open-weight previews on April 24, 2026, with both Pro and Flash released under the MIT license. On July 31, 2026, the smaller Flash model graduated to official status first, with DeepSeek's change log stating the Pro release "will follow soon". That promise is what materialized as the 0813 build.

    DeepSeek's own agent benchmarks at the time showed the reworked Flash outscoring the then preview V4 Pro on internal coding agent suites, a sequencing choice that made Flash the default agent workload model while Pro finished cooking in preview.

    Where It Sits Against the Competition

    Reported figures place V4 Pro Max close to several named frontier systems: it lands near 80.6% on SWE-bench Verified, and beats Anthropic's Claude Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench, though Opus remains ahead on SWE-bench Verified itself. An earlier preview build scored 45 on the Artificial Analysis Intelligence Index. None of these figures have been independently confirmed for the 0813 build specifically yet.

    Roan Weigert

    Roan Weigert

    DevRel @ GMI Cloud

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started