Skip to content
← All news
3 min read

DeepSeek shipped V4-Flash as MIT open weights at $0.14 per million tokens

A 284 billion parameter mixture of experts with a 1 million token context, re-post-trained for agent work, released under a license with no conditions attached.

A million token agentic model with no license conditions, at fourteen cents per million tokens.

On July 31, 2026 DeepSeek published the 0731 build of V4-Flash, turning what had been a preview into the official production model and putting the full weights on Hugging Face: 166.9 GB across 48 safetensors shards, under an MIT license with no commercial conditions attached. The model is 284 billion parameters total with 13 billion active, a 1 million token context, and up to 384,000 output tokens.

What actually changed

DeepSeek's own changelog is unusually plain about the scope: same architecture, same size, only re-post-trained, with what it calls significantly enhanced agent capabilities. The published numbers are Terminal Bench 2.1 at 82.7, Cybergym at 76.7, Toolathlon verified at 70.3, DSBench-FullStack at 68.7, DeepSWE at 54.4, and NL2Repo at 54.2. Two practical additions matter more than any of those: it natively supports the Responses API format, and it is specifically adapted for Codex. That is a model shaped for agent harnesses rather than for chat.

Read the benchmarks the way you read any vendor's

Every figure above is DeepSeek's own and none has been independently reproduced. That is not an accusation, it is the normal state of a launch-week model, and it is the same caution we applied to Alibaba's numbers. Artificial Analysis lists the model independently at $0.14 per million input tokens and $0.28 per million output, with cache hits at $0.003, and measures throughput around 104 tokens per second. Pricing and throughput are the parts an outside party can verify today. The agentic scores are the parts to re-test yourself on your own tasks.

Why a build studio cares

MIT is the whole story here. A frontier-adjacent agentic model at fourteen cents per million input tokens is interesting; one you can download, self-host, fine-tune, and ship inside a client product with no revenue threshold, no field-of-use restriction, and no negotiation is a different category of thing. That is worth stating plainly this week in particular, because the licensing direction elsewhere in open weights is moving the other way, and an unconditional license is becoming the exception rather than the default.

Next step: read the DeepSeek changelog and pull the weights from Hugging Face. If you want a self-hosted model wired into a product properly, write to us at hello@gattyworks.com.

DeepSeekOpen WeightsInference CostDeepSeekOpenWeightsMITLicenseHuggingFaceAgenticAILLMAICodingInferenceCostMachineLearningAI

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.