Stack Signal.
Technical news & guides across AI, programming and the open-source world
SATURDAY, OCTOBER 03, 2026 · 66 articles · RSS

AI & Machine LearningOct 03, 2026587 words

DeepSeek-V4-Flash: Open Weights That Are Actually Affordable to Run

DeepSeek has a habit: it ships frontier-scale open weights, and everyone else has to redo the math on whether they can actually run the things. The V4 preview follows the same pattern, but the model that deserves attention this time is not the biggest one. For developers who care about self-hosting and inference cost, the release reads as an efficiency story first.

Two models, one release

DeepSeek released the V4 series as a preview with two Mixture-of-Experts (MoE) open-weights models. DeepSeek-V4-Pro is the frontier play: 1.6 trillion total parameters with 49 billion activated per token. DeepSeek-V4-Flash is the efficient one: 284 billion total parameters, just 13 billion activated. Both support a one-million-token context window.

Both sets of weights are MIT-licensed and hosted on Hugging Face and ModelScope. That means fully open and fully self-hostable. The license matters as much as the architecture: if your team cannot ship training data or prompts to an external API, those two mirrored releases are the whole point.

Where the savings come from

The efficiency story is architectural. V4 pairs a hybrid attention scheme that combines two new mechanisms: Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). The payoff shows up at long context. In the 1M-token setting, V4-Pro needs only about 27% of the single-token inference FLOPs and roughly 10% of the KV cache next to DeepSeek-V3.2. That is the difference between "supports 1M tokens on paper" and "is actually practical to serve at 1M tokens."

Two quieter additions round it out. Manifold-Constrained Hyper-Connections (mHC) improve residual signal propagation, and the Muon optimizer speeds up convergence. Training itself ran on more than 32 trillion tokens, then a two-stage post-training pipeline took over: per-domain expert SFT plus GRPO reinforcement learning, followed by on-policy distillation to fold the domain skills back into a single model.

Pick your reasoning effort

V4 exposes three reasoning modes. Non-think is fast and intuitive. Think High does deliberate, conscious logical analysis. Think Max applies maximum reasoning, and with a recommended context window near 384K tokens, it leans on that long-context work. The dial matters for self-hosters: you do not have to commit to a fixed inference budget up front, and a local run of Think Max can stay inside that 384K window.

Start with V4-Flash

Flash is the model most teams should look at first. It claims near-Pro reasoning from a far smaller activated budget: 13 billion parameters versus Pro's 49 billion. Because only the activated parameters do the heavy lifting per token, that gap between 49 billion and 13 billion is exactly where the cost savings live.

That works out to a fraction of the inference cost of a 1.6-trillion-parameter frontier model, which is why Flash is aimed at high-volume, price-sensitive, and self-hosted deployments. Pro-Max still chases the very top of the open-source benchmark tables. If you pay per token or run your own hardware, Flash is the one to start with.

One naming caution. Several third-party write-ups are calling this release "V4.1 Flash" or quoting a 552-billion parameter count. Both are wrong. The official model is DeepSeek-V4-Flash at 284 billion total parameters, 13 billion activated. Trust the model card, not the blog headlines.

Why it matters

Open weights only matter if you can actually run them, and V4-Flash is a rare release that lowers the cost of strong reasoning instead of just raising the ceiling. It is MIT-licensed, self-hostable, and small enough that the one-million-token window reads as a practical feature rather than a spec-sheet boast. That combination is the reason to watch this one.