On September 10 DeepSeek shipped V4.1-Flash: the smallest model in its new architecture family, 552B parameters with just 8B active on input and 16B on output, weights on Hugging Face under MIT, native multimodal, 1M context, live on the API now. And in the same announcement the company said the flagship V4-Pro — the model that was the best thing DeepSeek made — starts routing to V4.1-Flash on September 14, at V4.1-Flash rates. The flagship is being retired by the thing it was supposed to outrun.
The architecture is the story. A new Causal Encoder–Decoder design splits the model asymmetrically: 8B active parameters on input, 16B on output, from a 552B-parameter MoE backbone. The KV cache needs 1/4 the HBM and 1/8 the SSD storage of the previous generation — and cache-hit charges are a large share of agent-workload costs, so compressing the cache compresses the bill. Off-peak rates are half of peak. The numbers make the argument: frontier-class results at commodity prices.
The open-weights release is the second half. Weights on Hugging Face under MIT license, tech report published alongside, and DeepSeek explicitly asking for help with large-scale inference deployments (“2,000 GPUs + a storage cluster? Let’s talk.”). The company is distributing the capability it just cannibalized its own flagship with, and inviting the world to run it.
The business context makes it sharper. Reuters reported September 9 that DeepSeek has tapped CITIC Securities to prepare a domestic IPO on the STAR Market — process to begin this year, valuation undecided. A company heading toward a public listing just released its best model as open weights and cut its own pricing floor. DeepSeek is pricing like it wants the category, not the margin — and it holds one cost advantage no US lab can match: the compute is already paid for.
The Commodity Frontier
Every capability announcement from the US labs now lands in a market where an open-weights alternative undercuts it within weeks. The frontier price is not a price — it is a ceiling that DeepSeek keeps pushing down. The interesting question is no longer whether open models catch the flagship; it is whether the flagship category survives being priced at the cost of inference.
The Takeaways
- DeepSeek V4.1-Flash launched September 10: 552B MoE, 8B active parameters on input and 16B on output, 1M context, native multimodal, weights MIT-licensed on Hugging Face.
- The KV cache needs 1/4 the HBM and 1/8 the SSD storage of the previous generation; off-peak API rates are half of peak, with new pricing effective September 10.
- DeepSeek’s own flagship V4-Pro routes to V4.1-Flash starting September 14 at Flash rates — the company is cannibalizing its best model.
- Reuters (September 9): DeepSeek tapped CITIC Securities to prepare a domestic IPO on the STAR Market, with the process to begin this year.

