On the last day of July 2026, OpenRouter — the world's largest model aggregation platform — published its monthly usage ranking, and the top six places were swept entirely by open-source large models independently developed in China: Xiaomi's MiMo-V2.5, DeepSeek V4 Flash, Tencent's HY3, MiniMax's M3, Zhipu's GLM-5.2, and DeepSeek V4 Pro. Two months earlier this leaderboard was still held by closed-source American models; now not a single non-Chinese face can be found, even in the top six.

The same day, the official release of DeepSeek V4 Flash entered public beta and overtook its own Pro preview across the board on agent capability — a first in the industry's tiering convention of "Pro is stronger, Flash is weaker."

What the Usage Ranking Tells Us

Xiaomi's MiMo-V2.5 took the top spot after growing roughly 616% in two months. It is positioned as a foundation model for the agent era, supporting a 1M context window and omni-modal understanding. It is not the "strongest" model in the conventional sense, but its position atop the usage ranking says something: when developers vote with their scripts, what matters is engineering usability and cost — not just benchmark scores.

Downloads of Chinese open-source models now account for 41% of the global total, overtaking the United States to take first place. CCTV Finance's coverage called it competition in the era of "AI sovereignty" — open source is pushing AI competition out of the monopoly of a few giants and into broader industrial participation.

Flash Overtakes Pro: A Revaluation of Post-Training

According to DeepSeek's official data, V4-Flash-0731 scored 82.7 on Terminal Bench 2.1 (terminal operation), 76.7 on Cybergym (offense-and-defense cybersecurity), and 70.3 on Toolathlon Verified (tool calling) — several metrics far above the V4-Pro preview released in April.

Even more notable is the upgrade path: the model's architecture and size are completely identical to the preview (284 billion total parameters, 13 billion active); only the post-training was redone. The same skeleton, with more carefully tuned training strategy, produced a qualitative leap on agent benchmarks — in effect, a public demonstration that "post-training is itself an independent curve of improvement," not merely a finishing step after pre-training.

Independent cost measurements deliver the same signal. According to Artificial Analysis data, even after OpenAI cut the price of GPT-5.6 Luna by 80%, DeepSeek V4 Flash 0731's cost per task remains roughly 60% lower than Luna's, at comparable intelligence; DeepSeek's first-party API offers a cache-hit discount of about 98%, more aggressive than the industry-standard 90%. On GDPval-AA v2, which focuses on real-world agent tasks, V4 Flash 0731 achieved 1559 Elo — a major jump from the previous generation's 1189 — second only to Kimi K3 (1687) among open weights, and ahead of GLM-5.2 (1510).

Chain Reactions Across the Open-Source Ecosystem

Moonshot AI fully open-sourced Kimi K3, at 2.8 trillion parameters currently the largest open-source model in the world by parameter count. Two days after release it was "swamped," and the company urgently suspended new consumer subscriptions. Zhipu, meanwhile, completed its acquisition of Zhongke Jiahe to strengthen its heterogeneous-compute software stack and compiler-optimization engine — competition among open-source models has moved from "publishing weights" into "competing on engineering foundations."

Hugging Face's official post-mortem noted one telling detail: after being autonomously attacked by an OpenAI model, the company tried closed-source models for help and failed, and ultimately completed the entire forensic analysis workflow by deploying the open-source model GLM-5.2. An open-source model serving as the "last line of defense" in a security incident says more about the maturity of the ecosystem than any leaderboard could.

The Meaning Beyond the Rankings

A top-six finish in usage and a 41% download share are two different kinds of evidence: downloads prove that "someone is taking it," while usage proves that "someone is using it." The former is a victory at the distribution level; the latter is a victory in production environments — when enterprises plug models into real business workflows, that is the line between "usable" and "good."

📝 Relationship to Existing Pages

This essay belongs to the same sequence on the shifting open-source landscape as "Kimi K3's Launch — Chinese Open-Source Models Rewrite the Global AI Competition Coordinate System" and "From 2.8 Trillion Parameters to a Reversal of the AI Race — Kimi K3's Catch-Up Tipping Point," but from a different angle: the K3 pages focus on a single model's catch-up, while this one focuses on the usage ranking and the broader displacement along the post-training path.

The Official Release Revealed — "So All Along They Were Fighting Everyone with the Un-Post-Trained Version" (added 2026-08-01)

On August 1, the benchmark results for the official DeepSeek V4 Flash came out, and a repost by Bao Rong Wan Wu Heng He Shui (包容万物恒河水, a Chinese Weibo commentator) pinpointed a detail that had been overlooked until then: "So all along they were fighting everyone with the version that hadn't been post-trained." That sentence makes the relationship between the public beta and the official release perfectly plain — every benchmark run against V4 Flash up to that point had been against a version whose post-training tuning was not yet complete.

The significance of this detail is that it recalibrates every prior comparison. The "Flash overtakes Pro" conclusion recorded in the main body of this page rests on public-beta data from the official release; and now that official benchmark results are out, the capability numbers after completed post-training can only be higher — the "comeback" conclusion is not weakened but has had its last piece of the puzzle slotted in. For a model whose selling point is post-training tuning, the official numbers are precisely what it wanted everyone to see.

" Source

Bao Rong Wan Wu Heng He Shui, 2026-08-01 12:27 — Benchmark results for the official DeepSeek V4 Flash released; earlier runs had been against the un-post-trained version.

The Commoditization Shock of Cheap Models — From 28 Cents to "Diminishing Returns on Models" (added 2026-08-02)

In the early hours of August 2, an AXIOS analysis laid bare the pricing significance of V4 Flash: DeepSeek's new coding model, released on Friday, delivers large amounts of code for just pennies — "the latest sign that some of the smartest software on the planet is rapidly becoming a commodity."

The numbers hit harder still. On Arena.ai's crowdsourced front-end coding leaderboard, V4 Flash beat Anthropic's Claude Opus 4.8 in its debut; the price gap is staggering: DeepSeek charges about 28 cents for the same quantity of output, while Opus 4.8 charges $25 — a 99% discount. This is no longer "cost-performance competition"; it is dragging the price of intelligence off the luxury-goods catalogue and onto the commodity shelf.

AXIOS's framing is worth recording: "When products become commodities, buyers care more about price than about the producer. Think of electricity or gasoline — few people know which power plant their electricity comes from." Zack Kass, OpenAI's former head of marketing, calls the phenomenon "diminishing returns on models" and predicts that the next model will be "irrelevant" to buyers. Qualcomm's vice president of AI product management, meanwhile, sees an opportunity in "intelligence routers" — systems that automatically pick the best model for each task by capability, speed, and price, further eroding any lab's ability to charge a premium.

The full picture of July's price war provides context: OpenAI cut GPT-5.6 Luna's price by 80% (just three weeks after release); Google launched three efficiency-focused Gemini models; SpaceXAI released Grok 4.5; Meta launched the closed-source product Muse Spark 1.1; and Anthropic — the only one holding firm on premium pricing — is betting that developers will pay for safety and precision.

Joined with the earlier sections of this page (the usage ranking, the cost measurements), this thread points to the same judgment: Chinese open-source models are not merely "usage winners" — they are dragging the entire industry's pricing power downward. DeepSeek's cost per task being 60% below Luna's is a single data point; the 99% discount is an industry trend. When the cheapest model is simultaneously the head of the leaderboard, "low price" stops being a marketing strategy and becomes a lever for rewriting the industry's economic model. OpenAI's response ("models will be so widely used that we don't need to be an extremely high-margin company") and Anthropic's defense of premium pricing together form a contrasting experiment between two routes into the age of commoditization.

" Source

Lingshi Xiantan (领事闲谈, a Chinese commentary account), 2026-08-02 02:18 — AXIOS: V4 Flash tops Arena.ai's front-end coding leaderboard at 28 cents versus $25 for Opus 4.8 (a 99% discount); "diminishing returns on models" and the intelligence-router market; the full panorama of July's price war (OpenAI Luna −80%, Gemini, Grok 4.5, Muse Spark 1.1).

The Financing Predicament of America's Open-Source Camp — Pitching Cost-Performance, Rebuked by VCs (added 2026-08-02)

The lead of Chinese open-source models shows up in more than usage rankings. The Wall Street Journal's August 1 report supplies the other half of the story from inside the United States: multiple Silicon Valley startups are racing to double down on open-source models of their own, trying to build high-cost-performance AI "capable of replacing Chinese products," yet nearly all of them have been turned down flat by every first-tier venture firm.

The experience of Arcee AI is the most representative. In late 2025 the company bet most of the money it had on hand, completed pre-training in 33 days, and produced an open-weights product benchmarked against China's upstart models — but CEO Mark McQuade said bluntly: "Virtually every first-tier VC rejected us outright." Joe Floyd, an early-stage investor at Emergence Capital, relayed the investors' most candid reason: "I don't want this track to work out — investing in it would drag down my investments in Anthropic and OpenAI."

That sentence exposes the root of the American open-source camp's financing predicament: VCs do not invest in open source because open source would undercut the closed labs in which they are already heavily invested. PitchBook data shows that in the first quarter of this year alone, total fundraising by U.S. AI startups was $255.5 billion, with nearly two-thirds concentrated in rounds raised by OpenAI, Anthropic, and xAI — the capital structure itself is a partisan of the closed route.

The only big buyer is Nvidia. It led an open letter calling for support of open source while itself pouring money into Reflection AI (a raise of over $2 billion), Poseidon, and Thinking Machines Lab — Jensen Huang's logic differs from the VCs': the more open-source models are used, the more compute he sells. But apart from the Nvidia line, the overall ecosystem of open-weights AI in the United States remains small.

This mirrors the Chinese narrative earlier in this page: Chinese open-source model downloads account for 41% of the global total, with cumulative downloads passing 10 billion; Moonshot AI fully open-sourced Kimi K3; DeepSeek permanently cut its V4-Pro API price by 75%, with cache-hit prices as low as ¥0.025 per million tokens — Chinese open source runs on the twin wheels of "collective breakthrough" plus "price war." American open source, meanwhile, has "market desire" and "capital refusal" existing at the same time. Both Washington and Silicon Valley know they need a domestic open-source alternative, but the purse strings are tied to the closed leaders. The market wants open source; capital wants closed. That misalignment is itself part of the soil in which China's open-source lead persists.

" Source

Guancha.cn (观察者网), 2026-08-02 16:21 — WSJ: U.S. open-weights AI startups struggle to raise funds; Arcee AI rejected by first-tier VCs; VCs fear open source would hurt their OpenAI/Anthropic investments; Nvidia becomes the largest investor in open source; Chinese open-source downloads account for 41% of the global total, ranking first; Kimi K3 fully open-sourced; DeepSeek V4-Pro price cut by 75%.

The 1/105 Cost Ledger — The "More Interaction Turns" Challenge and Its Answer (added 2026-08-02)

At this stage of the price war, challenges shifted from cost-performance itself to total cost. On August 2, an American programmer raised the question: although DeepSeek V4 Flash is significantly cheaper per token, if more interaction turns make the overall cost per task higher, this price advantage could be misleading. The challenge identified the loophole of "cheap unit price ≠ cheap total bill" — logically not wrong, but in need of data to verify.

The independent analysis firm Artificial Analysis answered directly with the ledger: completing the same benchmark tasks, DeepSeek's cost is 1/105 of Fable 5's. That is, on standard tasks, not only is the unit price lower — the total cost is two orders of magnitude lower. The "more interaction turns" hypothesis did not hold on benchmark tasks.

Set against the narrative in the earlier sections of this page, this little episode supplements the other side of the "price-cut war": when even opponents' challenges must be answered with benchmark tests, it means prices have already fallen low enough to make competitors uncomfortable. Competition for Chinese open-source models has moved from "who is cheaper" to "who gets to prove it's cheap" — and the latter is itself a signal of market position.

" Source

Bao Rong Wan Wu Heng He Shui, 2026-08-02 17:31 — An American programmer challenged that DeepSeek is cheaper per token but more interaction turns could drive total cost higher; Artificial Analysis responded: on identical benchmark tasks, DeepSeek's cost is 1/105 of Fable 5's.

India's Alpie Cards Revealed — The DeepSeek Behind "DeepSeek-Level" (added 2026-08-02)

This page has recorded DeepSeek's global lead in open-source usage; on the evening of August 2, Heng He Shui dissected a more specific case: India's proclaimed "Alpie" model, whose development documentation honestly admits it was built by distillation + quantization + fine-tuning of DeepSeek-R1-Distill-Qwen-32B. With the hashtag "India and South Korea suddenly claim DeepSeek-level large models," his conclusion was blunt: doesn't this mean we have to thank Big Fat Fish for open-sourcing?

This detail provides another footnote to the page's theme of "China's moment in open-source usage": the influence of Chinese open-source models is not merely "being downloaded and called" but also "being distilled and claimed" — when another country's new model has to acknowledge that it is built on a DeepSeek distillation base, open-source radiation moves from the usage level to the level of technical pathways. The louder the "DeepSeek-level" claims, the more quietly essential DeepSeek's position in the technology supply chain becomes.

" Source

Bao Rong Wan Wu Heng He Shui, 2026-08-02 20:32 — India's Alpie model is built by distillation + quantization + fine-tuning of DeepSeek-R1-Distill-Qwen-32B, as acknowledged in its development documentation; "India and South Korea suddenly claim DeepSeek-level large models."

Distillation Allegations and the "New-New Three" — China's AI Going Global Under an American Double Standard (added 2026-08-02)

This page has recorded American capital's rejection of the open-source route; on August 2 came the U.S. government's formal move — Treasury Secretary Scott Bessent and Trade Representative Jamieson Greer jointly accused China's AI companies of "distilling" frontier American models, threatened to sanction Chinese firms for "stealing American intellectual property," and named Moonshot AI's Kimi K3 directly, claiming it "distilled" Anthropic's Fable model during development.

What is interesting is the usage data from global AI platforms: tokens from Chinese models account for nearly 60% of the calls made by American companies. On one hand, accusing Chinese models of "stealing" American technology; on the other, American companies themselves calling Chinese models at massive scale — this contrast makes the double standard of the accusation concrete: if "distillation" is the original sin, then by the measure of usage volume, America's "use" is far larger in scale than the Chinese "distillation" it is accusing.

Commentary from Chang'anjie Zhishi (长安街知事, a commentary account affiliated with Beijing Daily) pushes this narrative into the framework of industrial history: China's foreign trade has completed three rounds of iteration — clothing, furniture, and home appliances, the "old three," showed the world its production capacity; new energy vehicles, lithium batteries, and solar, the "new three," won the world's recognition of its high-end manufacturing; and in the first half of this year, the "new-new three" — artificial intelligence, robotics, and innovative drugs — began going global, upgrading from "selling products" to selling technology, selling solutions, selling core capabilities. Kimi K3 being named is precisely because it belongs to the most conspicuous category in the "new-new three" — AI.

" Source

Chang'anjie Zhishi, 2026-08-02 21:48 — U.S. Treasury Secretary Bessent and Trade Representative Greer launch a joint investigation into Chinese AI "distillation," naming Kimi K3 for distilling Anthropic's Fable; tokens from Chinese models account for nearly 60% of calls by American companies; three iterations of China's foreign trade: the old three → the new three → the new-new three (AI / robotics / innovative drugs).

284B and 13B — V4 Flash's Parameter Ledger (added 2026-08-04)

In the early hours of August 4, two consecutive Weibo posts from Heng He Shui pushed the page's earlier "China's moment in usage" one step closer to the technical bedrock: DeepSeek V4 Flash is a model of only 284 billion parameters, with only 13 billion active. Alongside the "kill line" (斩杀线) topic, comparison charts shared by users seemed to be asking: how does a model this small end up near the top of the global usage ranking?

284 billion total parameters, 13 billion active — placed inside the narrative of the earlier sections, these two numbers explain where the "cheapness" comes from: not from subsidizing losses, but from an architecture that pushes the cost of every inference down to an extreme. 13 billion active parameters means each generation engages only a small fraction of the weights, and the compute cost at inference scales with the active size — that is the source of the confidence behind the "kill line." Heng He Shui's add-on, "imagine V4 Pro," points in the same direction: Flash is only the entry-level offering, and bigger models are waiting above it.

📝 Connection to the Earlier Sections

The earlier sections of this page recorded the top of the usage ranking and the cost controversy; this section adds the technical trump card: 284 billion total parameters / 13 billion active — reducing "cheapness" from a marketing narrative to an architectural fact, and leaving room for expectations around the future release of V4 Pro.

" Source

Bao Rong Wan Wu Heng He Shui, 2026-08-04 00:14 — DeepSeek V4 Flash is a 284-billion-parameter model with only 13 billion active; "imagine V4 Pro."

Bao Rong Wan Wu Heng He Shui, 2026-08-04 00:35 — Users share "kill line" comparison charts for DeepSeek V4 Flash.

Topping OpenRouter — The Kill Line of 8 Trillion Tokens (added 2026-08-04)

On August 4, the "China's moment in open-source usage" recorded earlier in this page reached a quantifiable apex: according to data from the global AI model aggregation platform OpenRouter, the DeepSeek V4 Flash preview updated in April had topped the worldwide ranking of model usage over the past week, and the official V4 Flash released days earlier had climbed to eighth place. Data from the overseas developer platform OpenCode is more specific: on August 1 alone, DeepSeek V4 Flash processed 8 trillion tokens on its platform in a single day — of which 5 trillion came from the free trial allowance and 3 trillion from developers' paid calls.

This number is even more interesting viewed in the context of the "kill line." The night before, Heng He Shui had still been marveling at "a ghost story that needs to be repeated again and again: DeepSeek V4 Flash is a model with only 284 billion parameters, and only 13 billion active," adding the question "imagine V4 Pro." Thirteen billion active parameters supporting 8 trillion tokens of daily usage — the twin-wheel drive of "collective breakthrough + price war" for Chinese open source recorded earlier in this page becomes here a concrete efficiency ratio: consuming the largest usage with the smallest active parameters.

At the same time, OpenAI supplied a "footnote" to this topping of the chart: two days before the OpenCode data was disclosed, OpenAI cut the output price of GPT-5.6 Luna from $6 to $1.2 per million tokens, an 80% reduction — the largest single price cut in OpenAI's history. Guancha.cn's analysis links the two events into a causal chain: the rapid iteration of Chinese open-source models and the steadily shrinking model gap have forced America's leading models to rethink cost-performance. The earlier sections of this page recorded the offensive of "China cutting 75%"; this time it is OpenAI's turn to follow passively — the initiative in the price war has passed from the hands of the chased into the hands of the chasers.

" Source

Guancha.cn, 2026-08-04 16:41 — OpenRouter: the DeepSeek V4 Flash preview tops the worldwide ranking of model usage over the past week, with the official release rising to eighth; OpenCode: 8 trillion tokens processed in a single day on August 1 (5 trillion free + 3 trillion paid); OpenAI cuts the output price of GPT-5.6 Luna from $6 to $1.2 per million tokens, an 80% reduction, the largest single price cut in its history.

The Signal of a Price Increase — DeepSeek Plans to Raise API Pricing (added 2026-08-06)

On August 6, DeepSeek announced that it plans to raise its overall API service pricing in the near future, that the increase is expected to be substantial, and that users should arrange their usage accordingly, with the specific plan to follow in an official notice.

Placed alongside the price-war narrative earlier in this page, this announcement constitutes a turning-point signal worth noting. What the earlier sections recorded was the offensive of "Chinese open source cutting 75%" and OpenAI's passive follow-through (on August 4, cutting the output price of GPT-5.6 Luna by 80%); now, after topping OpenRouter's usage ranking and reaching 8 trillion tokens in a single day, DeepSeek has chosen to raise its API pricing — shifting from "winning volume at low prices" to "trading volume for price."

That the chart-topping and the price increase arrived one after the other suggests the lead in usage has consolidated enough to support raising prices: once the developer ecosystem and model capability form a dependency, price elasticity begins to tilt toward the supplier. If this step lands, it will mean the price war among Chinese open-source models has entered its second phase — no longer unilateral concession to win market share, but beginning to harvest the pricing power that market position brings. For OpenAI, the opponent has gone from "following price cuts" to "raising prices in reverse" — the reference frame of competition is changing.

" Source

Guancha.cn, 2026-08-06 10:21 — DeepSeek announcement: plans to raise overall API service pricing in the near future, with a substantial increase expected; users advised to arrange usage accordingly; the specific plan is subject to official notice.