The story of DeepSeek

Liang Wenfeng was born in Guangdong, China in 1985. He co-founded a quantitative hedge fund called High-Flyer in 2015. By 2019 the fund was managing tens of billions of yuan through machine-learning-driven trading strategies.

Starting around 2021, Liang began doing something that was, at the time, considered odd by his industry peers. He started quietly accumulating Nvidia GPUs. Not to sell. Not to arbitrage. To stockpile.

By the time US export controls on advanced AI chips to China began tightening in 2022 and 2023, Liang had reportedly accumulated approximately 50,000 Nvidia A100s — a stockpile that would have been extraordinarily difficult, and eventually impossible, to acquire once the controls were fully in place.

In July 2023, Liang founded DeepSeek — an AI research company built explicitly on the compute infrastructure he had accumulated over the previous two years.

On January 20, 2026, DeepSeek released DeepSeek-R1 — a reasoning model that, on multiple benchmarks, matched or exceeded the frontier capabilities of leading US AI labs. The reported training cost: approximately $5.6 million. A tiny fraction of what US labs had spent to reach comparable capability.

On January 27, 2026, Nvidia lost approximately $600 billion in market capitalization in a single trading day. The market’s interpretation: if frontier AI could be trained for $5.6 million, the assumed structural moat around US AI infrastructure investment was substantially smaller than had been priced in.

The distillation accusation

In February 2026, Anthropic — the US AI lab that produces Claude — publicly stated that DeepSeek’s models appeared to have been trained substantially on outputs distilled from frontier US models, including Anthropic’s own.

The claim, if true, would substantially change the interpretation of DeepSeek’s $5.6 million training cost. Distilling from an already-trained frontier model is dramatically cheaper than training a frontier model from scratch. The impressive engineering achievement remains real. The pricing comparison to US labs becomes misleading.

Liang’s response, paraphrased across multiple interviews: every AI is trained on someone else’s work.

He is not wrong. But he is not fully right either. There is a specific difference between training on public data that was created by third parties for other purposes, and training on outputs that were specifically generated by another AI lab’s investment in frontier capability. The distinction matters for the economics of the industry.

The lessons for US AI executives

Lesson one: strategic patience wins. Liang’s 2021-2023 stockpiling of GPUs was, at the time, considered eccentric. In 2026, it is the foundation of the most disruptive AI event of the year. Strategic patience is systematically undervalued because it looks like inaction in the moment.

Lesson two: constraints produce ingenuity. DeepSeek’s engineering achievements — including highly efficient training approaches — were driven substantially by the compute constraints that US export controls imposed. Executives running well-resourced US labs should assume, structurally, that resource-constrained competitors will produce more efficient approaches than well-resourced incumbents.

Lesson three: the moat is smaller than you think. Whatever the truth of the distillation accusation, the market’s response to DeepSeek-R1 is instructive. The assumed durability of frontier US AI capability was priced higher than the actual reality supports. This is a systematic bias in most enterprise AI investment theses right now.

Three practical questions

One: what is your equivalent of Liang’s GPU stockpile? Some resource, capability, or positioning that would be extraordinarily difficult to acquire in five years but is available cheaply now. Are you accumulating it?

Two: what constraints could your organization impose on itself deliberately? DeepSeek’s efficiency was produced by external constraints they could not choose to avoid. Well-resourced organizations rarely produce comparable efficiency because they rarely impose comparable constraints. Would deliberate constraint improve your outcomes?

Three: what is your AI moat, actually? Not aspirationally. Actually. What specific capability or positioning would competitors be unable to replicate in eighteen months? If nothing, your moat is priced into your competitors’ plans, not your own.

The closing thought

Liang Wenfeng’s DeepSeek story is, at one level, a story about Chinese AI capability catching up to US frontier labs faster than most analysts expected. It is also a story about strategic patience, constraint-driven ingenuity, and the systematically overestimated durability of first-mover advantage in fast-moving technology.

The $600 billion Nvidia sell-off was not, primarily, a rational recalibration of Nvidia’s fundamentals. It was a recognition that the moats around US AI capability were smaller than the market had assumed.

Similar moats are being priced into every US enterprise’s AI transformation thesis right now. Most of them are similarly overstated.

The world has changed. The leaders who notice will be the ones the next decade is built around.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top