This is one of the talks that has affected me most.

The sentence I kept returning to was not about models, chips, or revenue. It was about choosing the one thing that actually matters. If you care about one sufficiently important thing, many of the problems at the edges stop being important. You can give them up. You can let other people have them. You can stop exhausting yourself by trying to win everywhere.

Liang Wenfeng’s argument is that restraint is not the opposite of ambition. It is what makes extreme ambition possible.

这是一次对我触动很大的讲话。

我反复想到的,并不是其中关于模型、芯片或者收入的判断,而是一个更简单的原则:如果你只在乎一件真正重要的事,那么许多边角上的问题就不再重要。你可以让出去,可以不争,也可以不再为了处处获胜而耗尽自己。

梁文锋真正表达的是:克制并不是野心的反面。恰恰因为有极大的野心,才必须克制。

1. A Company Is Organized by Its Vision

一家公司真正靠愿景组织

中文

我们最初创办这家公司时,并没有想最后要赚多少钱、去资本市场还是上市。最开始的几十个人也完全没有这么想。如果他们只想这些,就不会来。

我们是怀着对这个世界很大的善意来做这件事。我们觉得它对人类有用,是金钱以外的事情。后来利益变得很大,当然也会出现新的诱惑。但我们出发时的愿景,以及一直保持到现在的愿景,都不是商业利益最大化。

二十年前,我在管理上最崇拜杰克·韦尔奇。现在回头看,他讲的大部分东西可能已经不对了,但有一点是对的:一家公司最重要的是愿景。

大公司不是靠规章制度管理,而是靠愿景。愿景不是墙上的标语。愿景不是你怎么说,而是你实际怎么做。

我们是怎么把这么多人组织起来的?严格来说,我们没有组织。我们靠愿景组织。没有 KPI,也没有一套写在纸上的愿景。愿景存在于我们做事的方法,以及我们对待世界的态度里。每个人的理解可能不完全相同,但大方向是一致的:怀着善意,做一点对世界有用的事情。

这是真的,不是后来编出来的故事。否则无法解释我们做过的很多选择。

English

When we started the company, we were not thinking about how much money we would make, whether we would enter the capital markets, or whether we would go public. The first few dozen people did not think that way either. If they had, they would not have joined.

We began with a great deal of goodwill toward the world. We believed this work could be useful to humanity and that it was about something beyond money. Once the possible rewards became enormous, new temptations naturally appeared. But the vision we began with, and have kept until now, was never to maximize commercial gain.

Twenty years ago, the manager I admired most was Jack Welch. Looking back, much of what he taught may no longer be right. But he was right about one thing: the most important thing in a company is its vision.

A large company is not ultimately managed by rules. It is managed by vision. Vision is not a slogan on the wall. It is not what you say. It is how the company actually behaves.

How did we organize so many people? Strictly speaking, we did not. The vision organized us. We had no KPI system and no written statement of the vision. It existed in how we worked and in our attitude toward the world. Each person may have understood it differently, but the broad direction was shared: approach the world with goodwill and try to do something useful.

This was real, not a story invented afterward. Otherwise, many of our decisions make no sense.

2. Why Open Source Requires Restraint

为什么开源的前提是克制

中文

我们坚持开源,首先是因为愿景要求我们开源。其次,我们认为,如果真想在商业上把 AI 做成,开源反而有好处。

这似乎违反直觉。过去,一个软件市场可能一年只有几十亿美元。开源以后,原本的市场可能只剩几千万或几亿美元,开源与商业化天然冲突。

AI 不一样。它可能最终占人类 GDP 的百分之十,甚至更多。这个市场大到不可能由一个人独占。你越想独占,越会遇到阻力,越可能被历史抛弃。

所以你需要一套机制,确保自己只能获得有限的利益。想把 AI 做成,首先要学会克制。不能想着人类或者中国 GDP 的多少最终都归自己。越这样想,越做不成。

我们两年前没有很多钱,没有很多卡,没有知名度,也没有号召力。我们就是一群普通人。我更喜欢“一群普通人做出不普通的事”这个叙事,而不是“一群天才做出不普通的事”。克制和这个叙事是一体的。

AI 的利益足够大。最后哪怕只分到一点,也已经足够。因此现在最重要的问题不是怎么多拿一点,而是怎么提高做成 AGI 的概率。

开源是克制的一部分。低价也是。帮助竞争对手复现模型也是。它们表面上是让利,实际上能增强内部凝聚力,获得社会支持,并让团队更从容地接近真正的目标。

English

We insist on open source for two reasons. First, the vision itself requires it. Second, we believe open source improves our chances of building a viable AI business.

This sounds counterintuitive. In the past, a software market might have been worth only a few billion dollars a year. Open sourcing the product could reduce the commercial opportunity to tens or hundreds of millions. Open source and commercialization were naturally in tension.

AI is different. It may eventually account for ten percent of human GDP, perhaps more. The opportunity is too large for one company to monopolize. The more you try to own all of it, the more resistance you create, and the more likely history is to reject you.

You therefore need a mechanism that limits how much value you can capture. If you want to make AI happen, you first need restraint. You cannot begin by deciding what percentage of the world’s GDP, or China’s GDP, should belong to you. The more you think that way, the less likely you are to succeed.

Two years ago, we did not have much money, many chips, a famous name, or unusual influence. We were a group of ordinary people. I prefer the story of ordinary people doing something extraordinary to the story of geniuses doing something extraordinary. That story and our restraint come from the same place.

The value created by AI will be so large that even a small share will be enough. The important question is not how to capture a little more today. It is how to increase the probability that we can build AGI.

Open source is one form of restraint. Low pricing is another. Helping competitors reproduce our work is another. These choices look like concessions, but they strengthen internal cohesion, earn support from society, and let the team move toward the real objective with less friction.

3. Price for a Reasonable Return, Not the Maximum Return

定价只赚合理利润,而不是最大利润

中文

我们的 API 定价逻辑,是买回一批设备后,大约十个月收回设备成本。财务上,服务器可能按三年或五年摊销;但商业上,我们认为十个月回本已经是合理利润。

这并不是利润最大化。这个价格区间的需求弹性很低。价格即使提高一倍,token 消耗量可能也不会减少多少,因此收入很可能接近翻倍。

有一次,我们担心需求太多,最初把模型价格定得较高。后来把价格降到四分之一,公司群里很多人都在欢呼。大家花了那么多精力把模型做好,本来就是希望它便宜、好用,让更多人充分使用。降价让大家觉得工作的意义得到了实现。

竞争对手的员工可能不会因为降价而欢呼,因为价格降一半,ARR 就可能掉一半。这是我们的不同。

有人说十个月回本的利润仍然很高。确实还有降价空间,模型优化也会继续降低成本。但继续降价必须带来更多需求或者更多社会价值。如果所有人已经用得起,再降价既不增加收入,也不明显增加幸福感,就没有必要为了降价而降价。

克制是一种战略。有时舍弃一点短期收入,可以换来团队凝聚力、社会支持和更大的长期机会。

English

Our API pricing rule is roughly this: after buying a batch of equipment, we should recover the equipment cost in about ten months. A server may be depreciated over three or five years in the accounts, but commercially, we consider a ten-month payback sufficient.

This is not profit maximization. Demand is relatively inelastic in this price range. Even if we doubled the price, token consumption might not fall much, so revenue could nearly double.

At one point, we were worried that demand for a model would be too high, so we launched it at a relatively high price. Later, we cut the price to one quarter of that level. Many people in the company chat cheered. They had spent so much effort making the model good because they wanted it to be cheap, useful, and widely used. The price cut made the purpose of their work visible.

Employees at a competitor might not celebrate a price cut, because halving the price can halve ARR. That is one way we are different.

Some people say a ten-month payback is still very profitable. They are right. There is more room to reduce prices, and model optimization will keep lowering costs. But another price cut should create more demand or more social value. If everyone can already afford the service, a lower price may neither increase revenue nor materially improve anyone’s life.

Restraint is a strategy. Giving up some short-term revenue can buy cohesion, trust, and a larger long-term opportunity.

4. Do Not Stop for the Sesame Seeds

不要为了芝麻停下来

中文

去年春节,用户突然增加。我们没有想过怎么锁住这些用户、怎么变现,也没有想过做下一个超级 App、字节或者腾讯。

我们本来可以花很多钱去争 C 端流量,但选择了更克制的做法。因为后面可能有西瓜,前面的只是芝麻。现在看,去年没有全力投入 C 端可能是对的。我们用很低的成本维持服务,但没有停下来把它当作最重要的事情。

今年 API 和 B 端收入也可能达到很大规模。转写稿中的判断是,如果需求继续扩大、能获得更多显卡,ARR 可能达到数亿美元;如果 AI 收入达到十亿美元,公司现金流可能覆盖研发和其他费用。

这些事情值得做,也应该做好。但它们不是第一优先级。与 AGI 相比,它们不是同等重要的问题。

我们不是为了 C 端或者 B 端而做模型。我们为了 AGI 做模型,而通往 AGI 的每一个台阶恰好会产生 C 端产品、API 和 B 端服务。于是我们把这些副产品拿来商业化。

这是一种战略优势。站在更高的技术目标上,做低一级的产品,只需要很少的额外精力。去年我们几乎没有抢 C 端用户,但用户赶也赶不走。今年 B 端收入的增长也比较乐观,而我们没有销售,几乎没有客服,只用少数人维护 API。

有时你越想得到一个东西,越得不到;没有把它当成唯一目标,反而更容易得到。

English

During last year’s Spring Festival, the number of users suddenly surged. We did not think about locking them in, monetizing them, or becoming the next super app, ByteDance, or Tencent.

We could have spent heavily to fight for consumer traffic. Instead, we chose restraint. There may be watermelons later; the opportunities in front of us may be only sesame seeds. Looking back, not going all in on the consumer market was probably right. We kept the service running at low cost, but we did not stop and turn it into the company’s main objective.

This year, API and enterprise revenue may also become substantial. According to the transcript, if demand keeps growing and more GPUs become available, ARR could reach several hundred million dollars. At one billion dollars in AI revenue, operating cash flow might cover research and all other expenses.

These opportunities are worth taking, and the work should be done well. But they are not the first priority. Compared with AGI, they are not equally important problems.

We do not build models in order to serve consumers or enterprises. We build them to pursue AGI. Each step toward AGI happens to produce consumer products, APIs, and enterprise services, so we commercialize those by-products.

That creates a strategic advantage. If your real objective sits at a higher technical level, lower-level products require little extra effort. Last year we barely fought for consumer users, yet they would not leave. This year enterprise revenue is growing optimistically, even though we have no sales team, almost no customer service, and only a few people maintaining the API.

Sometimes the thing you want most becomes harder to get. The thing you do not treat as the only goal can arrive almost by itself.

5. Open Source Does Not Destroy the Business

开源并不会摧毁商业模式

中文

我们会继续开源,最强的模型也可能开源。因为我看不到闭源的必然好处。

即使把模型和原理都公开,真正用起来仍然有很高门槛。第三方不仅要部署成功,还要把成本做到很低。这需要大量工程和组织工作,不是知道原理就能完成。

创业公司太小,可能没有力量做这件事;大公司又可能很难组织。我们现在的规模处在一个甜蜜点。

如果我们按照十个月回本来定价,独立部署的第三方很难获得利润,因为它通常做不到我们的成本。因此开源不会影响收入。只有当你想赚一百倍利润时,开源才会阻碍你,因为成本更高的第三方也能用更低价格竞争。

我们甚至希望更多人部署我们的模型,会主动帮助开源社区。我们担心的不是别人抢走生意,而是他们部署不好、效果变差或者成本过高。

开源模型和我们自己部署的模型是同一个,不会公开较差版本、内部保留更好版本。过去一年,开源没有损害我们的 C 端服务,也没有削弱用户基础。

在我们的愿景下,开源与商业付费可以长期共存。

English

We intend to keep open sourcing our work, including perhaps our strongest models, because I do not see an inevitable advantage in keeping them closed.

Even if the model and its principles are public, using it well remains difficult. A third party must deploy it correctly and operate it at low cost. That requires extensive engineering and organizational work. Understanding the idea is not enough.

A startup can be too small to do this work. A large company can be too difficult to organize. Our current size sits in a kind of sweet spot.

If we price for a ten-month payback, an independent host will struggle to make money because it usually cannot match our costs. Open source therefore does not eliminate our revenue. It becomes a problem only if you want one-hundredfold margins, because even a less efficient third party can undercut you.

We want more people to deploy our models and will help the open-source community do it. Our worry is not that they will steal the business. It is that they will deploy the model badly, reduce its quality, or operate it at excessive cost.

The model we release is the same model we deploy ourselves. We do not open source an inferior version while keeping a better one inside. Over the past year, open source did not damage our consumer service or weaken the user base.

Under our vision, open source and paid services can coexist for a long time.

6. The Road to AGI Is a Staircase

通往 AGI 的路是一段一段台阶

中文

公司的长期目标是 AGI。每个人对 AGI 的定义可能不同,但不妨碍我们把它当作目标。

当前一代 AI 在一个前提下已经能超过人类:你必须把问题描述清楚,并提供完整的上下文和指令。真正困难的是,人类几乎从来不靠完整说明工作。

一个新员工可能用两个月熟悉公司。之后你只要说“叫小王过来”,他就知道小王是谁、在哪里、负责什么。AI 没有这两个月的经历。你必须重新告诉它全部背景。因此它能完成孤立任务,却还不能真正替代员工。

缺少的能力是持续学习。

AI 的发展像一段楼梯。语言模型是前面的台阶。去年跨过的是思维链,即 CoT:模型通过自己的推理把能力上限推高。今年的台阶是 Agent:模型借助工具和行动完成更大范围的任务。Agent 又建立在 CoT 和语言模型之上,每一步都没有白走。

Agent 也会达到上限。它能解决许多任务,却仍然无法像员工一样进入一个环境、持续学习并积累上下文。下一个明确的瓶颈是持续学习。

当模型能够持续学习,它就可能完成人类能做的大部分事情,并参与开发自己的下一代版本。所谓“奇点”可能并不是一个突然的临界点,而是一段连续但非线性的加速过程:AI 开始加速 AI 研究。

我们的理想顺序是:先解决持续学习,再让模型参与自我迭代,最后进入具身智能。这个顺序最省力,因为每一步都能帮助完成下一步。等到自我迭代能力足够强,具身智能和下一代机器人就不必完全靠人手开发。

如果顺序反过来,先做具身智能,会是极其辛苦、工程密集的路线。我们希望走轻松一点的路。

English

The company’s long-term objective is AGI. People may define AGI differently, but that does not prevent us from using it as the goal.

The current generation of AI can already outperform humans under one condition: the problem must be clearly described, with complete context and instructions. The difficulty is that human work almost never comes with complete instructions.

A new employee may spend two months learning the company. After that, you can say, “Ask Xiao Wang to come here,” and the employee knows who Xiao Wang is, where he sits, and what he does. AI does not have those two months of experience. You must explain the entire background again. It can complete isolated tasks, but it cannot yet replace an employee.

The missing ability is continuous learning.

AI develops like a staircase. Language models formed the earlier steps. Last year’s step was chain-of-thought reasoning, or CoT: allowing the model to reason pushed its ceiling higher. This year’s step is the agent: tools and actions let the model complete a wider range of tasks. Agents depend on CoT, which depends on language models. None of the steps is wasted.

Agents will also reach a limit. They may solve many tasks but still cannot enter an environment, keep learning, and accumulate context like an employee. The next visible bottleneck is continuous learning.

Once a model can learn continuously, it may perform most of the work humans can do and help develop its own next version. The so-called singularity may not be a sudden threshold. It may be a continuous but nonlinear acceleration in which AI begins to speed up AI research.

Our preferred sequence is continuous learning first, self-improvement second, and embodied intelligence last. This is the easiest route because each step helps build the next one. Once self-improvement is strong enough, embodied systems and future robots will no longer need to be designed entirely by humans.

Reversing the order and starting with embodiment would be brutally engineering-intensive. We would rather take the easier path.

7. The Only Core Interest Is Team Stability

唯一不能退让的核心利益是团队稳定

中文

我们可以在很多事情上克制,但必须知道什么不能让。

我们的核心利益只有一个:保持团队稳定。这甚至可以看作唯一的核心利益。

只要团队不散,我们就一定能继续接近 AGI。可能早半年,也可能晚一年;可能遇到挫折,但只要大家留下来,就能重新开始。钱、资源和其他要素最终都可以获得,团队一旦失去则很难重建。

最近一次融资在一定程度上降低了这个风险。老员工和关键员工持有的期权较多。只要他们稳定,其他人通常也不会轻易离开。大家并不完全为了钱而来,他们希望留在一个真正有机会做成 AGI 的环境里。

历史上,我们的人才流动一直比同行少,但团队稳定依然是最大的风险。公司做的很多事,最终都可以理解为在保护这一点。

除了团队稳定,其他东西都可以不要。我们不愿意把任何互联网大厂或创业公司当作必须打败的敌人。我们愿意帮助阿里、智谱、月之暗面以及其他竞争者做得更好,因为在核心利益上并不冲突。

当你只明确保护一件事,其他谈判会变得简单。你不需要在每条边界上都获胜。

English

We can practice restraint in many areas, but we must know what cannot be surrendered.

We have only one core interest: the stability of the team. It may be the company’s only true core interest.

As long as the team stays together, we can keep moving toward AGI. We may arrive six months earlier or a year later. We may suffer setbacks. But if everyone remains, we can begin again. Money, compute, and other resources can eventually be obtained. A lost team is much harder to rebuild.

The latest financing reduced this risk to some degree. Long-serving and critical employees hold meaningful options. If they remain stable, others are less likely to leave. People did not join only for money. They want to work in an environment that genuinely has a chance of building AGI.

Historically, our employee turnover has been lower than that of our peers. Even so, team stability remains the largest risk. Many of the company’s decisions can ultimately be understood as attempts to protect it.

Everything except team stability is negotiable. We do not want to treat any large internet company or startup as an enemy that must be defeated. We are willing to help Alibaba, Zhipu, Moonshot, and other competitors improve because our core interests do not conflict.

Once you know the one thing you must protect, every other negotiation becomes simpler. You no longer need to win at every edge.

8. Focus Means Refusing Good Businesses

聚焦意味着拒绝好生意

中文

我们只做通往 AGI 的主线:语言模型、CoT、Agent、持续学习。

AI 领域很广。3D、视频生成、世界模型、多模态、搜索都可能很重要,也可能是好生意。但“重要”不等于“现在是智能主线上的瓶颈”。

视频生成在 Sora 出现后非常热门,几乎所有公司都做。后来许多小公司又砍掉了。它在商业上可能很好,但与当前智能上限没有直接关系。我们不会因为它是好生意就做,只会因为它位于智能路线图上才投入。

多模态最终一定要做,对 C 端产品也很重要。V4 及后续版本会支持原生多模态。但在我们的理解里,多模态和搜索一样,是组件,不是智能主线本身。

世界模型的定义更模糊。我们的判断是,当前最重要的仍然是把训练做好,然后解决持续学习,再让模型自己提出问题。其他公司可以选择不同路线,没有绝对对错。

我们愿意帮助别人把这些技术带进生产环境,提高社会效率,但不需要所有事情都由自己完成。AI 足够大,会产生许多万亿美元级别的公司。我们只需要做好其中一块。

English

We work only on the main path toward AGI: language models, CoT, agents, and continuous learning.

AI is a broad field. 3D, video generation, world models, multimodality, and search may all be important and may all become good businesses. But “important” does not mean “the current bottleneck on the path of intelligence.”

After Sora appeared, video generation became fashionable and nearly every company started working on it. Many smaller companies later shut those projects down. Video may be commercially valuable, but it does not directly raise the current ceiling of intelligence. We will not pursue it merely because it is a good business. We will pursue something when it lies on the roadmap of intelligence.

Multimodality will eventually be necessary and is already important for consumer products. V4 and later versions are expected to support native multimodality. But in our view, multimodality is like search: a component, not the main line of intelligence itself.

“World model” is a less precise term. Our judgment is that the current priority remains better training, followed by continuous learning and then the ability of the model to ask its own questions. Other companies may choose different roadmaps. There is no universal answer.

We want to help others bring AI into production and improve social productivity, but we do not need to build everything ourselves. AI is large enough to create many trillion-dollar companies. We need to do only one part well.

9. Compute Is the Gap

真正的差距是算力

中文

我们与美国最大的差距不是人才,而是资源。

转写稿中提到,公司当时大约拥有两万张 H 等效算力,其中许多设备刚刚到货,未来几个月还会继续扩充。策略很简单:只要价格合理,能买多少卡就买多少。融资所得如果能在半年内全部变成显卡,反而是最理想的情况。卡能够产生服务和现金流,钱放在银行里只能获得很低利息。

人才不是瓶颈。中国和美国使用的是相近的人才池,中国每年也有大量新人进入行业。人才差距本质上仍然来自算力差距:卡少,能做的实验少,研究人员成长得也慢。

在最大的模型上,国内存在数量级差距。转写稿中的估算是,美国最前沿模型可能达到约 800B 激活参数,而国内仍主要在几十 B。训练同等规模模型可能需要五万张 GB300,或者约二十万张华为 950,而且这还没有算研究实验所需资源。

因此现阶段不应在完全相同的规模上硬拼。先在几十 B 激活参数的范围内把研究做充分,之后随着资源增加,再扩展到 150B、250B。

Scaling 还远没有在中国到头。模型、数据和训练成本继续扩大,收益仍然明显。硅谷讨论 Scaling 是否触顶,是因为他们已经走得很远;中国甚至还没有足够算力碰到那堵墙。

在算力总量有数量级差距时,全面超越不现实。但可以落后一到两年,用对方二十分之一的算力复现结果,再逐步把时间差缩短到六个月、三个月,并在少数重点方向超越。

English

The largest gap between us and the United States is not talent. It is resources.

According to the transcript, the company then had roughly twenty thousand H-equivalent GPUs, many of which had only recently arrived, with further expansion planned. The strategy was simple: buy as many GPUs as possible at a reasonable price. If the financing proceeds could all be converted into GPUs within six months, that would be an ideal outcome. GPUs can produce services and cash flow; money in the bank earns little.

Talent is not the bottleneck. China and the United States draw from similar talent pools, and China adds a large number of new people every year. Even the talent gap is largely a compute gap: with fewer chips, researchers can run fewer experiments and develop more slowly.

At the largest model scale, China faces an order-of-magnitude difference. The transcript estimated that frontier American models might reach about 800B active parameters, while Chinese models were still mostly in the tens of billions. Training at the same scale might require fifty thousand GB300s or roughly two hundred thousand Huawei 950s, before counting the compute needed for research experiments.

The rational response is not to compete at exactly the same scale today. First, do the research thoroughly in the tens-of-billions range. Then move toward 150B or 250B as resources expand.

China is nowhere near the end of scaling. Increasing model size, data, and training expenditure continues to produce visible returns. Silicon Valley can debate whether scaling has reached a wall because it has traveled much farther. China does not yet have enough compute to touch that wall.

With an order-of-magnitude compute gap, comprehensive superiority is unrealistic. But China can remain one or two years behind, reproduce the result with one-twentieth of the compute, then shrink the time gap to six months or three months and surpass the frontier in selected areas.

10. Cost and Time Are the Durable Advantages

最后真正留下来的差异是成本和时间

中文

大模型竞争到终局,模型之间可能不会有本质上的巨大差距。最后的差异主要有三个:成本、时间和用户体验。

成本排在第一。同样质量的服务,谁能以更低成本提供,谁就拥有真正的壁垒。时间排在第二。早几个月和晚几个月不是同一件事。用户体验也会产生粘性,但可能没有前两项本质。

比较模型效果,必须在相同成本下比较,就像比较同价位的汽车。一个模型好不好,不由某一个环节决定,而是整体系统决定。

Anthropic 当时在 Code Agent 上领先 OpenAI,但这个优势未必长期持续。OpenAI、Google 和 Anthropic 都很强,未来可能交替上升。先发优势会消失,效率会变得更重要。

中国公司的结构性优势可能就在成本和产品体验。美国公司资源充足,不必把低成本当成第一优先级;中国公司必须提高效率。中国也擅长把产品做得足够好、足够便宜。

在全球 AI 分工中,中国可能像其他制造业一样,成为产量最大的参与者之一,用系统性低成本提供与美国差距不大的产品。

English

At the end of the large-model competition, the models themselves may not differ dramatically. Three differences are likely to remain: cost, time, and user experience.

Cost comes first. If two companies provide the same quality of service, the one that can provide it more cheaply has a real moat. Time comes second. Arriving several months early is not the same as arriving several months late. User experience can create loyalty, but it may be less fundamental than cost and timing.

Models should be compared at the same cost, just as cars should be compared within the same price range. A strong model is not the result of one isolated component. It is the result of the whole system.

Anthropic was ahead of OpenAI in coding agents at the time of the conversation, but that lead might not persist. OpenAI, Google, and Anthropic are all strong and may alternate at the frontier. Early advantages fade; efficiency becomes more important.

China’s structural strengths may lie in cost and product experience. American companies have enough resources that low cost does not need to be the first priority. Chinese companies are forced to become efficient. China is also good at making products that are good enough and systematically cheaper.

In the global AI division of labor, China may play the role it has played in other industries: becoming one of the largest producers and offering products close to American quality at much lower cost.

11. Domestic Chips Have a Historical Opening

国产芯片迎来一个历史性窗口

中文

过去,国产 AI 芯片最大的问题是生态。卡买回来却不好用,软件栈远不如 CUDA。

这个问题正在变化。首先,AI 可以写代码,建立软件生态比过去容易。其次,TileLang 这样的高级语言可以用更少代码重写算子。再加上 AI 辅助,重新建立一套生态不再像过去那样困难。

CUDA 从游戏卡发展而来,长期兼容游戏卡的设计逻辑。过去 AI 计算市场较小,这很合理。现在计算卡市场已经大于游戏卡,专用芯片没有必要继续与旧架构耦合。无论华为还是英伟达,未来都会更多转向专用芯片。

V3 训练时使用英伟达卡,但已经尽量不依赖英伟达生态,而是通过 TileLang 构建自己的高级编译层。把这一套迁移到华为卡上,生态问题就可能被重新定义。

转写稿提到,DeepSeek 当时主要与华为合作,参与华为生态适配。华为给出的 950 配额约为一万六千张,而互联网大厂可能拿到十几万张。这个数量不足以训练更大一代模型,但足以验证生态。

梁文锋的判断是:国产芯片的生态问题有机会在一年内被事实扭转,真正剩下的瓶颈将是产能。他将中美芯片差距概括为“四倍加两年”:大约四张华为卡对应一张英伟达卡,同时产品代际落后约两年。

他对长期产能仍然乐观。未来几年可能持续受限,但很难相信五年后中国仍然卡在同一个产能问题上。

English

Historically, the biggest problem with Chinese AI chips was the ecosystem. Companies could buy the hardware but struggled to use it because the software stack was far weaker than CUDA.

That is changing. First, AI can write code, making it easier to build an ecosystem. Second, high-level languages such as TileLang allow kernels to be rewritten with much less code. Combined with AI assistance, rebuilding the stack is no longer as difficult as it once was.

CUDA evolved from gaming GPUs and inherited their design constraints. That made sense when AI computing was a smaller market than gaming. Now compute accelerators are becoming the larger market, and specialized chips no longer need to remain coupled to the old architecture. Both Huawei and Nvidia are likely to move further toward dedicated designs.

V3 was trained on Nvidia hardware but attempted to minimize dependence on Nvidia’s ecosystem by building a higher-level compiler layer with TileLang. Porting that layer to Huawei hardware could redefine the ecosystem problem.

The transcript states that DeepSeek was working mainly with Huawei on adaptation. Huawei’s allocation of 950-series chips to DeepSeek was said to be about sixteen thousand units, while large internet companies might receive more than one hundred thousand. That quantity was not enough to train a much larger next-generation model, but it was enough to validate the ecosystem.

Liang’s judgment was that the perception of an unusable domestic ecosystem could be overturned by evidence within a year. The remaining bottleneck would then be production capacity. He summarized the chip gap as “four times plus two years”: roughly four Huawei chips for one Nvidia chip, with the domestic product generation about two years behind.

He remained optimistic about long-term capacity. The constraint may persist for several years, but it is difficult to believe China will still face the same production bottleneck five years from now.

12. TileLang, CUDA, and Vertical Integration

TileLang、CUDA 与垂直整合

中文

TileLang 并不是为了牺牲效率换取国产化。梁文锋的判断相反:它会大幅提高开发效率。

过去离不开 CUDA 生态,现在可以用更简单的高级语言重写大量工作,代码量显著减少。当前 TileLang 仍主要由人编写,但未来可以用 AI 来写。底层执行效率可能损失百分之一到百分之二,他认为这是可以接受的。

这不只是弥补短板,而是技术发展带来的机会。

公司会继续自建大型集群,因为一直以来集群就是自己建设的。但是否自研芯片,要看收益。如果能以合理价格买到芯片,就没有必要自己造。运营发电厂的人不一定需要制造发电机。

DeepSeek 希望只做最擅长、最核心的一块,不希望向上游和下游吞下所有业务。AI 市场足够大,只做一小块也可能成为万亿美元级别公司。

To B 的上限最终由需求决定,而不是算力。当前技术能够创造多少真正有价值的任务,决定了企业服务收入的规模。随着技术突破,需求会继续扩大。

English

TileLang is not an attempt to trade efficiency for domestic compatibility. Liang argued the opposite: it can greatly improve development efficiency.

Work that once depended on the CUDA ecosystem can be rewritten in a simpler high-level language with much less code. TileLang was still largely written by humans at the time, but AI could eventually write it. A one-to-two-percent loss in low-level execution efficiency was considered acceptable.

This is not merely a way to compensate for a weakness. It is an opportunity created by technical change.

The company will continue building its own large clusters, as it has always done. But whether it should design its own chips depends on the return. If chips can be bought at a reasonable price, there is no need to manufacture them. An operator of a power plant does not necessarily need to build generators.

DeepSeek wants to do only the part it considers most central and most suited to the team. It does not want to absorb every upstream and downstream business. AI is large enough that a company can remain focused on one piece and still become a trillion-dollar company.

The ceiling of the enterprise business will ultimately be determined by demand, not compute. The amount of genuinely useful work the current technology can perform sets the size of enterprise revenue. As the technology improves, demand will expand.

13. A Research Organization Needs Slack

研究组织必须保持松弛

中文

公司的管理有两条线。

第一条从上到下。当公司需要发布 V4 或完成一个集体目标时,就进行正式分工,每个人负责一部分。

第二条从下到上。每个人可以决定自己想研究什么,没有 KPI,也不需要事事协调。只要公司有能力支持,研究者就可以按照自己对重要性的判断去探索。

原则上,正式任务最好不要占员工一半以上的时间。至少一半时间不被安排,用于自由研究。

公司一般不太加班。第一,研究需要松弛。一个人必须对问题产生自己的兴趣,并在平时持续思考。逼得太紧,就不可能真正探索。

第二,公司足够聚焦,因此要做的事情很少。克制意味着主动不做许多事情。任务少,每个人分到的工作也少,所以不需要靠加班填满时间。

产品里有很多不完善的部分,公司也没有急着补齐。这不是因为不知道,而是因为补齐它们不是当前最重要的事。

重大决策也不是梁文锋一个人推动。公司建立在共识之上。他的权威来自共识,而不是职位。一个决定只有成为团队共识,才真正推得下去。

English

The company is managed through two lines.

The first runs from the top down. When the company needs to release V4 or complete a collective objective, work is formally divided and each person takes responsibility for one part.

The second runs from the bottom up. People choose what they want to research. There are no KPIs, and not every experiment needs coordination. If the company can support it, researchers can explore whatever they believe matters.

In principle, formal assignments should consume no more than half of an employee’s time. At least half should remain unassigned and available for independent research.

The company generally does not work much overtime. First, research needs slack. People must develop their own interest in a problem and keep thinking about it outside a scheduled task. If the environment is too tight, genuine exploration becomes impossible.

Second, the company is highly focused, so there is less work to do. Restraint means deliberately refusing many tasks. When the company does fewer things, each person receives less assigned work, and overtime becomes unnecessary.

Many parts of the products remain unfinished. The company does not rush to complete them. This is not because the problems are invisible. It is because solving them is not currently the most important work.

Major decisions are not simply imposed by Liang. The company operates through consensus. His authority comes from that consensus, not from his title. A decision becomes executable only when the team broadly believes in it.

14. Q&A: Talent, Ecosystems, and Industry Consolidation

问答:人才、生态与行业收敛

中文

问:开源之后,行业是否有足够的人才复现模型、发展垂直模型和应用?未来生态会怎样形成?

答: 每个新行业早期都会缺人才,但这种短缺通常很快消失。网站开发、服务端开发,甚至过去被认为培养周期很长的飞行员,都没有永远短缺。AI 人才短缺也是阶段性的,两三年就会培养出大量人才,而且现在已经明显缓解。

真正的问题是国内做基础模型的公司太多。美国可能集中在三家,中国的资源却分散在许多公司,每家都重复做相近工作。最终一定会收敛。基础模型不需要那么多家,三四家充分竞争已经足够。

当大家意识到基础模型不会拥有异常高的利润率,一些公司自然会退出。任何长期高得不合理的利润都不符合客观规律。最终每家公司只能获得合理利润,成本控制好的人多赚一点,控制差的人少赚一点。

DeepSeek 愿意帮助生态伙伴,但精力有限。重要的是不存在根本利益冲突。大模型公司不可能拿走产业的大部分利润,因为长期差异主要只是时间和成本。

English

Question: After open sourcing the models, will the industry have enough talent to reproduce them, build vertical models, and develop applications? What will the ecosystem look like?

Answer: Every new industry begins with a talent shortage, and the shortage usually disappears quickly. Web developers, backend engineers, and even pilots were once considered scarce. None remained permanently scarce. The shortage of AI talent is also temporary. Large numbers of people can be trained within two or three years, and the constraint is already easing.

The larger problem is that China has too many foundation-model companies. The United States may concentrate resources in three companies, while China spreads them across many teams doing similar work. The field will consolidate. It does not need dozens of foundation-model providers. Three or four strong competitors are enough to create intense competition.

As companies discover that foundation models do not produce abnormally high margins, some will stop. Persistently excessive profits would violate the underlying economics. Each company will eventually earn a reasonable return. The efficient ones will earn somewhat more; the inefficient ones somewhat less.

DeepSeek wants to help ecosystem partners but has limited attention. The important point is that there is no fundamental conflict of interest. Foundation-model companies cannot capture most of the value in the industry because the durable differences between them will mainly be time and cost.

15. Q&A: Data, Synthetic Data, and Surpassing Humans

问答:数据、仿真数据与超越人类

中文

问:模型如果依赖人类产生的真实数据,智能是否会受限于人类过去的知识?仿真数据和模型自我生成的数据能否突破这个上限?

答: AI 可以在一定范围内超越人类。AlphaGo 下出过人类从未见过的棋。这说明模型可以从人类已有知识出发,产生人类过去没有明确表达过的结果。

它仍然可能存在上限和局限,但我们现在看不到那个上限。真实数据、仿真数据和生成数据都可能成为路径,不会只有一种方法。

数据几乎可以看作模型的一半。当前阶段,许多能力的提升仍然依赖高质量数据,尤其是后训练。

高端数据标注在中国没有明显成本优势。标注专家级数据,无论在中国还是美国都很贵。转写稿中提到,公司有接近一半的核心研究人员参与数据相关工作,先做成本较低、能够快速扩展的部分。

融资后会继续加大后训练投入,但真正瓶颈未必是资本,而是时间。OpenAI 和 Anthropic 开始得更早。国内高质量数据工作主要在近半年加速,需要积累过程。梁文锋判断,一年内这个问题有望得到明显改善。

English

Question: If models depend on real data produced by humans, will intelligence remain limited by humanity’s past knowledge? Can simulated or self-generated data break that ceiling?

Answer: AI can surpass humans within a defined domain. AlphaGo played moves no human had seen before. This suggests that a model can begin with human knowledge and produce results humans have never explicitly expressed.

There may still be a ceiling, but we cannot see it yet. Real data, simulated data, and generated data may all become useful. There will not be only one method.

Data is almost half of the model. At the current stage, many gains still depend on high-quality data, particularly in post-training.

China has no obvious cost advantage in expert data labeling. High-end labeling is expensive in both China and the United States. The transcript says that close to half of the company’s core researchers were involved in data work, beginning with the portions that were cheaper and easier to scale.

Financing will allow more investment in post-training, but the true bottleneck may be time rather than capital. OpenAI and Anthropic started earlier. China accelerated its high-quality data work only recently and must accumulate experience. Liang expected the gap to improve materially within a year.

16. Q&A: Hardware Depreciation and the American Lead

问答:硬件折旧与美国领先

中文

问:今天建设的万卡集群,两三年后会不会变成落后资产?模型变得更聪明、硬件变得更高效,这两条曲线如何相遇?

答: 英伟达卡大致可以按五年折旧,华为卡最多按三年。华为 950 今年和明年仍然有价值,再往后可能因为能耗而变得不经济。它的生命周期较短,是因为产品代际本来就落后约两年。

只要价格合理,今天能买到 B200 仍然划算。真正的问题不是要不要买,而是买不到足够数量。

中国用三种方式消化算力差距。第一,承受模型规模更小、时间上更晚。第二,用更高算法效率完成相同任务。第三,集中资源,在少数关键方向做到更好。

可以把当前叙事概括为:落后一到两年,但只用美国二十分之一的算力做出结果。下一步是用几分之一的算力,把落后缩短到六个月或三个月。

English

Question: Will clusters built today become obsolete in two or three years? How will better models and better hardware meet each other?

Answer: Nvidia chips can roughly be depreciated over five years. Huawei chips may deserve no more than three. A Huawei 950 can remain useful this year and next year, but later its power consumption may make it uneconomic. Its useful life is shorter partly because the generation is already about two years behind.

At a reasonable price, buying every available B200 still makes sense. The real problem is not whether to buy, but whether enough units can be obtained.

China absorbs the compute gap in three ways. First, it accepts smaller models and later delivery. Second, it uses more efficient algorithms to perform the same work. Third, it concentrates resources and outperforms in selected areas.

The current story can be summarized as follows: reproduce the result one or two years later with one-twentieth of the American compute. The next goal is to use a fraction of the compute while reducing the delay to six months or three months.

17. Q&A: Continuous Learning and Recursive Improvement

问答:持续学习与递归改进

中文

问:持续学习什么时候会突破?最大的技术难点是什么?

答: 全世界都还没有找到真正有效的方法。持续学习不是一项单独技术,而是一个问题,可能需要许多工程和算法共同解决。现在有很多有希望的思路,但没有一条完全走通。

对投资人来说,最容易看到的是 Agent;对研究人员来说,更重要的问题已经变成学习本身。

公司内部有一个重要目标:下一版模型首先要帮助 DeepSeek 开发再下一版模型。第一目标不一定是让所有用户觉得好用,而是让研究团队自己觉得好用。只要模型能提高内部研发效率,就会缩短实现 AGI 的时间。

当前 Agent 的主要限制是不能有效持续学习。一旦解决,AI 对研究的帮助会大幅增加,通用智能的其他问题也会变得容易。否则,完全靠人继续推进 AGI,会是一条数据密集、人力密集而且很辛苦的路。

持续学习研究本身不一定消耗大量卡和固定人力。它更像“摸奖”:门槛不高,许多人都可以思考和尝试,但谁能找到有效方法并不确定。公司与其他公司的区别,是愿意把它当作重要问题,持续讨论并给研究者时间。

English

Question: When will continuous learning break through, and what is the largest technical obstacle?

Answer: No one in the world has found a truly effective method yet. Continuous learning is not one isolated technology. It is a problem that may require many engineering and algorithmic components. There are promising ideas, but no complete solution.

Investors see agents because they are visible. Researchers increasingly see learning itself as the more important problem.

The company has an internal objective: the next model should first help DeepSeek build the model after it. The first goal is not necessarily that every user finds it useful. It is that the research team finds it useful. If the model improves internal research productivity, it shortens the path to AGI.

The main limitation of today’s agents is their inability to learn continuously and effectively. Once that is solved, AI’s contribution to research will increase dramatically, making the remaining AGI problems easier. Without it, humans must continue pushing AGI through a data-intensive, labor-intensive, and exhausting process.

Research on continuous learning does not necessarily consume vast amounts of compute or fixed staffing. It resembles searching for a winning ticket: the barrier to trying is low, many people can think about it, but no one knows who will find the working method. What distinguishes the company is its willingness to treat the question as important and give researchers time to keep returning to it.

18. Q&A: When Does AGI Arrive?

问答:AGI 什么时候到来

中文

问:AGI 还需要多久?国产硬件能否在这段时间追上?

答: 在当前范式下,中国模型一两年内可以接近国外,甚至今年就可能出现能平替国外模型的产品。但这仍然不是 AGI。至少,AGI 应该具备持续学习能力。

国产硬件可能需要几年。首先解决生态信心,再解决产能。未来一两年仍然会被产能限制,但很难相信五年后仍然无法改善。

AGI 的到来不是一个明确临界点,而是渐进但非线性的过程。AI 会加速 AI 研究,所以后期进展不会保持线性。

语言模型 Scaling 目前还没有显示出明确上限。最终仍然需要具身智能,因为普通人的真实需求是衣食住行和现实世界中的劳动,而不是只在电脑里完成任务。

在没有具身之前,一个有意义的 AGI 标准是:它能够帮助开发下一版模型。进入具身以后,对应的标准则是:它能够帮助开发下一代机器人。

English

Question: How long until AGI, and can Chinese hardware catch up within that period?

Answer: Under the current paradigm, Chinese models may approach foreign models within one or two years, perhaps even produce substitutes this year. But that is still not AGI. At minimum, AGI should be able to learn continuously.

Domestic hardware may need several years. The ecosystem-confidence problem must be solved first, followed by production capacity. Capacity will remain a constraint over the next few years, but it is difficult to believe that no progress will be made within five.

AGI will not arrive at one clean threshold. It will be gradual but nonlinear. AI will accelerate AI research, so later progress will not remain linear.

Language-model scaling has not yet shown a clear upper bound. Embodied intelligence will ultimately be necessary because ordinary human needs exist in food, housing, movement, care, and physical work, not only inside computers.

Before embodiment, a meaningful standard for AGI is that it can help develop the next model. After embodiment, the corresponding standard is that it can help develop the next generation of robots.

19. Q&A: Taste, Hallucinations, and Coding Agents

问答:品位、幻觉与 Coding Agent

中文

问:如果 AI 能够自我进化,研究中的品位、直觉和 taste 还重要吗?

答: AI 现在不缺品位和直觉,缺的是持续学习。让它写文章,它已经能表现出相当的品位。真正限制它的是无法长期积累经验。

问:大模型幻觉影响用户体验,应该怎样解决?

答: 幻觉可以通过更好的后训练逐步改善,是一个能够解决的长期问题。但从公司的当前优先级看,它更接近产品问题,不是智能主线最关键的瓶颈。会解决,但不是重点。

问:应该优先做垂直行业 Agent,还是通用 Agent?

答: 现阶段最合理的路径是先全力做通用 Agent,尤其是 Coding Agent。Coding 能完成大量任务,也能直接提高 AI 研发效率。金融、法律、医疗等垂直 Agent 的优先级可以稍后。

国内未来的商业模式可能与美国不同,现在还很难提前判断。技术变化太快,过早规划产品线和商业化路径,往往是在为一个很快消失的世界做计划。

English

Question: If AI can improve itself, will taste and intuition still matter in research?

Answer: AI does not currently lack taste or intuition. It lacks continuous learning. Ask it to write an essay and it can already show considerable judgment. What limits it is the inability to accumulate experience over long periods.

Question: Hallucinations damage the user experience. How should they be solved?

Answer: Hallucinations can be improved through better post-training. It is a long-term but solvable problem. From the company’s current perspective, however, it is closer to a product problem than the central bottleneck of intelligence. It should be addressed, but it is not the highest priority.

Question: Should the company prioritize vertical agents or a general agent?

Answer: The most rational path today is to focus first on general agents, especially coding agents. Coding covers a large range of tasks and directly improves AI research productivity. Financial, legal, medical, and other vertical agents can come later.

China’s eventual business model may differ from America’s, and it is too early to predict. The technology changes so quickly that detailed product and commercialization plans often optimize for a world that disappears before the plan is executed.

20. Q&A: Research, Commercialization, and the Capital Market

问答:科研、商业化与资本市场

中文

问:如何平衡纯粹科研、AGI 愿景、商业化和资本市场责任?

答: 公司一直在商业化,只是没有以商业化为最终目标。C 端用户和 B 端收入已经构成商业基础。如果 B 端需求继续扩大,公司可能接近正现金流甚至净利润。

即使技术从此冻结,最坏情况也可以全力销售 API、把服务做好,支撑一家上市公司。公司有保底的商业业绩,同时仍然可以追求更大的目标。

DeepSeek 不是 Bell Labs。Bell Labs 可以明确不以商业化生存为前提,而 DeepSeek 本质上仍然是一家公司,政府不会无条件提供资金。公司必须活下去,也必须决定赚什么钱、什么时候赚钱、赚多少。

历史上有许多伟大公司拥有利润以外的追求。这些追求并没有阻碍商业化,反而帮助它们建立更强的商业基础。

利润以外的目标和公司身份并不矛盾。真正的问题是取舍。

English

Question: How can pure research and the pursuit of AGI be balanced with commercialization and responsibility to the capital markets?

Answer: The company has always commercialized its work. Commercialization simply was not the final objective. Consumer users and enterprise revenue already form a commercial base. If enterprise demand keeps expanding, the company may approach positive cash flow or even profitability.

Even if technical progress froze today, the fallback would be to sell APIs aggressively and operate the service well. That alone might support a public company. The company has a defensible commercial floor while still pursuing a much larger objective.

DeepSeek is not Bell Labs. Bell Labs could exist without commercial survival as its first constraint. DeepSeek remains a company. The government will not fund it unconditionally. It must survive and must decide what money to earn, when to earn it, and how much to take.

Many great companies in history pursued something beyond profit. That pursuit did not prevent commercialization. It often gave them a stronger commercial foundation.

A goal beyond profit does not contradict being a company. The real question is what the company chooses not to do.

21. Q&A: What Should the Organization Become?

问答:组织最终会变成什么

中文

问:DeepSeek 的组织方式是否有历史上的模仿对象?未来人员扩大后会怎样变化?

答: 没有模仿对象。每一步都从现实情况出发,分析利弊后选择。这个组织是时代和具体环境的产物,不是复制某家公司。

随着人员增加,完全松散的结构需要调整。有些部门必须建立更严谨的层级和流程,另一些研究部门可以继续保持扁平和松散。调整已经开始,因为某些工作没有组织结构就无法推进。

但未来的组织仍然应同时保留两种能力:一部分能够可靠执行集体任务,另一部分能够自由探索尚未定义的问题。

研究组织最危险的事情,是为了提高表面效率,把所有空白时间都变成计划,把所有人都塞进明确项目。那样短期看起来更整齐,长期却失去发现下一条路线的能力。

English

Question: Is there a historical organization DeepSeek is trying to imitate? How will the structure change as the company grows?

Answer: There is no model to copy. Each decision was made by examining the actual situation and choosing among the tradeoffs. The organization is a product of its time and circumstances, not an imitation of another company.

As headcount grows, a completely loose structure must change. Some departments need clear hierarchy and rigorous processes, while research groups can remain flatter and less structured. This adjustment has already begun because certain kinds of work cannot proceed without organization.

The future company should preserve two capabilities at once: one part must reliably execute collective tasks, while another remains free to explore problems that have not yet been defined.

The greatest danger for a research organization is turning every blank space into a plan and placing every person inside a formally defined project in the name of efficiency. The company becomes tidier in the short term and loses its ability to discover the next path.