Concise Summary简洁概述
Nvidia's moat, per Huang, isn't chip lock-in but the sheer difficulty of turning electrons into tokens efficiently — backed by CUDA's ecosystem breadth, installed base, and cross-cloud portability.
On China, Huang argues export controls can't stop AI progress since China already has sufficient 7nm capacity and cheap energy — restrictions only accelerate Chinese chip self-sufficiency.
老黄认为 Nvidia 的护城河不是芯片锁定,而是把电子高效转化为 Token 这件事本身足够难,再叠加 CUDA 生态广度、装机规模与跨云可移植性三重支撑。
关于对华政策,老黄认为出口管制挡不住中国 AI 发展,因为中国已有足够的 7nm 产能和廉价能源,限制只会加速中国芯片走向完全自主。
Infographic信息图
Electrons to tokens
电子到 Token
Huang frames Nvidia's moat as the electron→token conversion itself — a poorly-understood, hard-to-commoditize engineering problem, not a chip SKU. This reframes 'is Nvidia getting commoditized by AI' as a category error.
老黄把 Nvidia 的护城河定义为“电子转 Token”这一过程本身——一个远未被理解透彻、难以商品化的工程问题,而非某款芯片。这把“AI 会不会让 Nvidia 商品化”这个问题直接框定为伪命题。
CUDA's three-layer moat
CUDA 三层护城河
Ecosystem breadth (Triton, vLLM, SGLang all run on CUDA), installed base (hundreds of millions of GPUs across every cloud), and cross-cloud portability — not technical lock-in — are Huang's stated reasons customers stay, plus on-site engineers who deliver 2-3x speedups.
老黄给出的留存理由不是技术锁定,而是生态广度(Triton、vLLM、SGLang 均跑在 CUDA 上)、装机规模(跨所有云的数亿块 GPU)与跨云可移植性,外加驻场工程师带来的 2-3 倍提速。
Export controls backfire thesis
出口管制反噬论
Huang argues China already has enough 7nm capacity (equivalent to old Hopper chips) plus near-free energy to brute-force compute via scale; restrictions only accelerate full chip self-sufficiency while forfeiting the world's second-largest tech market.
老黄认为中国已拥有足够的 7nm 产能(相当于老款 Hopper 芯片)和近乎免费的电力,足以靠堆规模弥补算力代差;出口限制只会加速中国芯片完全自主,同时白白让出全球第二大科技市场。
Never pick winners, never price-gouge
不押赢家,不趁机涨价
Nvidia refuses to auction scarce GPUs to the highest bidder or bet on which AI lab wins, a philosophy Huang traces to nearly failing among 60 rival 3D-graphics startups early in his career — yet the piece notes this coexists with $30B+ in strategic bets on OpenAI and Anthropic.
Nvidia 不搞 GPU 竞价、不押注某个 AI 实验室会赢,这源自老黄早年在 60 家 3D 图形公司混战中险些出局的教训——但文章指出,这与其向 OpenAI、Anthropic 砸下合计超 300 亿美元的战略投资其实并不完全自洽。
Detailed Summary详细解读
The interview's opening move sets its tone: Dwarkesh asks whether Nvidia, as a software company, will itself be commoditized by AI. Huang's answer — that turning electrons into tokens is a complex, poorly-understood process — is less a factual claim than a rhetorical pivot that recasts every subsequent challenge (TPU competition, CUDA erosion, margin pressure) as evidence of unsolved complexity rather than vulnerability. This framing does real work throughout the piece: it lets Huang treat competitors' gains as narrow point solutions against Nvidia's broader claim to 'accelerated computing' generally.
The supply-chain section is the piece's strongest evidence for Huang's non-obvious claim that bottlenecks resolve in 2-3 years. It names concrete mechanisms — the Micron HBM conversation, TSMC packaging co-development, the $2B Lumentum investment — rather than asserting confidence abstractly. This matters because it shows Nvidia's dominance is partly manufactured through coordination costs competitors can't easily replicate, not just chip design superiority.
The GPU-vs-TPU exchange is the interview's technical core. Dwarkesh's systolic-array critique — that matrix multiplication doesn't need general-purpose transistors — is a legitimate efficiency argument that specialized ASIC advocates make regularly. Huang's rebuttal leans on the Blackwell 50x efficiency gain (independently confirmed by SemiAnalysis, not just Nvidia's own number) to argue that algorithm/architecture co-design, not raw specialization, drives the largest gains — implicitly conceding that pure matmul efficiency isn't where Nvidia wins.
Huang's admission about Anthropic is the interview's most candid moment: he directly states Nvidia lacked the capital to back Anthropic early, forcing its reliance on Google/AWS chip ecosystems. This is a real concession, but the piece is right to flag that calling it 'an exception, not a trend' is doing a lot of load-bearing work — Anthropic's TPU commitment grew from 1GW to 3.5GW within a year, and OpenAI is separately building Triton/AMD relationships, which weakens the one-off framing.
The China section is where the interview's stakes are highest and Huang's position least reconciled. He simultaneously argues controls are futile (China has capacity and energy to route around them) and that the US must maintain permanent AI leadership — without ever specifying what restriction level, if any, he'd support. The piece correctly identifies this as an unresolved tension rather than a coherent policy stance, and notes his 'DeepSeek debuting on Huawei chips would be a disaster for America' line as the most quotable but least examined claim in the whole exchange.
The closing critique section is the piece's editorial value-add: it doesn't just relay Huang's talking points but tests them against Nvidia's own actions (the 'do less' philosophy versus $30B to OpenAI, $10B to Anthropic, $20B for Groq) and against unresolved questions (what export-control level Huang actually endorses, whether the no-price-gouging promise survives real competition). This is the section that elevates the piece from transcript summary to analysis.
访谈开局就定下基调:Dwarkesh 问身为软件公司的 Nvidia 会不会被 AI 自身商品化。老黄的回答——把电子转 Token 描述为一个复杂、远未被理解的过程——与其说是事实陈述,不如说是一次修辞转向,把后续所有挑战(TPU 竞争、CUDA 被侵蚀、毛利承压)都重新定义为“复杂性尚未攻克”而非“护城河出现裂缝”。这个框架在全文反复起作用:它让老黄可以把竞争对手的进展说成是针对具体环节的窄点突破,而 Nvidia 主张的是更广泛的“加速计算”。
供应链一节是文章为老黄“瓶颈两三年内必解”这一非直觉论断提供的最有力证据。它给出了具体机制——说服美光押注 HBM、与台积电共研封装、20 亿美元投资 Lumentum——而非空泛地表达信心。这一点很重要,因为它说明 Nvidia 的优势有一部分来自协调成本本身(竞争对手难以复制),而不仅是芯片设计上的领先。
GPU 对 TPU 的交锋是访谈的技术核心。Dwarkesh 的脉动阵列质疑——矩阵乘法不需要通用晶体管——是专用 ASIC 支持者常提的合理效率论点。老黄的反驳借助了 Blackwell 相比 Hopper 50 倍能效提升这一数字(经 SemiAnalysis 独立验证,非 Nvidia 自报数据),主张架构与算法协同设计而非纯粹专用化才是最大增益来源——这其实间接承认了纯矩阵乘法效率并非 Nvidia 的取胜战场。
老黄关于 Anthropic 的坦白是访谈中最真诚的一刻:他直接承认早年 Nvidia 资金不足,无法支持 Anthropic,才导致后者依赖 Google/AWS 芯片生态。这是一次真实的让步,但文章准确指出“特例而非趋势”这句话承担了过重的论证负担——Anthropic 对 TPU 的承诺一年内从 1GW 涨到 3.5GW,OpenAI 也在另建 Triton/AMD 关系,这都削弱了“一次性个案”的说法。
对华一节是整场访谈利害关系最高、老黄立场最难自洽的部分。他一方面主张管制毫无意义(中国有产能和能源可以绕过),另一方面又坚持美国必须保持永久领先——却始终没有说明自己到底能接受何种程度的限制。文章准确指出这是一处未解决的张力,而非连贯的政策立场,并点出“如果 DeepSeek 首发用了华为芯片,那对美国是灾难”这句最具传播力却最少被推敲的论断。
结尾批评一节是本文最具编辑价值的部分:它没有止步于转述老黄的话术,而是用 Nvidia 自身的行动去检验其说法(“尽量少做”哲学 vs 向 OpenAI 投 300 亿、Anthropic 投 100 亿、收购 Groq 花 200 亿),并指出仍悬而未决的问题(老黄到底能接受什么程度的出口管制、不涨价承诺在真正竞争来临时能否维持)。正是这部分让文章从访谈摘要升级为分析。
FAQ常见问答
Does Huang actually explain why Nvidia won't be commoditized, or just assert it?老黄真的解释了 Nvidia 为什么不会被商品化,还是只是断言?
He gives three concrete mechanisms — ecosystem breadth, installed base, cross-cloud portability — plus on-site engineers delivering measurable 2-3x speedups, not just a bare assertion.
他给出了三个具体机制——生态广度、装机规模、跨云可移植性——外加驻场工程师带来可衡量的 2-3 倍提速,并非空泛断言。
Is the Blackwell 50x efficiency claim independently verified or just Nvidia marketing?Blackwell 50 倍能效提升是独立验证的,还是 Nvidia 自己的营销数字?
The piece notes SemiAnalysis independently confirmed the 50x figure after Nvidia's initial 35x claim was met with skepticism, giving it more credibility than a self-reported benchmark.
文章提到 Nvidia 最初宣称 35 倍时外界并不信,后经 SemiAnalysis 独立分析确认实际达到 50 倍,因此比自报数字更可信。
Does Huang ever say what level of export control he'd actually accept?老黄有没有明确说过自己能接受什么程度的出口管制?
No — the piece explicitly flags this as unanswered. He argues against restriction while also insisting the US must stay ahead, without reconciling the two positions.
没有——文章明确指出这是他始终未回答的问题。他反对限制出口,同时又坚持美国必须保持领先,两种立场之间的矛盾未被化解。
How strong is the 'Anthropic is an exception' argument, really?“Anthropic 是特例”这个说法到底站不站得住?
Weaker than Huang presents it: Anthropic's TPU commitment grew from 1GW to 3.5GW in a year and OpenAI is separately building Triton/AMD ties, suggesting a pattern rather than an isolated case.
比老黄说得要脆弱:Anthropic 对 TPU 的承诺一年内从 1GW 涨到 3.5GW,OpenAI 也在另外发展 Triton/AMD 关系,这更像是一种趋势而非孤立个案。
Why did Nvidia acquire Groq if its own GPU architecture is supposedly superior?如果 Nvidia 自家 GPU 架构本就更优,为什么还要收购 Groq?
Huang frames it as responding to a new tiered-inference pricing market — customers now pay premiums for faster token response — not as an admission that GPU architecture underperforms.
老黄将其解释为应对新出现的“推理分层定价”市场——客户愿为更快的 Token 响应付溢价——而非承认 GPU 架构本身不如人。
In-depth Analysis · Pros & Cons深入解读 · 优缺点
Dwarkesh Patel's two-hour interview pushes Jensen Huang past his usual talking points, forcing direct answers on TPU competition, CUDA lock-in, and China export controls. The piece reconstructs the sharpest exchanges and then holds Huang's claims up against the interview's own internal contradictions.
Dwarkesh Patel 长达两小时的专访把黄仁勋逼出了惯常话术,让他就 TPU 竞争、CUDA 锁定与对华出口管制给出直接回应。文章还原了最激烈的交锋段落,并回过头用访谈本身暴露的矛盾去检验老黄的说法。
- Names concrete mechanisms给出具体机制Rather than vague praise for Nvidia's moat, the piece cites named episodes — the Micron HBM conversation, the Lumentum investment, Blackwell's independently-verified 50x gain — that ground abstract claims in checkable specifics.文章没有泛泛夸赞 Nvidia 的护城河,而是引用了具体事件——与美光的 HBM 对话、对 Lumentum 的投资、经独立验证的 Blackwell 50 倍能效提升——让抽象论断落到可核实的细节上。
- Preserves the adversarial texture保留了针锋相对的辩论质感The piece keeps Dwarkesh's pushback intact rather than flattening it into a monologue — the systolic-array critique, the 7nm/EUV counterpoint on China, giving readers both sides of each exchange.文章保留了 Dwarkesh 的追问和反驳,而非把访谈压缩成老黄的独白——脉动阵列的质疑、关于中国 7nm/EUV 差距的反驳都完整呈现,让读者能看到交锋双方的论点。
- Ends with genuine critique, not summary结尾是真正的批评而非总结The closing section explicitly tests Huang's stated philosophy against his actual capital allocation and flags the unanswered export-control question — an editorial layer most interview recaps skip.结尾部分明确用老黄的实际资本配置去检验他所宣称的哲学,并点出悬而未决的出口管制问题——这是大多数访谈摘要会跳过的编辑层分析。
- Distinguishes verified vs. self-reported claims区分了独立验证与自我陈述的数字By noting SemiAnalysis independently confirmed the 50x Blackwell figure, the piece signals awareness that not all of Huang's numbers carry equal evidentiary weight.文章特意指出 50 倍数字经 SemiAnalysis 独立验证,说明作者意识到老黄给出的各项数字证据分量并不相同。
- No independent fact-check of China capacity claims未独立核实中国产能说法Huang's claims about China's 7nm capacity, energy abundance, and Huawei's interconnect capabilities are presented largely unchallenged with hard figures; the piece flags Dwarkesh's pushback but doesn't independently verify either side's numbers.老黄关于中国 7nm 产能、能源充裕度和华为互联能力的说法基本未附具体数据佐证;文章记录了 Dwarkesh 的反驳,但并未独立核实双方数字的准确性。
- Margin durability claim untested毛利可持续性未经检验The 'never price-gouge, always predictable' promise is explicitly noted as resting on 70%+ gross margins that haven't yet faced real competitive pressure — the piece flags this but can't resolve it since the test hasn't happened yet.“绝不涨价、永远可预测”的承诺被文章明确指出建立在尚未经历真正竞争压力的 70%+ 毛利率之上——文章点出了这一点,但由于该压力测试尚未发生,无法给出定论。
- Relies on a single, PR-adjacent source信息源单一且偏公关性质The entire piece derives from one podcast interview where Huang controls the framing; there's no cross-reference to Anthropic, Google, or independent analysts' rebuttals beyond the interview's own back-and-forth.全文仅依据一场由老黄主导话语框架的播客访谈,除了访谈内的现场交锋外,没有交叉引用 Anthropic、Google 或独立分析师的其他反驳材料。
- 'Do less' philosophy critique could go further对“少做事”哲学的批评仍可更进一步The piece notes the contradiction between Huang's stated minimalism and his $60B+ in combined investments/acquisitions, but stops short of asking whether this pattern signals genuine platform strategy or defensive consolidation of the AI stack.文章指出了老黄“少做事”哲学与其合计超 600 亿美元投资并购之间的矛盾,但未进一步追问这种模式究竟是真正的平台战略,还是对 AI 全栈的防御性整合。
Read this if you want the sharpest available window into Nvidia's official narrative on CUDA moats, supply chains, and China policy — but treat every number as Huang's framing, not neutral fact, and note that the piece's own closing critique is more reliable than his talking points on export controls and the 'do less' philosophy.
如果你想了解 Nvidia 官方叙事中关于 CUDA 护城河、供应链与对华政策最锋利的一面,这篇文章值得一读——但每个数字都应视为老黄的立场陈述而非中立事实,文章结尾的批评部分在出口管制和“少做事”哲学上,反而比老黄本人的说法更可信。
Original Text原文
The English text on this side is an AI translation provided for convenience; the authoritative version is the source in the other language.
Jensen Huang recently sat down for a dense, two-hour interview with Dwarkesh Patel. There was no small talk — just a rapid-fire clash of ideas. If you don't have time to watch the whole thing, remember this one line: the man who controls the lifeline of global AI infrastructure summed up Nvidia's mission this way:
"The input is electron, the output is tokens. That is in the middle, Nvidia."
The whole conversation was intense and blunt. Host Dwarkesh pressed hard on every topic, and when it came to chip exports to China, the two of them went at it for a full twenty minutes. Huang strongly rejected the analogy of AI chips as weapons, while Dwarkesh kept zeroing in on the newly released Anthropic Claude Mythos model and its startling cyberattack capabilities, pressing the point relentlessly.
Source: Dwarkesh Patel Podcast (early April 2026) Original video: YouTube link
Key takeaways
Nvidia's operating philosophy is to "do less, but make everything unique." That's why Nvidia doesn't run its own cloud, doesn't bet on a single winner, and doesn't ration GPUs through bidding.
Supply-chain bottlenecks will be resolved within two to three years at most — the real long-term constraint is energy policy, not chip capacity.
CUDA's moat isn't technical lock-in — it's the massive installed base of GPUs, the rich ecosystem, and portability across different cloud platforms. Nvidia also sends engineers directly to help AI companies optimize their models, typically boosting speed by 2-3x with ease.
Anthropic's use of TPUs and Trainium is a special case, not a trend — the result of Nvidia failing to invest in Anthropic early on, which forced the company to rely on Google's and Amazon's chip ecosystems instead.
China already has enough 7nm chip capacity and abundant energy — export restrictions can't actually stop AI development. Instead they'll push China toward full chip self-sufficiency, and the U.S. will simply lose the world's second-largest tech market for nothing.
Nvidia's acquisition of Groq wasn't because its own GPU architecture falls short — it's because inference-market tokens have become expensive enough to be priced in tiers.
From Electrons to Tokens: How Does Nvidia Define Itself?
Right out of the gate, Dwarkesh threw a sharp question: if what Nvidia does is essentially write software, and AI is gradually commoditizing software, won't Nvidia itself eventually get commoditized too?
Huang flatly rejected the framing: "Turning electrons into tokens is, by itself, almost impossible to commoditize." That's because the conversion process is not only complex but still far from fully understood. He emphasized Nvidia's philosophy:
"We should do as much as needed, as little as possible."
This line ran through the entire interview, explaining why Nvidia doesn't build its own cloud, doesn't favor any particular winner, and doesn't allocate GPUs by bidding.
On the counterintuitive question of whether AI will devalue software companies, Huang's view was actually the opposite. He believes the number of AI agents will grow exponentially, and these agents will need to make heavy use of existing software tools — tools that in the past could only be operated by a limited pool of human engineers.
He gave a vivid example: Synopsys's Design Compiler, a chip-design tool, could see explosive growth in usage instances in the future — because AI agents will use these tools in bulk for design exploration. Far from wiping out software companies, this will unleash unprecedented demand.
Huang summed it up: "Today's agents aren't yet good at using these tools. Going forward, either the tool companies build their own agents, or the agents get smart enough on their own to use the tools efficiently. Both things will happen at the same time."
The Supply-Chain "Evangelist": How Does Huang Drive the Whole Industry?
In the interview, Dwarkesh went straight for what looks like Nvidia's unbreakable moat — citing the hundred-billion-dollar purchase commitments from the latest earnings report, as well as SemiAnalysis's even higher $250 billion valuation, and questioning whether Nvidia was locking out competitors by simply buying up the entire market.
Huang didn't deny it, but stressed it's not that simple. The real reason Nvidia dares to spend that much is that downstream demand is large enough to give suppliers the confidence to expand capacity:
"More demand → upstream dares to invest → capacity grows → the market gets bigger → even more demand" — that's the flywheel actually running behind Nvidia.
But that flywheel doesn't spin on its own. Huang described spending enormous effort acting as the supply chain's "evangelist," going CEO by CEO up and down the chain to explain clearly why the AI wave is coming, when it's coming, and how big it will be:
"I need them to see the future the way I see it."
He specifically recounted a pivotal conversation with Micron CEO Sanjay Mehrotra, laying out in detail why the HBM (high-bandwidth memory) market was about to explode. It later proved that Micron's bet on HBM and LPDDR was an extremely successful call.
Further upstream, in optical communications, Nvidia directly led a reshaping of the supply chain — co-developing packaging technology with TSMC, creating new processes, sharing patents with partners, and directly investing to help them expand capacity. The recent $2 billion investment in Lumentum is a prime example.
Huang described Nvidia's annual GTC conference as a collective "mind upgrade" for the whole supply chain, upstream and downstream alike:
"People keep telling me, 'Jensen, your keynotes sometimes feel like a lecture,' but that's exactly the point. I want everyone across the supply chain to see the huge opportunity AI is bringing as clearly as I do."
Capacity Bottlenecks? They Don't Exist
At this point, Dwarkesh pressed another sharp question: Nvidia already takes up the vast majority of TSMC's 3nm capacity — reportedly 60% in 2026 and as much as 86% in 2027. Given that enormous base, how could Nvidia possibly double again?
Huang's answer was strikingly optimistic. He said flatly:
"Every capacity bottleneck lasts two to three years at most. Once you can build one, you can build a million."
At any given moment, instantaneous demand will outrun supply capacity, and the bottleneck can show up somewhere you'd never expect — even plumbers. "We're actually inviting plumbers to next year's GTC."
He used CoWoS — TSMC's advanced packaging technology for integrating chips with high-bandwidth memory — as an example: two years ago it was the biggest bottleneck for AI chips, but after Nvidia put in serious effort, capacity has since multiplied several times over.
More importantly, Nvidia starts anticipating bottlenecks years in advance. Take silicon photonics (using light to transmit data): Nvidia not only develops the core technology itself but also invests in and partners with TSMC and others to expand capacity — keeping full control of the supply chain.
But Huang stressed that what genuinely worries him isn't these hardware bottlenecks — it's downstream energy policy:
"Without energy, you can't build any industry, let alone re-industrialize America. Chips, EVs, robots, AI factories — they all consume energy, and energy isn't a problem you solve in two or three years."
GPU vs. TPU: F1 Race Car or Cadillac?
At this point, Dwarkesh raised another pointed observation: the world's two strongest AI models — Claude and Gemini — were both trained on Google's TPUs. Doesn't that mean Nvidia has already fallen behind?
Facing this challenge, Huang quickly widened the frame: "What Nvidia builds isn't a 'tensor processing unit' — it's broader 'accelerated computing.'" Beyond AI, Nvidia's GPUs also cover dozens of fields, including molecular dynamics, fluid dynamics, quantum computing, and data processing. That kind of broad applicability is something a specialized chip (ASIC) can never match.
Even more critically, Nvidia serves as general-purpose infrastructure across the cloud — anyone can use it, and it runs on every major cloud platform, including Google, Amazon, Azure, and OCI. TPUs and Trainium, by contrast, can only be used by their specific cloud providers.
Dwarkesh clearly wasn't convinced by this explanation. Speaking for some AI researchers, he pointed out that a TPU is essentially a systolic array purpose-built to optimize matrix multiplication — simple and efficient — while a GPU's generality is actually wasteful: "AI is just repeated matrix computation. Why build a GPU that can also do other things, when that just wastes transistor area?"
Huang firmly pushed back: "Matrix multiplication matters, sure, but it's not all of AI. If you want to try new attention mechanisms or fuse different model architectures, you need a general-purpose computing platform like a GPU."
Then he brought out the numbers directly — the new Blackwell architecture delivers a full 50x improvement in energy efficiency over Hopper:
"Moore's Law gets you at most 25% a year. But if you want a 10x or even 100x leap, the only way is to keep changing the algorithms and the way you compute."
Huang specifically noted that when Nvidia first announced Blackwell was 35x more energy-efficient than Hopper, nobody believed it. It wasn't until SemiAnalysis did an independent analysis that they found the actual improvement was 50x! And that breakthrough didn't come from a process-node upgrade — it came from an all-around leap in processor architecture, algorithms, distributed computing strategy, and networking technologies like NVLink and Spectrum-X.
Finally, Huang closed with a vivid analogy:
"A GPU is like an F1 race car, a CPU is more like a Cadillac. Anyone can drive a Cadillac to 100 mph, but pushing an F1 car to its limit requires real professional driving skill."
And Nvidia itself has that kind of "professional driving skill" — using AI to automatically generate optimal compute kernels, easily helping customers double performance or more. "Given how many Hopper and Blackwell units are already deployed worldwide, that 2x performance boost directly translates into doubled revenue for our customers."
CUDA: Ecosystem Convenience or Technical Lock-in?
Dwarkesh pressed again, sharply: if 60% of Nvidia's revenue comes from a handful of giant customers who have every resource to write their own kernels, how much of an edge does CUDA really still provide? Anthropic and Google are already pushing their own chips, and even OpenAI — which relies on Nvidia GPUs — has built the Triton framework so that kernel programming no longer depends on CUDA.
Facing this question, Huang gave a three-part answer in one breath:
First, the richness of the ecosystem. CUDA already supports nearly every mainstream framework — from OpenAI's Triton to vLLM and SGLang, and now the newest reinforcement-learning frameworks Verl and NeMo RL. Nvidia itself is actively involved in Triton's low-level development, which means that if something breaks, you at least know clearly the problem is in your own code, not in some bottomless underlying implementation.
Second, the massive installed base. "As a developer, what do you care about most? Obviously, how many users have it installed!" Nvidia has hundreds of millions of GPUs deployed worldwide, from the aging A10 to the latest Blackwell, spanning every cloud platform and every vertical. Whether you're building cloud services or robotics, you want your code to run anywhere, anytime — and CUDA guarantees exactly that.
Third, freedom across cloud platforms. Nvidia is the only chip company present simultaneously on every major cloud service — Google, Amazon, Azure, and OCI. AI companies can't be sure which cloud provider they'll end up choosing in the future, so Nvidia offers them maximum flexibility and peace of mind.
But Dwarkesh wasn't letting go: do these advantages really matter that much to top-tier customers? As AI gets better and better at writing efficient kernels on its own, won't Nvidia end up as just another chipmaker competing purely on performance and price? At that point, could Nvidia still sustain gross margins above 70%?
Huang responded with total confidence: "Nvidia's engineers don't just optimize kernels — they optimize a customer's entire technology stack."
He said no one understands Nvidia's architecture better than Nvidia itself: "Our performance-per-watt and total cost of ownership (TCO) are the best in the world, no exceptions."
He even publicly called out competitors like TPU and Trainium: "I encourage them to step up and prove their inference performance using a recognized benchmark like Inference Max. But the reality is, they simply don't dare to."
Finally, Huang pointed to an important detail that's easy to overlook: while 60% of Nvidia's revenue does come from a handful of hyperscale cloud providers, most of those GPUs ultimately serve external customers. Cloud providers are willing to buy Nvidia at massive scale precisely because Nvidia itself brings in the most customers. That's Nvidia's real ecosystem moat.
Anthropic's Chip Choices: Huang Admits a Mistake
One of the most interesting moments in the interview was when Huang openly explained why Anthropic makes heavy use of Google's TPUs and Amazon's Trainium.
Dwarkesh mentioned Anthropic's just-announced TPU deal with Google and Broadcom, totaling 3.5GW of compute capacity, and sharply asked: "If Nvidia's price-performance is truly the best in the world, why would Anthropic still choose someone else?"
Jensen's response was very direct: "Anthropic is an exception, not a trend."
He even went so far as to say, exaggeratedly:
"Without Anthropic, where would TPU's growth come from? 100% it's Anthropic. Without Anthropic, where would Trainium's growth come from? Also Anthropic."
While reality isn't quite as absolute as Jensen put it, his core point is clear enough—among all of Nvidia's major customers, Anthropic is the only one that has visibly leaned toward other chip ecosystems.
More interestingly, Jensen then voluntarily admitted to a major misstep of his own in the past:
"A long time ago, we really weren't capable of doing this. I underestimated how hard it is to build a large model lab like OpenAI or Anthropic, and I underestimated how much massive financial backing from suppliers they would need. Back then Nvidia simply couldn't come up with 5 or 10 billion dollars for Anthropic, but Google and AWS could."
He said bluntly:
"Our misstep forced Anthropic to go find others. But even so, I'm still happy that Anthropic exists—Anthropic is good for the world."
Afterward, having learned his lesson, Jensen resolved not to repeat the same mistake, which is why Nvidia later made major investments in OpenAI ($30 billion) and Anthropic ($10 billion).
On why Nvidia doesn't run its own cloud service, Jensen again emphasized his core business philosophy:
"If we don't take the risk of building the computing platform, truly no one will. Without NVLink, CUDA, and the building and investment of the whole ecosystem, the AI industry wouldn't be thriving the way it is today. But cloud services are different—there are plenty of people in the world who can do that. If we don't do it, someone else naturally will."
He stressed that Nvidia would not personally get into the financing business, because "there are already plenty of people doing financing in the market—we'd rather partner with all the people who do financing."
As for why Nvidia never bets on a single winner, Jensen recalled a lesson from his early days starting the company: "When we were just starting out, there were 60 3D graphics companies across the whole industry. If you'd taken a vote back then on who was most likely to fail, we'd have been ranked first, because our architectural direction was flat-out wrong."
Finally, Jensen summed it up:
"I'm humble enough to know not to pick winners. Either everyone figures out how to survive on their own, or we simply take care of everyone."
GPU Pricing Philosophy: Why No Price Hikes, No Bidding Wars
During the interview, Dwarkesh threw out a question that intuitively feels "reasonable": "When GPUs are scarce, why not just sell to the highest bidder?"
Jensen gave an unexpected but firmly confident answer:
"We've never done bid-based GPU allocation. That's not Nvidia's style—we're only responsible for setting a reasonable price, and customers decide for themselves whether to buy."
Jensen stressed that even when the market is on fire, they would never raise prices to take advantage: "Other chip companies might, but we absolutely won't. We want to be the trustworthy cornerstone of the whole industry, so customers never have to guess or worry about whether we'll take advantage of the moment to gouge them."
He even debunked a rumor in passing: "There's a rumor going around that Larry Ellison and Elon Musk once begged me at a dinner to allocate them GPUs. That part is true, but they didn't need to beg—they just needed to place an order."
Even more surprising, Jensen revealed that in nearly 30 years of cooperation between Nvidia and TSMC, they have never signed a formal legal contract:
"There's always been a kind of tacit understanding between us—sometimes I come out ahead a bit, sometimes I lose out a bit, but overall it's fair. I can trust them and rely on them completely."
Finally, he pointed out another of Nvidia's advantages—extremely strong predictability:
"This year we're delivering the Vera Rubin architecture, next year it's Vera Rubin Ultra, the year after that it's Feynman, and the year after that there's a new, as-yet-unannounced architecture."
He summed up with pride:
"Look around the world—which ASIC team dares to promise: a stable new architecture every year, and token costs dropping by an order of magnitude every year? Only Nvidia can do that. We're as punctual and reliable as a clock."
Export Controls on China: Jensen's Fiercest Exchange of the Whole Interview
Midway through the interview, Dwarkesh kicked off the most intense 20-minute debate of the session. He said outright that he's used to playing "devil's advocate"—he had previously challenged Dario Amodei, who supports export controls, and now, facing Jensen, who opposes controls, he flipped the same sharp question around:
"If Chinese companies and the government got hold of AI chips capable of training a top-tier model like Anthropic's Claude Mythos, wouldn't that threaten US national security?"
Claude Mythos was just released recently, and it has already discovered thousands of zero-day vulnerabilities, even able to autonomously find high-risk vulnerabilities in mainstream operating systems and browsers. Precisely because of this, Anthropic doesn't dare release it publicly, and only makes it available in limited fashion to organizations like Google and Microsoft to patch vulnerabilities.
Jensen pushed back quickly and firmly: "The compute needed to train Mythos isn't actually anything special—it's already widespread in China. They have over 60% of the world's chip manufacturing capacity, they have top-tier computer scientists, and 50% of the world's AI researchers also come from China."
He put it even more sharply: "If you're really worried about China, using the worst possible approach—turning them into victims, into enemies—is definitely not a good idea."
Jensen pointed out that right now, the biggest thing missing between the US and China is genuine dialogue between AI researchers. He believes both sides should discuss openly and directly, and make clear which areas AI should not touch.
Dwarkesh disagreed with Jensen's view. He stressed that although China has a lot of chips, its advanced compute is only one-tenth of America's, because China's process node is still stuck at 7nm, without EUV equipment. This means the US has a precious window in which it can reach the level of Mythos faster than China and plug the vulnerabilities beforehand.
But Jensen pointed to a reality that gets overlooked:
"The free energy China has access to is genuinely staggering. AI is fundamentally a massive parallel computing problem—if Chinese chip compute isn't advanced enough, they can absolutely stack up a supercomputer using huge numbers of cheap chips and nearly free electricity. They even have empty ghost cities, ghost data centers, that could be scaled up rapidly."
He added further: "7nm chips are basically our old Hopper chips, and the vast majority of the world's advanced AI models were trained using Hopper. Huawei, moreover, set the largest single-year chip shipment record in history last year."
In Jensen's view, export controls will only push China to become chip-self-sufficient faster, while the US will, as a result, needlessly give up the world's second-largest tech market.
Dwarkesh tried to keep pressing with the "memory bandwidth" issue, but Jensen simply responded: "Huawei is essentially a networking company—they're fully capable of using advanced interconnect technology to string together huge numbers of ordinary chips into a supercomputer. They've even already demonstrated the ability to turn low-end chips into a giant supercomputer using silicon photonics."
Then Jensen dropped the most controversial line of the whole interview:
"If in the future a top model like DeepSeek launches first on Huawei chips, that would be a disaster for the United States."
His logic is very clear: if open-source AI models get optimized for a non-US tech stack, then as those models spread across the Global South, the Middle East, and Southeast Asia, American technology standards and hardware ecosystems will lose their competitiveness, and the US will lose its global dominance in AI technology.
Dwarkesh immediately countered: "Anthropic's models can run on GPUs, TPUs, and Trainium at the same time—this kind of cross-platform compatibility won't easily disappear."
Jensen was dismissive: "Try it—take a model optimized for Nvidia and move it to another platform, and see how the performance turns out. Nvidia's success is the best proof of this. AI models are created on our tech stack, and they achieve their best results on our tech stack too—there's no doubt about that."
Next, Dwarkesh cited a pointed analogy that Dario Amodei once made at Davos: "Nvidia selling chips to China is like Boeing proudly saying, 'North Korea's nuclear missile casings were made by Boeing, so this counts as supporting America's tech ecosystem.'"
This immediately infuriated Jensen, who fired back forcefully:
"Comparing AI to whatever you just mentioned is absolutely absurd."
Dwarkesh wasn't done yet: "But if their chips can run AI models capable of attacking all American software, isn't that a weapon?"
Jensen calmly pointed out: "The real solution is for the US and China to reach a clear consensus through communication, to make sure no country misuses AI technology. What's more, China is the world's largest contributor to open source, and AI safety depends on a global open-source ecosystem. We can't strangle it."
Dwarkesh seized on a subtle contradiction in Jensen's logic and pressed further: "On one hand you say Nvidia chips are the strongest and will definitely win in the China market; on the other hand you say that even without selling them chips, China can still achieve the same thing."
Jensen stressed that the two points aren't contradictory: "If there's a better chip on the market, naturally people choose the better one. If there isn't, they can still make do with what's available—that's completely logical."
By the end of the debate, both sides' positions were very clearly drawn:
Jensen insisted that AI technology is like a "five-layer cake," and every layer must be competed for and won. Sacrificing the chip layer (that is, not selling chips to China) to prevent them from training top-tier models is an extreme and short-sighted approach. He cited the example of the US telecom industry: strict export controls in the past once caused American companies to completely lose the global telecom market, and ultimately left the US "no longer in control of its own telecom industry."
Dwarkesh fiercely pressed Jensen on whether this kind of reasoning amounted to a "loser mentality." Jensen went all in on the spot:
"The person in front of you is not someone who wakes up in the morning ready to lose. That kind of loser attitude, loser assumption, means nothing to me."
Finally, Jensen calmly summed up his core view:
"No one is saying it has to be either fully open or fully closed. America must always stay ahead and have the best technology. But at the same time we should also actively compete globally and win the market. These two things can absolutely be done at the same time. The world was never black and white."
Behind the Groq Acquisition: Inference Enters an Era of Tiering
In the final stretch of the interview, Dwarkesh pulled the topic back from sensitive geopolitics to technology, asking Jensen: "Why doesn't Nvidia try different chip architectures? Like Cerebras's wafer-scale chips, or a large-package structure like Dojo, or even release a version that doesn't rely on CUDA at all?"
Jensen's response was simple and direct: "Of course we could, but it's been proven that these architectures aren't actually better. We've already tested them repeatedly in our simulators, and the results were all worse than our current approach."
That said, he admitted that Nvidia has indeed recently taken a new step in the inference chip space—by acquiring Groq at a high price.
But Jensen stressed that acquiring Groq wasn't because GPU architecture isn't good enough—it's because the inference market itself has undergone a major shift:
"A few years ago, tokens produced by inference were basically worthless—you could even call them free. But now different customers have different requirements for tokens; they're willing to pay a higher price for faster response speed."
He gave the example of his own software engineers: "If we can give engineers faster-responding tokens and double their productivity, of course we're willing to pay extra for that." This is what's called the "premium inference market"—it didn't exist in the past, but it's forming rapidly now.
Jensen continued to explain: "We've now entered an 'era of inference tiering.' The same model can be priced differently based on response speed, which means high throughput is no longer the only standard. Faster response speed, even with somewhat lower overall throughput, can still command a higher average selling price (ASP)."
As for whether they'd consider going back to older process nodes (like 7nm) to ease chip supply pressure, Jensen was firm that it's unlikely:
"In theory you could do that, but it makes no economic sense at all. We can afford to move forward, but we can't afford to move backward. Each generation's architectural progress isn't just about the process node—it's also about packaging, stacking technology, numerical precision, and overall system architecture innovation. Unless there really comes a day when we truly can't increase capacity any further, we will never go backward."
If the AI Revolution Had Never Happened, What Would Nvidia Be Doing?
Toward the end of the interview, Dwarkesh raised a hypothetical, somewhat philosophical question: "If the deep learning revolution had never happened, what would Nvidia be doing today?"
Jensen didn't hesitate, answering flatly: "We'd still be doing accelerated computing—that's what Nvidia has been doing all along." He stressed that the era of general-purpose computing has run its course, and the shift of the computing world toward accelerated computing isn't necessarily tied to AI.
"Even if AI didn't exist, Nvidia would still be a very large company."
He said that even without AI, fields like computational lithography, quantum chemistry, data processing, and image generation still need powerful accelerated computing capability. He mentioned that a large portion of the topics at GTC aren't centered on AI at all, yet remain critical to the industry.
"Tensor computing isn't all of computing. We want to be able to help every field of computing."
But in the end Jensen also candidly admitted a small feeling deep down: "But if there really were no AI in the world, I would feel very sad."
Lightning Round
Q: Will Nvidia become a commoditized company? Jensen: No. Because the conversion from electrons to tokens is inherently very complex, and the engineering and scientific problems are still far from being fully understood. The explosion of AI agents will also bring enormous growth in demand for software tools.
Q: What is CUDA's greatest value, really? Jensen: It's not technical lock-in, but the installed base of hundreds of millions of GPUs worldwide, an extraordinarily rich ecosystem, and convenience across every cloud platform. Beyond that, Nvidia's engineers can also easily help customers achieve 2-3x performance gains.
Q: Why did Anthropic choose TPUs over GPUs? Jensen: Because in its early years Nvidia lacked the financial capacity to invest in Anthropic in time, which forced them to rely on Google's and AWS's chip ecosystems. This is a special case, not a long-term trend.
Q: Should Nvidia sell AI chips to China? Jensen: Yes. China already has ample 7nm chip capacity and abundant energy, and export restrictions will only accelerate the buildup of China's own domestic ecosystem. The US needs to compete aggressively across every technology layer, rather than trying to "win" by sacrificing the market.
Q: Why doesn't Nvidia run its own cloud service? Jensen: "Because if we don't do it, someone else will." Nvidia's philosophy is to only do the things that no one else could do if we didn't do them. Cloud infrastructure clearly doesn't fall into that category.
Old Huang's Self-Contradictions and the Test of Reality
This nearly two-hour interview surfaced several contradictions and open questions worth continuing to watch.
First, Jensen's so-called philosophy of "doing as little as possible" has long since drifted from its literal meaning in practice. $30 billion invested in OpenAI, $10 billion in Anthropic, $20 billion spent acquiring Groq, plus backing emerging cloud providers like CoreWeave... Nvidia's tentacles have in fact already reached deep into every link of the AI industry chain. In hindsight, this principle looks more like a retroactively polished narrative than a rule that actually constrains its actions.
Second, whether the claim that "Anthropic is a special case" can withstand the test of reality remains an open question. Over the past year, Anthropic's TPU collaboration with Broadcom and Google has scaled up several times over (from 1GW to 3.5GW), while at the same time OpenAI has begun developing its own Triton framework and partnering with AMD. If more such "special cases" emerge in the future, the foundation of Jensen's argument will inevitably be shaken.
Then there's the most sensitive topic: exports to China. Throughout the debate, Jensen never directly answered the most crucial question: "What level of export restriction is he actually willing to accept?" He repeatedly stressed that exports should not be restricted, while also insisting that the US must stay perpetually ahead. When these two positions come into conflict, which one takes priority? He gave no clear answer. His example of the "failure of the US telecom industry" is thought-provoking, but whether this analogy fully applies to a technology as strikingly dual-use as AI chips remains debatable.
As for business strategy, Jensen repeatedly emphasized that Nvidia, as the industry's "cornerstone," will never raise prices or allocate GPUs through bidding wars, and can precisely forecast a new architecture release every year. But the credibility of this commitment rests heavily on Nvidia's current gross margin of over 70%. Once real competition arrives, will Nvidia still be able to afford such generosity and composure?
Going forward, the actual mass-production progress and performance of the Vera Rubin architecture, how Groq's 3 LPX performs in real-world inference workloads, and the true scale of deployment of China's homegrown chips over the coming months will all be the best indicators for testing the many claims made in this interview with Jensen.
Full interview video: Dwarkesh Patel Podcast
黄仁勋最近接受了 Dwarkesh Patel 长达两个小时的高密度专访。这场访谈没有客套寒暄,只有密集的观点碰撞。如果你没时间看完,那就记住这一句话——掌控全球 AI 基础设施命脉的老黄,用这样一句话定义了 Nvidia 的使命:
“输入是电子,输出是 Token,中间是 Nvidia。” (“The input is electron, the output is tokens. That is in the middle, Nvidia.”)
整场对话的气氛激烈又直白。主持人 Dwarkesh 在每个话题都紧追不放,尤其涉及中国芯片出口时,两人更是直接“杠”了整整二十分钟。老黄强烈反对把 AI 芯片当作武器的类比,而 Dwarkesh 则紧盯着刚发布不久、网络攻击能力惊人的 Anthropic Claude Mythos 模型,穷追猛打。
访谈来源:Dwarkesh Patel Podcast(2026 年 4 月初) 原始视频:YouTube 链接
要点速览
- Nvidia 的经营哲学是“少做,但每件事都是独一无二”。这也是为什么 Nvidia 不做云、不押注赢家、不竞价分配 GPU。
- 供应链的瓶颈最多两三年就能解决,真正制约未来的是能源政策,而非芯片产能。
- CUDA 的壁垒不是技术锁定,而是庞大的 GPU 装机量、丰富的生态系统,以及在不同云平台之间的可移植性。Nvidia 还直接派驻工程师帮 AI 公司优化模型,通常能轻松提升模型速度 2-3 倍。
- Anthropic 使用 TPU 和 Trainium 是特殊个案而非趋势,原因在于早年 Nvidia 没及时投资 Anthropic,逼得后者只能依靠 Google 和 Amazon 的芯片生态。
- 中国早已拥有足够的 7nm 芯片产能和大量能源,限制出口根本挡不住 AI 发展,反而会推动中国芯片的完全自主化,让美国白白丢掉全球第二大科技市场。
- Nvidia 收购 Groq 不是因为自家 GPU 架构不行,而是因为推理市场的 Token 已经贵到可以分档收费了。
【1】从电子到 Token:Nvidia 如何定义自己?
一开场,Dwarkesh 就抛出一个尖锐的问题:如果 Nvidia 的工作本质是写软件,而 AI 又正在把软件逐渐商品化,那 Nvidia 自己会不会也被商品化?
老黄干脆直接把问题推翻重来:“把电子转化成 Token,这事儿本身几乎没法被商品化。”因为这个转化过程不仅复杂,而且远未被充分理解。他强调 Nvidia 的哲学是:
“做必要的事情,越少越好。” (“We should do as much as needed, as little as possible.”)
这句话贯穿了整个访谈,解释了 Nvidia 为什么不自己做云,不偏向某个赢家,也不搞竞价分配 GPU。
而针对“AI 会不会让软件公司贬值”这个反直觉问题,老黄的观点恰好相反。他认为未来 AI 智能体的数量将指数级增长,这些智能体需要大量使用现有的软件工具,而过去这些工具只能由有限数量的人类工程师操作。
他举了个生动的例子:芯片设计公司 Synopsys 的设计编译器(Design Compiler),未来的使用实例可能会爆炸式增长——因为 AI 智能体将成批地使用这些工具进行设计探索。这不仅不会淘汰软件公司,反而会带来史无前例的需求大爆发。
老黄总结道:“今天智能体还不擅长使用这些工具,未来要么工具公司自己开发智能体,要么智能体自动变聪明,学会高效使用工具。这两个过程会同时发生。”
【2】当供应链“布道者”:老黄如何推动整个产业链?
访谈中,Dwarkesh 直接戳中了 Nvidia 看似牢不可破的护城河——他援引了最新财报里的千亿美元采购承诺,以及 SemiAnalysis 更高的 2500 亿美元估值,质疑 Nvidia 是不是靠“买空市场”卡住了竞争对手。
老黄没否认这点,但强调事情没这么简单。Nvidia 敢砸出这么多钱的根本原因,是下游需求足够大,给了供应商十足信心去扩产:
“需求多 → 上游敢投 → 产能增 → 市场更大 → 需求更多”, 这是 Nvidia 背后真实运转的飞轮。
但这飞轮并不是自动转起来的。老黄形容自己花了大量精力做供应链“布道者”,挨个儿向上下游 CEO 讲清楚 AI 大潮为什么会来、什么时候来、以及将有多大:
“我需要让他们看到我看到的未来。”
他特意讲了与美光(Micron)CEO Sanjay Mehrotra 的关键对话,详细告诉对方为何 HBM(高带宽内存)市场马上会爆炸式增长。后来事实证明,美光押注 HBM 和 LPDDR 是个极为成功的决策。
在更上游的光通信领域,Nvidia 直接牵头重塑了供应链:和台积电一起开发封装技术,创造新工艺,还主动把专利分享给合作伙伴,并直接投资帮他们扩大产能——最近对 Lumentum 的 20 亿美元投资,就是典型案例。
老黄把 Nvidia 每年举办的 GTC 大会,定义为产业链上下游的集体“思想升级”:
“有人总跟我说『Jensen,你演讲有时像在上课』,但这正是我要做的事。我希望产业链上的所有人,都能像我一样清晰地看到 AI 即将到来的巨大机会。”
【3】产能瓶颈?不存在的
访谈到这里,Dwarkesh 再次追问了一个尖锐问题:Nvidia 已经占据了台积电 3nm 产能的绝大部分,2026 年甚至达到 60%,2027 年更要占到 86%。在如此巨大的基数下,Nvidia 怎么可能再翻倍增长?
老黄的回答异常乐观,他直截了当地表示:
“所有产能瓶颈最多持续两三年。一旦你能造一个,就能造一百万个。”
任何时候,瞬时需求都会超越供应能力,瓶颈可能出在你完全想不到的环节,甚至可能是水管工人——“明年 GTC 我们还真邀请了水管工参加。”
他以 CoWoS(台积电用来集成芯片和高带宽内存的先进封装技术)为例,两年前它还是 AI 芯片的最大瓶颈,Nvidia 花了大力气解决,现在产能已经翻了数倍。
更重要的是,Nvidia 提前数年就开始主动预判瓶颈。比如硅光子(用光传输数据)领域,Nvidia 不仅亲自开发关键技术,还投资并与台积电及合作伙伴联手扩产,完全掌握供应链的主动权。
但老黄强调,真正让他担忧的并非这些硬件瓶颈,而是下游能源政策:
“没能源,你什么产业都建不了,更别提再工业化美国了。芯片、电动车、机器人、AI 工厂,这些都吃能源,而能源问题可不是两三年就能解决的。”
【4】GPU vs TPU:F1 赛车还是凯迪拉克?
访谈到这里,Dwarkesh 又抛出了一个犀利的观点:世界上最强的两个 AI 模型——Claude 和 Gemini,都是用 Google 的 TPU 训练出来的。这是不是意味着 Nvidia 已经落后了?
面对这个挑战,老黄迅速把讨论格局拉大:“Nvidia 做的不是『张量处理单元』,而是更广泛的『加速计算』。”除了 AI,Nvidia 的 GPU 还能覆盖分子动力学、流体力学、量子计算、数据处理等数十个领域。这种广泛的适用性,是专用芯片(ASIC)无论如何也追不上的。
而且更关键的是,Nvidia 是云计算领域的通用基础设施,任何人都能操作,能跑在 Google、Amazon、Azure、OCI 等所有主流云平台上。但 TPU 和 Trainium 这些芯片就不同,只能被特定的云服务商使用。
Dwarkesh 显然对这个解释不买账。他代表一些 AI 研究者指出,TPU 本质上就是专门优化矩阵乘法的脉动阵列(systolic array),简单、高效,而 GPU 的通用性反而成了浪费:“做 AI 就是反复的矩阵计算,你非得弄个能做其他事情的 GPU,晶体管面积不是白瞎了吗?”
对此,老黄坚决反驳:“矩阵乘法当然重要,但不是 AI 的全部。如果你想试试新的注意力机制、融合不同的模型架构,你就需要 GPU 这种通用计算平台。”
然后他直接给出了数据——新一代的 Blackwell 架构相比 Hopper 架构能效提高了整整 50 倍:
“靠摩尔定律,每年最多提升 25%。但你要实现 10 倍甚至 100 倍的飞跃,唯一的方法就是不断改变算法和计算方式。”
老黄特意提到,最初宣布 Blackwell 比 Hopper 能效高 35 倍时,没人相信。直到 SemiAnalysis 独立分析后,才发现实际提升居然达到 50 倍!而这一突破的背后,靠的不是制程升级,而是处理器架构、算法、分布式计算策略,以及 NVLink、Spectrum-X 等网络技术的全面创新。
最后,老黄以一个生动的比喻结束:
“GPU 就像 F1 赛车,CPU 更像凯迪拉克。凯迪拉克人人都能开到时速 100 英里,但想把 F1 赛车推到极限,你必须拥有专业的驾驶技术。”
而 Nvidia 自己就拥有这种专业的“驾驶技术”——用 AI 自动生成最优计算内核,帮客户轻松将性能提升 2 倍甚至更多。“考虑到现在 Hopper 和 Blackwell 在全球的装机量,这 2 倍的性能提升,直接等于客户收入翻倍。”
【5】CUDA:生态便利还是技术锁定?
访谈中 Dwarkesh 再次犀利提问:如果 Nvidia 60% 的收入都来自少数几个巨头客户,而这些客户完全有资源自己写内核,那 CUDA 到底还有多大优势?Anthropic 和 Google 已经开始主推自己的芯片,连依赖 Nvidia GPU 的 OpenAI 都开发了 Triton 框架,让内核编程不再依赖 CUDA。
面对这个问题,老黄一口气给出了三个层面的答案:
第一,是生态的丰富程度。 CUDA 已经支持了几乎所有主流框架,从 OpenAI 的 Triton,到 vLLM、SGLang,再到最新的强化学习框架 Verl 和 NeMo RL。Nvidia 自己也积极参与 Triton 的底层开发,这意味着如果出了问题,你至少明确知道问题出在自己代码上,而不是无底洞般的底层实现。
第二,是庞大的装机规模。 “作为开发者,你最在乎什么?当然是用户装机量!” Nvidia 在全球拥有数亿块 GPU,从老旧的 A10 到最新的 Blackwell,横跨各个云平台、每个垂直领域。无论你是做云服务还是机器人,你都希望自己的代码随时随地能跑起来,而 CUDA 就能保证这一点。
第三,是跨云平台的自由。 Nvidia 是唯一一个能同时存在于 Google、Amazon、Azure 和 OCI 等所有主流云服务上的芯片公司。AI 公司无法确定自己未来会选哪家云服务,因此 Nvidia 能提供最大的灵活性和安全感。
但 Dwarkesh 还是不依不饶:这些优势对那些顶级客户真的有那么重要吗?当 AI 越来越擅长自己写高效内核时,Nvidia 会不会变成单纯拼性能和价格的芯片商?到那个时候,Nvidia 还能维持超过 70% 的高毛利率吗?
对此老黄非常自信地回应:“Nvidia 工程师优化的不只是内核,而是客户整个技术栈。”
他说,没有人比 Nvidia 自己更懂 Nvidia 的架构,“我们的每瓦性能和总拥有成本(TCO)都是全球最好的,没有例外。”
他甚至公开喊话 TPU 和 Trainium 等竞争对手:“我鼓励他们站出来,用 Inference Max 这种公认的基准测试证明自己的推理性能。但现实是,他们根本不敢来。”
最后,老黄指出了一个容易被忽视的重要细节: 虽然 Nvidia 60% 的收入确实来自几家超大型云厂商,但这些 GPU 大部分最终服务于外部客户。云厂商愿意大规模采购 Nvidia,是因为 Nvidia 本身带来了最多的客户。这才是 Nvidia 真正的生态护城河。
【6】Anthropic 的芯片选择:老黄承认了一个错误
访谈最有意思的地方,莫过于老黄公开解释 Anthropic 为什么大量使用 Google 的 TPU 和 Amazon 的 Trainium。
Dwarkesh 提到了 Anthropic 刚宣布的与 Google 和 Broadcom 总计 3.5GW 算力规模的 TPU 交易,尖锐地问道:“如果 Nvidia 性价比真的全球第一,Anthropic 为什么还要选别家的?”
对此,老黄的回应非常直接:“Anthropic 是一个特例,不是趋势。”
他甚至夸张地表示:
“没有 Anthropic,TPU 哪来的增长?100% 是靠 Anthropic。没有 Anthropic,Trainium 的增长从哪来?还是 Anthropic。”
虽然现实情况并没有老黄说得这么绝对,但他的核心意思其实很明确——在 Nvidia 所有的重要客户中,只有 Anthropic 明显地偏向了其他芯片生态。
更有意思的是,接下来老黄主动承认了自己过去的一个重大失误:
“很久以前,我们确实没能力这么做。我低估了建立一家像 OpenAI 或 Anthropic 这样的大模型实验室有多难,也低估了它们对供应商巨额资金支持的需求。当时 Nvidia 根本拿不出 50 亿、100 亿美元给 Anthropic,但 Google 和 AWS 能做到。”
他直言不讳地说:
“我们的失误导致 Anthropic 不得不去找别人。但即使这样,我仍然为 Anthropic 的存在感到高兴——Anthropic 对世界是有益的。”
后来老黄痛定思痛,决心不再犯同样的错误,因此 Nvidia 后续大手笔投资了 OpenAI(300 亿美元)和 Anthropic(100 亿美元)。
关于 Nvidia 为什么不自己做云服务,老黄再次强调了他核心的经营哲学:
“如果我们不冒险打造计算平台,真的就没人做了。如果没有 NVLink、CUDA 和整个生态的搭建与投入,AI 产业根本不会有今天的繁荣。但云服务不同,世界上有很多人能做。我们不做,自然会有人去做。”
他强调 Nvidia 不会亲自做融资业务,因为“融资业务市场上已经有很多人在做了,我们宁愿跟所有做融资的人合作。”
至于为什么 Nvidia 从来不去押注某个赢家,老黄回忆起创业时的教训:“我们刚起步的时候,全行业有 60 家 3D 图形公司。要是那时候投票选谁最可能失败,我们绝对排第一,因为我们的架构方向根本就是错的。”
最后老黄总结道:
“我足够谦逊地知道,不要去挑选赢家。要么大家自己想办法活下去,要么我们干脆照顾好所有人。”
【7】GPU 定价哲学:不涨价、不竞标的理由
访谈中,Dwarkesh 又抛出了一个让人直觉上觉得“这才合理”的问题:“GPU 紧缺时,为什么不直接卖给出价最高的人?”
对此老黄给出了一个出人意料但底气十足的回答:
“我们从来不做竞价分配 GPU 的事。这不是 Nvidia 的风格,我们只负责定好一个合理的价格,客户自己决定买不买。”
老黄强调,即使市场火爆到爆炸,他们也绝不会趁机涨价:“其他芯片公司可能会,但我们绝不。我们想做整个行业可信赖的基石,让客户永远不需要猜测或担心我们会不会趁机割韭菜。”
他甚至顺带辟了个谣:“坊间传闻 Larry Ellison 和 Elon Musk 曾经在一次晚餐上苦苦求我分 GPU 给他们,这事确实有,但他们根本不用求,只要下订单就可以了。”
更让人吃惊的是,老黄透露 Nvidia 与台积电近 30 年的合作,从来没签过正式法律合同:
“我们之间一直存在一种默契,有时候我占点便宜,有时候吃点亏,但总体上公平。我可以完全信任他们、依赖他们。”
最后,他点出了 Nvidia 的另一个优势——超强的可预测性:
“今年我们交付 Vera Rubin 架构,明年是 Vera Rubin Ultra,后年是 Feynman,再下一年还有未公布的新架构。”
他带着骄傲地总结道:
“你放眼全球,有哪家 ASIC 团队敢拍胸脯承诺:每年稳定推出新架构、每年 Token 成本持续下降一个数量级?只有 Nvidia 能做到,我们像钟表一样准时可靠。”
【8】对华出口管制:老黄全场最激烈的交锋
访谈进行到中段时,Dwarkesh 开启了本场最激烈的 20 分钟辩论。他直言,自己习惯当“魔鬼代言人”——此前他曾挑战过支持出口管制的 Dario Amodei,现在面对反对管制的老黄,他反过来问了同样尖锐的问题:
“如果中国企业和政府拥有了训练出类似 Anthropic Claude Mythos 这种顶级模型的 AI 芯片,会不会威胁美国国家安全?”
Claude Mythos 近期刚发布,就已发现了数千个零日漏洞,甚至能在主流操作系统和浏览器中自主发掘高危漏洞。正因如此,Anthropic 不敢公开发布,只限量提供给 Google、微软等机构修补漏洞。
对此老黄迅速而坚定地反驳:“Mythos 训练所需的算力其实并不特殊,在中国早已普遍存在。他们有世界 60% 以上的芯片产能,拥有最顶尖的计算机科学家,全球 50% 的 AI 研究者也都来自中国。”
他更犀利地提出:“如果你真担心中国,用最糟糕的方式——把他们变成受害者、敌人,肯定不是好主意。”
老黄指出,现在美中最大的缺失,是 AI 研究者之间的真正对话。他认为双方应该公开、直接地讨论,明确 AI 哪些领域不能涉及。
Dwarkesh 不认同老黄的观点。他强调,中国虽然芯片多,但先进算力只有美国的十分之一,因为中国的制程还卡在 7nm,没有 EUV 设备。这意味着美国有一个宝贵的窗口期,可以比中国更快达到 Mythos 的级别,并提前堵上漏洞。
老黄却一针见血地指出了被忽略的现实:
“中国拥有的免费能源实在太惊人了。AI 就是一个巨大的并行计算问题,如果中国芯片算力不够先进,他们完全可以用大量便宜芯片和几乎免费的电力拼成超算。他们甚至有空置的鬼城、鬼数据中心,可以迅速规模化部署。”
他进一步补充道:“7nm 芯片其实就是我们过去的 Hopper 芯片,而全球绝大部分先进 AI 模型,就是用 Hopper 训练出来的。华为去年更是创下了史上最大规模的单年芯片出货量。”
在老黄看来,出口管制只会促使中国更快走向芯片自主,而美国则将因此白白放弃全球第二大的科技市场。
Dwarkesh 试图用“内存带宽”问题继续施压,但老黄干脆回应:“华为本质上是一家网络公司,他们完全有能力通过先进的互联技术,把大量普通芯片串联成超级计算机。他们甚至已经展示过用硅光子技术把低端芯片变成巨型超算的能力。”
随后,老黄抛出了访谈中最具争议的一句话:
“如果未来 DeepSeek 这样的顶尖模型,首发选在华为芯片上,那对美国来说将是灾难。”
他的逻辑非常清晰:如果开源 AI 模型被优化到非美国的技术栈上,当这些模型传播到全球南方、中东和东南亚地区时,美国的技术标准和硬件生态将不再有竞争力,美国将丢掉全球 AI 技术主导权。
Dwarkesh 随即反驳道:“Anthropic 的模型同时能跑在 GPU、TPU 和 Trainium 上,这种跨平台兼容性不会轻易消失。”
老黄不以为然:“你试试看,把一个为 Nvidia 优化的模型搬到其他平台上去,性能会怎么样?Nvidia 的成功就是最佳证明。AI 模型在我们的技术栈上被创造,也在我们的技术栈上达到最好效果,这点毫无疑问。”
接下来,Dwarkesh 引用了 Dario Amodei 曾在达沃斯论坛上的尖锐比喻:“Nvidia 卖芯片给中国,就像波音自豪地说,朝鲜的核武器导弹外壳是波音造的,所以这是支持美国技术生态。”
这句话立刻激怒了老黄,他强烈回击:
“你把 AI 和你刚刚提到的任何东西比较,简直是荒谬透顶。”
Dwarkesh 仍不罢休:“但如果他们的芯片可以跑出攻击所有美国软件的 AI 模型,这难道不算武器?”
老黄冷静地指出:“真正解决之道,是美中通过沟通达成明确共识,确保所有国家都不滥用 AI 技术。更何况,中国是全球最大开源贡献者,AI 安全依赖于全球开源生态。我们不能掐死它。”
Dwarkesh 抓住老黄逻辑上的微妙矛盾追问:“你一边说 Nvidia 芯片最强,在中国市场一定能赢;一边又说,就算不卖芯片,中国照样可以做到一样的事。”
老黄强调,这两点并不矛盾:“如果市场上有更好的芯片,自然选更好的。如果没有,也能用已有的方案,这完全合乎逻辑。”
辩论到最后,双方立场都非常鲜明:
老黄坚持,AI 技术像一块“五层蛋糕”,每一层都必须竞争并取胜。牺牲芯片层(也就是不卖芯片给中国)来防止他们训练出高端模型,是一种极端且短视的做法。他举了美国电信行业的例子:过去严格的出口管制曾让美国公司彻底丢掉全球电信市场,最终让美国“不再控制自己的电信产业”。
Dwarkesh 激烈地质问老黄,这种说法是不是一种“输家心态”。老黄当场火力全开:
“你面前这个人,不是早上醒来准备输的人。这种输家的态度、输家的假设,对我来说毫无意义。”
最后,老黄平静地总结了自己的核心观点:
“没有人说要么全部开放、要么全封闭。美国必须永远领先,拥有最好的技术。但同时我们也应该积极参与全球竞争并赢下市场。这两件事是完全能同时做到的。世界从来不是非黑即白的。”
【9】收购 Groq 背后:推理市场进入分层时代
访谈最后阶段,Dwarkesh 将话题从敏感的地缘政治拉回到技术上,问老黄:“Nvidia 为什么不尝试不同的芯片架构?比如 Cerebras 那种晶圆级芯片,或者像 Dojo 那样的大封装结构,甚至推出完全不依赖 CUDA 的版本?”
老黄简单又直接地回应:“我们当然能做,但事实证明这些架构并没有更好。它们早就在我们的模拟器里反复验证过了,效果都不如现在的方案。”
不过他承认,Nvidia 最近的确在推理芯片领域迈出了新的一步——高价收购了 Groq。
但老黄强调,收购 Groq 并非因为 GPU 架构不够优秀,而是因为推理市场本身出现了重大变化:
“几年前,推理产生的 Token 基本不值钱,甚至可以说免费。但现在不同客户对 Token 的要求不同,他们愿意为更快的响应速度支付更高的价格。”
他举了自家软件工程师的例子:“如果我们能给工程师提供更快响应的 Token,让他们的生产力翻倍,我们当然愿意为此额外付费。”这就是所谓的“高端推理市场”,过去并不存在,但如今正在快速形成。
老黄继续解释:“我们现在已经进入了一个『推理分层时代』。同一个模型可以根据响应速度不同来定价,这意味着吞吐量高不再是唯一标准。更快的响应速度,即使整体吞吐量低一些,也可能获得更高的平均售价(ASP)。”
对于是否会考虑回到旧制程(比如 7nm)来缓解芯片供应压力的问题,老黄果断表示不太可能:
“理论上可以这么做,但经济上完全不划算。我们能承担得起向前发展,但负担不起向后退步。每一代新架构的进步,不仅仅是制程,还有封装、堆叠技术、数值精度和整体系统架构的革新。除非有一天真的再也无法提高产能,否则我们绝不会往回走。”
【10】假如 AI 革命从未发生,Nvidia 还会做什么?
访谈的尾声,Dwarkesh 提出了一个假设性的、稍显哲学的问题:“如果深度学习革命从来没有发生,今天的 Nvidia 会做什么?”
老黄没有犹豫,干脆地回答:“我们还是会做加速计算,这本来就是 Nvidia 一直在做的事。”他强调,通用计算的时代已经走到了尽头,计算世界转向加速计算的趋势和 AI 并不必然相关。
“即使 AI 不存在,Nvidia 依然会是一家非常大的公司。”
他说,即便没有 AI,计算光刻、量子化学、数据处理和图像生成等领域,依旧需要强大的加速计算能力。他提到 GTC 大会上有很大一部分话题并非围绕 AI 展开,但依然对产业至关重要。
“张量计算不是计算的全部。我们希望能帮助所有计算领域。”
但老黄最后也坦率承认了自己内心深处的一点小情绪:“但如果世界上真的没有 AI,我会感到非常伤心。”
最后的快问快答
Q:Nvidia 会不会变成商品化公司? Jensen:不会。因为从电子到 Token 的转化本身非常复杂,工程和科学问题还远未被理解透彻。AI 智能体的爆发还会为软件工具带来巨大的需求增长。
Q:CUDA 最大的价值到底是什么? Jensen:不是技术锁定,而是全球数亿块 GPU 的装机量、极其丰富的生态系统,以及跨每一家云平台的便捷性。此外,Nvidia 的工程师还能帮客户轻松实现 2-3 倍的性能提升。
Q:Anthropic 为什么选择 TPU 而不是 GPU? Jensen:因为 Nvidia 早年缺乏财务能力,没有及时投资 Anthropic,导致他们不得不依靠 Google 和 AWS 的芯片生态。这是一个特例,而不是长期趋势。
Q:Nvidia 应不应该向中国出售 AI 芯片? Jensen:应该。中国已经有了充足的 7nm 芯片产能和丰富的能源,出口限制只会加速中国自主生态的建立。美国需要在所有技术层面积极竞争,而不是通过牺牲市场来“赢”。
Q:Nvidia 为什么自己不做云服务? Jensen:“因为如果我们不做,总有人会做。”Nvidia 的哲学是只做那些如果我们不做,就没人能做的事情。而云基础设施显然不属于这一类。
老黄的自相矛盾与现实考验
这场近两个小时的访谈,呈现出了几个值得持续关注的矛盾与悬念。
首先,老黄所谓“尽可能少做事”的哲学,在现实中早已偏离了字面含义。投入 300 亿美元给 OpenAI,100 亿美元给 Anthropic,花 200 亿美元收购 Groq,又扶持 CoreWeave 等新兴云服务商……Nvidia 的触手实际上早已深入 AI 产业链的每一个环节。如今看来,这个原则更像是一种事后美化的叙事,而非真正约束行动的法则。
其次,“Anthropic 是特例”这一论断,能否扛住现实的检验也充满悬念。Anthropic 在过去一年与 Broadcom 和 Google 的 TPU 合作规模翻了数倍(从 1GW 增至 3.5GW),与此同时 OpenAI 也开始发展自己的 Triton 框架并和 AMD 合作。如果未来出现更多这样的“特例”,那么老黄的说法基础势必动摇。
再来看最敏感的对华出口话题。整场辩论下来,老黄始终没有直面回答一个最关键的问题:“他到底愿意接受什么程度的出口限制?”他反复强调不应该限制出口,又同时表示美国必须永远领先。一旦这两种立场发生冲突,究竟哪个更重要?他没有给出明确答案。他举出的“美国电信产业失败”的案例虽有启发性,但这种类比是否能完全适用于 AI 芯片这种显著的双重用途技术,仍有待商榷。
而在商业策略上,老黄一再强调 Nvidia 作为“行业基石”绝不涨价、不搞竞价分配 GPU,能精准预测每年发布新架构。但这个承诺的可信度,很大程度上依赖于 Nvidia 目前超过 70% 的超高毛利率。当竞争真正来临时,Nvidia 是否还能如此慷慨与淡定?
接下来 Vera Rubin 架构的实际量产进度与性能表现、Groq 3 LPX 在真实推理任务中的表现,以及中国自研芯片在未来几个月的真实部署规模,都将是检验老黄此次访谈诸多观点的最佳指标。
完整访谈视频:Dwarkesh Patel Podcast
See all posts