BY baoyu.io — Bilingual Study Editionbaoyu.io 最新 50 篇精读
All ↩目录 ↩
#32baoyu.io宝玉 · 2026-04-01 · baoyu.io

OpenAI President Greg Brockman: AI self-improvement, the Super App gamble, the road to AGI, and the compute expansionOpenAI 总裁 Greg Brockman:AI 自我改进、Super App 豪赌、通往 AGI 之路、算力扩张

Reading Greg Brockman's 90-minute reframing against the numbers he left out把 Brockman 90 分钟的叙事重新包装,放回被他省略的数字里检验

01

Concise Summary简洁概述

Brockman uses a 90-minute interview to reframe three reactive OpenAI decisions — killing Sora, taking $110B, betting on the Spud model — as proactive strategy.

The reframing mostly holds up logically but consistently omits the commercial pressure (burn rate, lost deals, market-share loss) driving each decision.

Brockman 用 90 分钟访谈,把 OpenAI 三个被动决策——砍 Sora、拿 1100 亿融资、押注 Spud 模型——重新包装成主动战略。

这套重新包装在逻辑上大体自洽,但系统性地回避了驱动每个决策的商业压力:烧钱速度、丢掉的交易、市场份额流失。

02

Infographic信息图

3
reactive decisions reframed as strategy
3 个被重新包装为战略的被动决策
50%→27%
OpenAI enterprise share decline vs. Anthropic's rise to ~40%
OpenAI 企业份额从 50% 降至约 27%,Anthropic 升至约 40%
18-24个月
compute lock-in period vs. estimated real runway
算力锁定周期与实际资金跑道均约 18-24 个月
🎬

Killing Sora: strategy or bleeding stopped

砍掉 Sora:战略选择还是止血

Brockman frames the Sora shutdown as a compute-allocation choice using a 'vector alignment' metaphor — scattered bets cancel out, focus creates momentum. He omits that Sora was burning ~$1M/day with users down from 1M to under 500K, and that Disney's $1B licensing deal collapsed as a result.

Brockman 用「向量对齐」的比喻解释砍 Sora:力量分散会互相抵消,集中才能产生动能。但他没提 Sora 当时每天烧掉约 100 万美元、用户从百万跌破 50 万,也没提 Disney 原定 10 亿美元的授权交易因此告吹。

🧩

Super App vs. the Microsoft problem

Super App 与绕不开的微软问题

OpenAI plans to merge ChatGPT, Codex, and the Atlas browser into one unified 'AI layer' product within months. The pitch overlaps almost exactly with Microsoft 365 Copilot — and since OpenAI is simultaneously Microsoft's model supplier and its direct competitor, the interviewer never raises this tension, and Brockman never resolves it.

OpenAI 计划几个月内把 ChatGPT、Codex 编程 Agent 和 Atlas 浏览器整合成统一的「AI 层」产品,这个构想和微软 365 Copilot 几乎重叠。OpenAI 既是微软的模型供应商又是直接竞品,这个张力访谈全程没人提起,Brockman 也没有回应。

🌫️

Vibes instead of benchmarks

用「感觉」代替数据

Across the hardest questions — Spud's capabilities, AGI progress, catching up to Anthropic — Brockman substitutes qualitative language ('big model smell', AGI as 'a vibe') for numbers. This isn't accidental: pre-launch secrecy and investor pressure to look ahead collide, leaving impressionistic claims as the only safe register.

面对最难的问题——Spud 的能力、AGI 进度、是否追上 Anthropic——Brockman 反复用「大模型的味道」「AGI 是一种氛围」这类感性描述代替具体数字。这不是偶然:发布前保密和向投资人证明领先的双重压力,让「感觉」成了唯一安全的表达方式。

💰

$110B runway math doesn't fully add up

1100 亿美元的跑道算不清账

Only ~$25B of the $110B raise is committed cash; the rest is GPU credits or milestone-gated. With an estimated $57B annual cash burn by 2027 and compute locked in 18-24 months ahead, Brockman's 'I disagree' to Amodei's YOLO-spending critique never engages the actual runway math.

1100 亿美元融资中只有约 250 亿是确定的现金,其余是 GPU 额度或附带里程碑条件。据估计到 2027 年 OpenAI 年现金消耗将达 570 亿美元,而算力采购要提前 18-24 个月锁定。Brockman 对 Amodei「YOLO 式烧钱」批评的回应只是一句「我不同意」,没有正面回应跑道测算。

The argument, step by step
论证推进链条
1
Set the stage: three reactive events (Sora shutdown, $110B raise, Spud pretraining) all need explaining within the same 90 minutes.
铺垫背景:砍 Sora、拿 1100 亿融资、Spud 预训练完成,三件事同时发生,都需要在这 90 分钟里给出解释。
2
Test the Sora explanation: the 'tech tree branch' framing is coherent but omits burn rate and the collapsed Disney deal.
检验 Sora 的解释:「技术树分支」框架自洽,但回避了烧钱速度和 Disney 交易告吹的事实。
3
Examine the Super App pitch and expose the unaddressed overlap with Microsoft 365 Copilot.
审视 Super App 构想,揭出它和微软 365 Copilot 重叠却无人追问的问题。
4
Check the 'caught up to Anthropic' claim against Synergy Research market-share data — the comparison quietly narrows to capability, not share.
用 Synergy Research 的市场份额数据核验「已追上 Anthropic」的说法——对比范围被悄悄限定在能力层面,回避了份额。
5
Probe the AGI/Spud claims: qualitative 'vibe' language stands in for benchmarks, and the 70-80% figure contradicts Brockman's own definition of AGI as undefinable.
追问 AGI 与 Spud 的说法:定性的「氛围」语言取代了具体指标,70-80% 的数字与他自己「AGI 无法精确定义」的说法自相矛盾。
6
Synthesize into three structural pressures — product competition, rhetorical evasion, and financial runway risk — that will be tested over the next six months.
收束为三层结构性压力——产品竞争、表达回避、财务跑道风险——并指出接下来六个月的实际交付才是检验标准。
03

Detailed Summary详细解读

The piece opens by setting the frame: Brockman walks into the interview right after OpenAI killed Sora, closed a $110B raise, and finished pretraining Spud — three events each demanding explanation, and the interview's task is to recast all three as proactive strategy. The author then tests each claim: Sora's 'different branch of the tech tree' framing is internally coherent, but WSJ reporting on the burn rate and the collapsed Disney deal reveals the other half of the story. This pattern — accept the polished frame, then check it against external data — runs through the whole piece.

On the Super App, the piece grants the vision — merging ChatGPT, Codex, and the Atlas browser into one 'AI layer' — real coherence, but flags that the interviewer never raised its near-total overlap with Microsoft 365 Copilot. OpenAI is simultaneously Microsoft's model supplier and its direct competitor, and that structural tension goes entirely unaddressed. Given Sora's fresh failure as a consumer product, the 'delivery within months' promise is treated as a genuine open question rather than a given.

The Anthropic-catch-up section delivers the sharpest contrast: Brockman admits past weakness in 'last-mile usability' and claims OpenAI has now caught up, with users preferring it in head-to-head comparisons. The author counters with Synergy Research figures — Anthropic's enterprise share rising from a low base to ~40% while OpenAI's fell from 50% to ~27% — showing Brockman quietly confined 'catching up' to model capability, sidestepping the harder market-share metric.

On Spud and AGI progress, the piece zeroes in on a logical gap: Brockman calls the definition of AGI 'more a vibe than a science,' yet simultaneously offers a precise self-rating of '70-80% there.' The article argues this contradiction — quantifying a concept explicitly declared unquantifiable — mirrors an industry-wide problem where 'vibes' masquerade as measurement across every major AI lab's AGI claims.

The financing section is the piece's most tightly argued: Brockman likens compute spend to hiring salespeople — more investment, more revenue — and answers Dario Amodei's 'YOLO spending' critique with a bare 'I disagree.' The author breaks down the $110B raise, showing only ~$25B is immediate cash, the rest GPU credits or milestone-gated, and pairs this with an estimated $57B annual burn by 2027 and an 18-24 month compute-lockin cycle — making the real financial cushion far thinner than the headline number suggests, and never actually answering Amodei's core point.

The closing section pulls the three threads into a 'three layers of pressure' framework: on the product side, the Super App faces a three-way squeeze from Microsoft, Anthropic, and Google; on the rhetorical side, Brockman's reliance on vibes over data is the inevitable product of colliding secrecy and investor-narrative pressures; on the scale side, the headline $110B masks a runway that may be as short as 18-24 months. This synthesis reframes all the piece's point-by-point checks into one verdict: only actual delivery over the next six months will tell whether this was foresight or spin.

文章开篇设定框架:Brockman 上场前 OpenAI 刚完成三件大事——砍 Sora、拿 1100 亿融资、跑完 Spud 预训练,每一件事都需要解释,而访谈的任务就是把这三件被动应对重新讲成主动选择。作者随后逐一验证:Sora 的「技术树分支论」听起来自洽,但补上 WSJ 披露的烧钱数据和 Disney 撤资细节后,故事的另一半浮现出来。这种「给出漂亮框架,再用外部数据检验」的写法贯穿全文。

关于 Super App,作者指出这个愿景(合并 ChatGPT、Codex、Atlas 浏览器为统一「AI 层」)本身有说服力,但采访者从未追问它和微软 365 Copilot 几乎重叠的问题——OpenAI 既是微软的模型供应商又要做直接竞品,这个结构性矛盾被完全绕开。加上 Sora 刚证明 OpenAI 做消费者产品的记录并不光彩,「几个月内交付」的承诺被作者标为存疑。

追赶 Anthropic 一节是全文最尖锐的对比:Brockman 承认此前在「最后一公里可用性」上落后,声称现在已经追上,用户在正面对比中更倾向 OpenAI。但作者引入 Synergy Research 数据——Anthropic 企业市场份额从 2023 年低位升至约 40%,OpenAI 从 50% 降到约 27%——指出 Brockman 悄悄把「追上」限定在模型能力层面,回避了市场份额这个更硬的指标。

关于 Spud 和 AGI 进度,作者聚焦一个逻辑漏洞:Brockman 说 AGI 的定义「更像是一种氛围」,无法用科学方式界定,却同时给出「70-80% 已实现」的精确自评。文章指出,一个承认无法精确定义的概念,却被量化到具体百分比,这种矛盾折射出整个行业在 AGI 话语上的困境——每个人都在用「感觉」冒充测量。

资金部分是全文论证最扎实的一段:Brockman 把算力比作雇佣销售——投入越多收入越高,回应 Dario Amodei「YOLO 式烧钱」批评时只说「我不同意」。作者拆解 1100 亿融资结构,指出仅约 250 亿是即时现金,其余为 GPU 额度或附带里程碑,配合预估的 570 亿年现金消耗和 18-24 个月的算力采购提前期,实际财务缓冲远比数字表面显得脆弱,而这正是 Amodei 论证的核心,Brockman 从未正面回应。

结尾把三条线收拢为「三层压力」:产品端 Super App 面临微软、Anthropic、Google 三面夹击;表达端 Brockman 用感性语言替代数据是保密与融资叙事双重压力下的必然产物;规模端 1100 亿美元听起来安全,实际跑道可能只有 18-24 个月。这个总结性框架把全文的逐项核验拉回到一个统一的判断:接下来 6 个月的实际交付,才是检验这套叙事的唯一标准。

04

FAQ常见问答

Did Brockman ever directly address Dario Amodei's core financial critique?Brockman 有没有正面回应 Amodei 对财务风险的核心质疑?

No. He said 'I disagree' to the charge of reckless spending but never engaged Amodei's specific claim that a one-year revenue misprediction could bankrupt the company — the article treats this as the interview's clearest dodge.

没有。他只用「我不同意」回应「烧钱过于激进」的指控,从未正面回应 Amodei「收入预判偏差一年就可能破产」的具体论证——文章认为这是访谈中最明显的回避。

Is the claim that Sora's shutdown was purely a technical/focus decision credible?「砍 Sora 纯粹是技术聚焦决策」这个说法可信吗?

Only partially. The vector-alignment logic is coherent, but it omits that Sora was burning ~$1M/day, users had fallen from 1M to under 500K, and Disney's $1B licensing deal fell through as a direct result.

只是部分可信。「向量对齐」逻辑本身自洽,但它回避了 Sora 每天烧掉约 100 万美元、用户从百万跌破 50 万,以及 Disney 10 亿美元授权交易因此告吹的现实。

How solid is the '70-80% to AGI' figure?「AGI 完成 70-80%」这个数字有多可靠?

Not very — Brockman himself says AGI's definition is 'more a vibe than a science,' which makes a precise percentage self-contradictory. The article flags this as symptomatic of an industry-wide habit of quantifying an admittedly unquantifiable concept.

不太可靠——Brockman 自己说 AGI 的定义「更像是一种氛围」,这就让一个精确的百分比自相矛盾。文章指出这反映了整个行业量化一个自认无法精确定义的概念的通病。

Does OpenAI actually have more compute runway than the $110B figure suggests?OpenAI 的实际算力资金跑道,是否真的比 1100 亿这个数字看起来更充裕?

Likely less. Only ~$25B is committed cash; the rest is GPU credits or milestone-gated funding, and estimated annual burn could reach $57B by 2027 — implying an effective runway closer to 18-24 months.

很可能更少。1100 亿中只有约 250 亿是确定现金,其余是 GPU 额度或附带里程碑的资金,而预估到 2027 年年现金消耗可达 570 亿美元,实际跑道可能只有 18-24 个月。

Is the Super App's overlap with Microsoft 365 Copilot a real strategic problem?Super App 和微软 365 Copilot 的重叠,是不是一个真实的战略问题?

Yes — OpenAI is Microsoft's model supplier while building a directly competing unified-workflow product, and the interview never raises or resolves this conflict of interest, leaving it as an open structural risk.

是。OpenAI 一边是微软的模型供应商,一边在打造一个直接竞争的统一工作流产品,访谈全程没有提出或解决这个利益冲突,这仍是一个悬而未决的结构性风险。

05

In-depth Analysis · Pros & Cons深入解读 · 优缺点

This piece deconstructs a 90-minute Greg Brockman interview line by line, separating the polished 'strategic choice' framing from the commercial and competitive pressures each decision actually reflects. It uses external data (WSJ, Synergy Research, The Information) to test Brockman's claims against reality rather than take them at face value.

这篇文章逐句拆解了 Greg Brockman 长达 90 分钟的访谈,把访谈里包装成「战略选择」的说辞,和每个决策背后真实的商业与竞争压力剥离开来。作者引入 WSJ、Synergy Research、The Information 等外部数据,把 Brockman 的说法放到事实里检验,而不是照单全收。

Strengths亮点 / 优点
  • Data-checked skepticism
    用数据核验的怀疑态度
    Nearly every major claim from Brockman is paired with an external data point (WSJ, Synergy Research, The Information) rather than accepted at face value, giving the critique real evidentiary weight.
    Brockman 的几乎每个重要说法都配有外部数据核验(WSJ、Synergy Research、The Information),而非照单全收,让批评具备真实的证据分量。
  • Identifies the logical self-contradiction on AGI
    精准指出 AGI 论述的自相矛盾
    Catching the tension between 'AGI is a vibe, not a science' and 'we're 70-80% there' is a sharp, specific piece of analysis rather than generic skepticism.
    抓住「AGI 是氛围不是科学」与「已完成 70-80%」之间的矛盾,是具体而犀利的分析,而不是泛泛的质疑。
  • Surfaces the unasked Microsoft question
    揭出访谈中被回避的微软问题
    Flagging that neither the interviewer nor Brockman addresses the Super App's direct overlap with Microsoft 365 Copilot exposes a real structural risk the original interview left invisible.
    指出访谈者和 Brockman 都没有触及 Super App 与微软 365 Copilot 的直接重叠,揭示了一个原访谈中被完全隐藏的结构性风险。
  • Balanced — credits what holds up
    保持平衡,肯定站得住脚的部分
    The piece doesn't dismiss Brockman wholesale — it grants that the vector-alignment logic for Sora and the compute-as-revenue-center argument are internally coherent before pointing out what's missing.
    文章没有全盘否定 Brockman——在指出缺失之前,先承认「向量对齐」逻辑和「算力是收入中心」的论证在内部是自洽的。
Limits & Critiques局限 / 批评
  • Second-hand critique
    批评基于二手信源
    The article's counter-evidence (WSJ, Synergy Research, The Information estimates) is itself aggregated reporting, not primary verification — some figures (e.g. $57B burn by 2027) are analyst estimates, not disclosed facts.
    文章用来反驳 Brockman 的证据(WSJ、Synergy Research、The Information 的估算)本身也是二手报道而非一手核实——部分数字(如 2027 年 570 亿现金消耗)是分析师估算,并非公开披露的事实。
  • Unverifiable anecdotes on both sides
    双方都存在无法验证的轶事
    Both Brockman's internal-use case studies (physicist example, Codex/Slack integration) and some of the author's rebuttal points (Abilene water usage skepticism) rely on unverifiable claims presented with similar confidence.
    Brockman 举的内部案例(物理学家案例、Codex 接 Slack)和作者部分反驳论点(对 Abilene 数据中心用水量的质疑)同样缺乏独立验证,却以相近的笃定语气呈现。
  • Market-share snapshot may be stale
    市场份额数据可能已过时
    The Synergy Research enterprise-share figures are a point-in-time snapshot; fast-moving competitive dynamics in AI coding tools mean the 40%/27% split could shift materially within months of publication.
    Synergy Research 的企业市场份额数字是某个时间点的快照;AI 编程工具赛道变化极快,40% 对 27% 的格局可能在文章发表后几个月内就发生明显变化。
  • No counter-interview with Brockman
    缺少对 Brockman 的追问反驳环节
    The critique is entirely retrospective analysis of a completed interview; OpenAI or Brockman never gets a chance to respond to the specific data points raised, so some rebuttals remain one-sided by construction.
    整篇批评是对一场已完成访谈的事后分析,OpenAI 或 Brockman 从未有机会针对文章提出的具体数据点做出回应,部分反驳因此天然是单方面的。
Bottom line
总评

Worth reading for anyone tracking OpenAI's strategic narrative versus its underlying financial and competitive reality — but read it as a corrective to the original interview, not a standalone profile, since its rebuttal evidence is itself second-hand and time-sensitive.

适合关注 OpenAI 战略叙事与真实财务、竞争处境之间落差的读者,但应把它当作对原访谈的核验补充来读,而非独立的完整报道——因为文中用来反驳的证据本身也是二手且带时效性的。

06

Original Text原文

The English text on this side is an AI translation provided for convenience; the authoritative version is the source in the other language.

When Greg Brockman sat down in the Big Technology Podcast studio, OpenAI had just done three big things: killed Sora, secured $110 billion in funding, and finished pretraining its next-generation model, Spud. Cutting a marquee product, raising a sky-high round, and betting on a new model — three things happening at once, each demanding an explanation. Over the next 90 minutes, the OpenAI co-founder and president had a difficult task: repackage every seemingly reactive decision as a forward-looking strategic choice.

Did he pull it off? Mostly yes, but the cracks in the narrative hide some clues worth examining closely.

For the past 18 months, Brockman has mainly run the company's "Scale" division, overseeing GPU infrastructure, data centers, and supply chains. This interview spans an enormous range of topics: from why Sora was cut, to what the Super App will look like, to what the next-generation model can do, to how the $110 billion will be spent, to why he donated $25 million to MAGA Inc.

Original video: https://www.youtube.com/watch?v=J6vYvk7R190

Key takeaways

Was killing Sora "focusing on the core technical line" or stopping the bleeding? Brockman says Sora and GPT are different branches of the technology tree, and with limited compute you can only pick one. He didn't mention that Sora was burning about a million dollars a day and that its user base had fallen below 500,000.

The Super App will merge ChatGPT, the coding agent Codex, and the browser Atlas into one unified app, expected to ship within the next few months — but this vision directly overlaps with Microsoft 365 Copilot, and no one in the interview asked about that.

The next-generation model Spud crystallizes roughly two years of research breakthroughs. Brockman says it will raise both the ceiling and the floor of AI capability, but the descriptions are all impressionistic, with no concrete metrics.

He self-rates AGI progress at 70-80%, but also says the definition of AGI is "more like a vibe" — he never resolves the contradiction of precisely quantifying something that can't be precisely defined.

He admits OpenAI once lagged behind Anthropic on "last-mile usability" for coding tools, and claims they've now caught up — though third-party data shows Anthropic's enterprise market share is still growing rapidly.

Responding to Anthropic CEO Dario Amodei's criticism that OpenAI's infrastructure investment is "too aggressive": "I disagree." But he didn't address Amodei's specific argument that a one-year miss on revenue projections could lead to bankruptcy.

Killing Sora: a technical choice or a business bleed-stop?

Kantrowitz opened by asking: OpenAI leads in the consumer market, so why suddenly pull resources away from video generation?

Brockman's explanation framework differs quite a bit from outside speculation. He said Sora's model and the GPT series are "completely different branches on the technology tree," built in fundamentally different ways. In a world of limited compute, pushing two technology paths at once is extremely costly. He used a vector metaphor:

The sum of random vectors is zero — only aligned direction lets you move forward.

Scattered force cancels itself out; only concentrated force generates momentum. He believes the problem in deep learning isn't a shortage of opportunities but the opposite — too many — so spreading investment thin means not going far in any direction.

This explanation is internally consistent on technical logic, but it's only half the story. According to WSJ investigations and TechCrunch reporting, Sora had already failed commercially: its user count had plunged from a peak of one million to under 500,000, burning roughly $1 million a day in compute costs. Per WSJ reporting, Disney's planned $1 billion investment and character licensing deal with Sora was also called off, and Disney learned of the news less than an hour before it was announced publicly. Shutting down Sora had both a technical-focus rationale and real commercial pressure to stop the bleeding — Brockman chose to talk about only the former.

He stressed this isn't "shifting from consumers to enterprise" but a forced choice. He believes the two most important applications right now are: a personal assistant that understands you and is aligned with your goals, and an AI that can solve hard problems for you. Just these two things already exceed OpenAI's current compute capacity.

One technical detail: the image generation feature in ChatGPT is unaffected. Image generation is built on the GPT architecture, on the same technical line as text and voice; Sora uses a diffusion model, a different branch entirely. The Sora research team wasn't disbanded — it pivoted to world-simulation research for robotics.

Kantrowitz asked further: Google DeepMind's Demis Hassabis believes the model closest to AGI isn't a text model but an image generator, because it must understand how objects interact and how the world works. By abandoning that path, is OpenAI betting on the wrong track?

Brockman said this is a real risk: "Absolutely. In this field you have to make choices, you have to place bets." But he then pivoted, saying OpenAI is confident because it has "already seen the path to AGI." He gave an example: a physicist handed a long-unsolved problem to an OpenAI model, and 12 hours later got a solution — the physicist said it was the first time he felt the AI was "thinking." However, Brockman didn't reveal which physicist or which problem, so the case can't be independently verified. [Note: in public information, mathematician Terence Tao published a paper in March 2026 whose core proof was completed by ChatGPT Pro, but that was in mathematics and doesn't directly correspond to the case described here.]

The Super App: a grand vision, but who answers the Microsoft question?

Having cut video generation, OpenAI is placing its bet elsewhere. Kantrowitz asked what the Super App actually is.

Brockman said: integrating ChatGPT (chat), Codex (coding agent), and the browser into one app.

The computer should adapt to the person, not the other way around.

He compared the Super App to a laptop: is your laptop for personal use or work use? Both. The Super App is the same — both a personal assistant and a work tool. The personal side includes everything ChatGPT already does, plus deeper memory and context capabilities. ChatGPT already has a feature called Pulse (personalized content push) that delivers content daily based on what the AI knows about you.

But Brockman said what users see is just "the tip of the iceberg." What matters more is unifying the underlying technology: over the past two years, AI development has shifted from "just looking at the model" to "looking at the whole system" — how the model gets context, how it connects to the outside world, what actions it can take. OpenAI internally had multiple separate implementations of these things; now they're being consolidated into one general-purpose "AI layer." In general, users shouldn't need a dedicated finance AI or legal AI — the Super App should be general enough. If this vision is realized, it could genuinely redefine what an AI product looks like.

The problem is that Microsoft 365 Copilot is already doing almost exactly the same thing — embedding AI into workflows, unifying entry points, connecting various tools. OpenAI is simultaneously Microsoft's model supplier and its direct competitor — how is that tension resolved? Kantrowitz didn't ask, and naturally Brockman didn't answer. Add to that the fact that Sora just proved OpenAI's track record on consumer products isn't exactly stellar, and the "ship within a few months" timeline deserves a question mark.

Codex: how far is the road from programmers to everyone

If the Super App's ambitions are to be realized, Codex is one of the key pieces. It was originally designed for software engineers, but OpenAI has already seen a lot of spontaneous use by non-engineers internally. Brockman said Codex is fundamentally two things: a general-purpose agent framework (capable of calling tools) plus an AI that can write code. Connect it to spreadsheets or Word documents, and it can do knowledge work.

Specific examples: someone used Codex for video editing, and the AI wrote an Adobe Premiere plugin itself to automatically chapter the footage; someone on the internal communications team connected Codex to Slack and email to consolidate feedback. If true, these cases suggest AI's capability as "universal glue" connecting different tools is taking shape. But these are all internal use cases recounted by Brockman, with no external user verification, and no mention of whether they're publicly available.

Brockman said that for non-programmers, Codex's usability is currently low — if you hit an error during setup, a developer knows how to handle it, but an ordinary user is just stumped. But he believes the hardest part is already done (building a truly intelligent AI), and what remains is the "much easier" part: lowering the barrier to entry. Whether this holds depends on how you define "hard." For people who've built consumer products, the "last mile" is often harder than the ninety-nine miles before it.

He also mentioned OpenClaw founder Peter Steinberger joining OpenAI, as a signal of Codex expanding to non-programmers. Steinberger is an Austrian developer who, in late 2025, built an open-source AI agent that can autonomously manage email, order food, and control smart-home devices, reaching 196,000 GitHub stars and 2 million weekly active users. In February 2026, Altman announced Steinberger had joined OpenAI to lead "next-generation personal agents" — the role itself signals that OpenAI is taking "getting AI beyond the programmer crowd" seriously.

Catching up to Anthropic: admitting falling behind, claiming to have caught up

Kantrowitz asked: Anthropic already has Claude chat, Claude Cowork, and Claude Code — essentially having built its own Super App first. Did OpenAI only wake up to this after seeing what Anthropic did?

Brockman admitted that OpenAI has historically had the best competition results in coding, but underinvested in "last-mile usability." A model can be brilliant in coding competitions but has never seen the messy state of real-world codebases.

Around the middle of last year, OpenAI formed a dedicated team to tackle this: studying all the "dirty work" of real-world codebases and building training environments that simulate the messy scenarios disrupted in strange ways. He said that at this point they've caught up, and in head-to-head comparisons users increasingly prefer OpenAI. But he also said the front-end experience still lags.

Third-party data paints a different picture. According to reporting from firms like Synergy Research, Anthropic's enterprise market share rose from a low base in 2023 to about 40% by late 2025, while OpenAI's fell from 50% to about 27%. Claude Code has earned extremely strong word of mouth among software engineers. When Brockman says they've "caught up," what exactly is being compared — model capability or market share? He cleverly confined the battlefield to the former.

The bigger change is internal. Brockman said OpenAI used to treat research and deployment as "almost two separate things," and is now integrating them into one process — thinking about how the product will use it even while doing the research. This might be the single most substantively informative line in the whole conversation, hinting that OpenAI's prior organizational structure really did have problems.

"We won" is the scariest moment

Shifting from product competition to internal culture, Kantrowitz changed angle. He asked: OpenAI has been leading since 2022, and now competition has intensified — has the internal mood changed?

Brockman told a story. At the holiday party after ChatGPT launched, he sensed a "we won" atmosphere pervading the company. That was the moment that scared him most.

No, we're the challenger — we always have been.

This is standard tech-executive rhetoric — a $300 billion company calling itself a challenger, much like Google once called itself a startup; believe it or not as you like. But he followed up with something more grounded: when things are going well, don't believe people telling you how great you are; when things are going badly, don't believe people telling you how bad you are either.

Spud: it feels right, but where's the data?

Beyond culture, next comes technology. On the next-generation model Spud, Brockman confirmed it's a new pretrained foundation model, crystallizing "roughly two years of research results."

He first corrected a framing: what matters isn't any single model, but the whole "engine of progress." Model development happens in stages: first pretraining, training a base model on vast amounts of data — the most expensive stage; then post-training, using reinforcement learning to teach the model to solve various problems; and finally tuning behavior and usability. Spud is the result of a new round of pretraining, and will go through extensive post-training optimization afterward.

As for what Spud can actually do, Brockman gave no concrete metrics. He coined a concept called "big model smell" — when a model is truly strong enough, you can "smell" that quality. You ask it a question and it doesn't go off track, doesn't need you to explain things repeatedly. He said Spud will raise both the ceiling (solving more open-ended, longer-horizon problems) and the floor (usefulness on everyday tasks) at the same time.

No benchmark scores, no comparisons with competitors, no task-specific demo data. This purely qualitative description may be for confidentiality reasons—keeping models under wraps before launch is standard in the industry. But at a time when competitors are frequently showing off benchmarks, relying on "feel" alone has limited persuasive power. OpenAI deliberately withheld certain benchmark details when GPT-4 launched too, and later won back the narrative through actual product performance. Whether Spud can follow the same path depends on how it performs after launch.

December 2025: The Inflection-Point Narrative

Kantrowitz asked: what happened in December 2025? Coding agents seemed to move from theory to practical use.

Brockman said that model release pushed AI from being able to do "20% of your tasks" to "80%." It went from a "nice-to-have" to something where "you have to reorganize your workflow around AI."

He shared a personal test he'd run for years: having AI help him build a website—one that took him months to build back when he was learning to code. For most of 2025, it took four or five hours and multiple rounds of back-and-forth. In December, one prompt, done in one shot.

Slowly, slowly, slowly, then all at once.

This phrasing comes from Hemingway's The Sun Also Rises, where the original line was about "how did you go bankrupt." It's an apt way to describe the nonlinear growth of AI capability.

Another example: an engineer he works closely with, doing low-level systems engineering, couldn't use AI at all in the GPT-5.2 era. By 5.3, given a design document, the AI could implement the code, add monitoring metrics, run a profiler, and optimize on its own—producing exactly what the engineer wanted.

On the point that "public reaction to the GPT-5 launch was somewhat disappointing," Brockman said every release produces two camps: one that sees a night-and-day difference, and another whose use cases weren't intelligence-bottlenecked, so they don't notice a difference. The real shift is that people's mental models of AI capability update more slowly than the technology itself.

AGI: 70-80% There, But the Definition Is a "Vibe"

Nvidia CEO Jensen Huang recently said, "I think we've already achieved AGI." Kantrowitz asked Brockman if he agreed.

Brockman said AGI means different things to different people. Current technology is "very jagged": it already surpasses humans at tasks like writing code, but AI still fails at basic tasks that humans handle easily. Where you draw the line is "more of a vibe and a feeling than a science."

He gave his own self-assessment: "70-80% of the way there."

This self-contradictory statement actually reflects a dilemma facing the whole industry. If the definition of AGI isn't even scientific, more of a "vibe," then how was 70-80% calculated? Everyone is quantifying a concept they themselves admit they can't precisely define. Huang's definition of AGI is extremely narrow—he himself admits that 100,000 agents couldn't build another Nvidia; Lex Fridman has proposed "can it found and run a billion-dollar company" as a test. Brockman is more cautious than Huang, but falls into the same definitional trap.

He believes AGI will arrive within the next few years, with AI able to do nearly all the intellectual work you do on a computer, though capability will remain jagged.

Automated AI Researcher: Using AI to Accelerate AI

OpenAI plans to launch an automated AI researcher this fall.

Brockman calls the current stage "takeoff"—a term with a specific meaning in AI safety circles, referring to the stage where AI enters rapid recursive self-improvement. Using it in this context at least shows he's aware of its weight. His meaning has two layers: technically, the better AI gets, the faster it can accelerate its own improvement; in application, chip makers are ramping up investment and companies across every sector are exploring AI applications—all that energy is accumulating.

As for the AI researcher specifically, Brockman's definition is: taking the complete end-to-end work of an OpenAI research scientist and running it on silicon.

But he stressed this doesn't mean just letting AI run on its own. He used an analogy: it's like managing a junior researcher—leave them unsupervised too long and they'll wander down useless paths. Senior researchers provide direction and review; AI does the execution.

In the Agent Era, Human Responsibility Can't Be Outsourced

Kantrowitz quoted something Brockman had said before: when you use AI agents to get things done, "you become the CEO of a fleet of thousands of agents executing your goals and vision, but you don't understand exactly how everything gets solved. This new way of working might make you feel like you've lost the pulse on the problem."

Brockman said this cuts both ways, and the key is to "acknowledge the strengths, mitigate the weaknesses." AI agents give you enormous leverage, but ultimately there's a "responsible party." If your agent messes up a website and affects users, that's not the agent's fault—it's yours.

Kantrowitz asked: but how do you reconcile "losing the pulse" with "being responsible"?

Brockman used a home renovation analogy: you hire a general contractor to renovate your house, and there are details you genuinely don't need to worry about because you trust the professionals. But if something goes wrong, you should know about it. "You can't just blindly accept 'I can stop paying attention.' You need to actively stay on top of the problem."

The $110 Billion Compute Gamble: Revenue Engine or Casino Chips?

Back from technical roadmap and product vision to business reality.

Brockman compares compute to hiring salespeople: as long as the product sells and there's a scalable way to sell it, more salespeople means more revenue. Compute isn't a cost center, it's a revenue center. He recounted an internal conversation on the day ChatGPT launched:

"How much compute should we buy?" "All of it." "No, no, seriously, how much should we buy?" "No matter how much we build, I know it won't keep up with demand."

He said every year since then has proven that call right. ChatGPT currently has over 400 million weekly active users, and consumer subscription revenue keeps growing—so from the demand side, that judgment does hold up. But the challenge is that compute purchases have to be locked in 18 to 24 months in advance, meaning you have to predict the future with precision.

As for revenue mix, consumer subscriptions are still the largest source for now, but enterprise knowledge work is growing extremely fast. He doesn't think future revenue will simply split into "consumer" and "enterprise" categories—it's more that users have an entry point into a digital world, and revenue comes from that entry point.

Kantrowitz brought up Dario Amodei's suggestion that some competitors are burning cash "YOLO-style," turning the "risk dial too far."

Brockman's response: "I disagree. We've always been extremely deliberate, and we were among the first to recognize this trend and position ourselves ahead of it."

But after "I disagree," he didn't actually engage with Amodei's core argument: if revenue projections are off by even a year, the company could go bankrupt. That's a question that needs numbers, not attitude—and the numbers aren't entirely on Brockman's side: of the $110 billion in financing, only about $25 billion is confirmed as immediate cash; the rest is either compute (Nvidia's $30 billion is mostly GPUs) or comes with milestone conditions (Amazon's $35 billion). According to estimates from outlets like The Information, OpenAI's annual cash burn will reach $57 billion by 2027. The financing figure looks large, but the actual runway may only be 18-24 months. By comparison, Anthropic has likewise committed to $50 billion in infrastructure, but is pacing it strictly to revenue growth.

On whether even larger-scale pretraining is still needed, Brockman's position is clear: improving pretraining makes every downstream step easier. But he also said there's been an important shift over the past 24 months: no longer just chasing raw capability, but also considering inference efficiency at the same time, aiming to find the optimal point of "intelligence × cost."

Safety and Public Trust

On AI risk, Brockman proposed a "resilience" framework: many participants jointly develop the technology, while also building social infrastructure to ensure the technology heads in the right direction. He used electricity as an analogy: many people generate electricity, and electricity is dangerous too, but we've built a whole system of safety standards, regulation, and inspectors around it.

On negative public sentiment toward AI (YouGov data shows Americans who think AI has had a negative effect on society outnumber those who think it's had a positive effect by three to one), Brockman said the key is showing people concretely how AI helps them. He gave a medical example: a family's child had symptoms like headaches, and their insurance denied an MRI; using ChatGPT to research the symptoms, they found arguments that persuaded the insurance company, got the MRI, discovered a brain tumor, and treated it in time to save the child's life. He said stories like this "happen every day."

On the data center controversy, Brockman claimed that OpenAI's Abilene facility—one of the world's largest supercomputers, with over 100,000 GPUs deployed—uses "about as much water annually as a single household." This claim is questionable. For comparison, Microsoft's 2024 environmental report disclosed global data center water use of about 7 billion liters annually, and Google's figures are on the same order of magnitude. Even with a state-of-the-art air-cooling setup, the heat load from 100,000 GPUs far exceeds what a single household's water use could plausibly handle, and this claim has not been independently verified. On electricity, he pledged that OpenAI would "foot the bill itself" rather than driving up local residents' electricity rates, citing North Dakota as an example where the arrival of data centers actually helped lower electricity costs—a claim likewise lacking third-party verification.

Political Donations: The $25 Million Controversy

On the controversy over his $25 million donation to MAGA Inc, Brockman said he and his wife have also donated to bipartisan super PACs. He describes himself as a "single-issue donor" focused on supporting politicians who "genuinely embrace AI technology." But MAGA Inc. is a PAC that broadly supports Trump's agenda, not a dedicated tech-policy organization—there's a clear tension between the "single-issue" framing and the recipient of the donation. Brockman also donated $25 million to Leading the Future, a bipartisan PAC promoting AI development. The $25 million was the largest single donation MAGA Inc. received in the second half of 2025, and it sparked the #QuitGPT boycott movement in early 2026.

What This Interview Really Reveals

Unpacking the narrative Brockman built in this interview, you can see OpenAI facing pressure on three fronts.

The first is product pressure. Cutting Sora did stop the bleeding, and ChatGPT remains the world's largest AI product by user count, but the next bet, the Super App, faces a far harsher competitive environment than Sora did. Microsoft 365 Copilot on the enterprise side, Anthropic on the developer side, Google on the consumer side—OpenAI has to build a "unified entry point" while being squeezed from all three directions. Brockman's response to this is "ship it within months," but the lesson of Sora is that having the technology doesn't mean you can execute a good product.

The second is a communication bind. Throughout the interview, whenever Brockman was asked the hardest questions, he repeatedly substituted emotional language for data. Spud's capabilities are judged by "feel," AGI progress by "vibe," and the evidence for catching up to Anthropic is an undisclosed internal comparison. This reflects a structural contradiction facing OpenAI right now: as a company that just raised $110 billion, it needs to prove to investors that it's technologically ahead, but it can't leak specific metrics before a product launch. So "feel" becomes the only safe mode of expression. The problem is that feelings can't substitute for facts, especially when competitors are speaking in benchmarks and market share.

The third is scale risk. The $110 billion financing figure sounds intimidating, but it's less secure once you break it down. Brockman's rebuttal to Amodei's "YOLO" criticism stopped at "I disagree," never directly addressing the core question: what happens if demand growth comes in a year slower than expected? He himself said compute purchases have to be locked in 18-24 months ahead—which happens to be exactly the runway analysts estimate OpenAI actually has.

Brockman's advice to AI skeptics is "just try it first." But whether a $110 billion bet is the right call isn't a question a personal trial experience can answer. Over the next six months, whether Spud delivers on its promises and whether the Super App can ship while squeezed between Microsoft and Anthropic will determine whether this narrative turns out to be vision or spin.

Original video: https://www.youtube.com/watch?v=J6vYvk7R190

Greg Brockman 坐进 Big Technology Podcast 的演播室时,OpenAI 刚做完三件大事:砍掉 Sora,收到 1100 亿美元融资,完成下一代模型 Spud 的预训练。砍一个明星产品、拿一笔天价融资、押注一个新模型,三件事同时发生,每一件都需要解释。接下来 90 分钟里,这位 OpenAI 联合创始人兼总裁要完成一项高难度任务:把每一个看似被动的决策重新包装为前瞻性的战略选择。

他做到了吗?大部分时候做到了,但叙事的缝隙里藏着一些值得细看的线索。

Brockman 过去 18 个月主要负责公司的“规模”(Scale)部门,管 GPU 基础设施、数据中心和供应链。这场访谈覆盖的话题跨度极大:从 Sora 为什么被砍,到 Super App 长什么样,到下一代模型能做什么,到 1100 亿美元怎么花,到他为什么给 MAGA Inc 捐了 2500 万美元。

原始视频:https://www.youtube.com/watch?v=J6vYvk7R190

要点速览

  1. 砍掉 Sora 是“聚焦技术主线”还是止血? Brockman 说 Sora 和 GPT 是技术树的不同分支,算力有限只能选一条。但 Sora 日烧百万美元、用户跌破 50 万的现实他没提。
  2. Super App 将合并 ChatGPT、编程 Agent Codex 和浏览器 Atlas 为一个统一应用,预计未来几个月内交付,但这个愿景和微软 365 Copilot 有直接重叠,访谈中没人问这个问题。
  3. 下一代模型 Spud 凝聚了约两年的研究突破,Brockman 称其将同时提升 AI 能力的天花板和地板,但能力描述全是感性的,没有具体指标。
  4. AGI 进度自评 70-80%,但 Brockman 同时说 AGI 的定义“更像是一种氛围”,无法精确定义的东西如何精确量化,这个矛盾他没有解决。
  5. 承认在编程工具的“最后一公里可用性”上曾落后于 Anthropic,声称已追赶上来,第三方数据显示 Anthropic 企业市场份额仍在快速增长。
  6. 回应 Anthropic CEO Dario Amodei 对 OpenAI 基础设施投资“过于激进”的批评:“我不同意。” 但没有回应 Amodei 关于“收入预判偏差一年就可能破产”的具体论证。

砍掉 Sora:技术选择还是商业止血?

Kantrowitz 开场就问:OpenAI 在消费者市场领先,为什么突然把资源从视频生成转走?

Brockman 给出的解释框架和外界猜测相当不同。他说 Sora 的模型和 GPT 系列是“技术树上完全不同的分支”,构建方式根本不同。在算力有限的世界里,同时推进两条技术路线代价极大。他用了一个向量比喻:

随机向量之和为零,对齐方向才能前进。

力量分散就互相抵消,集中才能产生动能。他认为深度学习领域的问题不是机会不够,恰恰是太多,分散投入等于哪个方向都走不远。

这套解释在技术逻辑上自洽,但只讲了故事的一半。据 WSJ 调查和 TechCrunch 报道,Sora 在商业上已经失败:用户量从峰值一百万迅速跌至不足 50 万,每天烧掉约 100 万美元算力成本。据 WSJ 报道,Disney 原计划投资 10 亿美元并将角色授权给 Sora 的交易也随之取消,Disney 在公开宣布前不到一小时才得知消息。关停 Sora 既有技术聚焦的逻辑,也有商业止血的现实压力,Brockman 选择只谈前者

他强调,这不是“从消费者转向企业”,而是必须做选择。他认为当前最重要的两个应用是:一个了解你、与你的目标对齐的个人助手,和一个能替你解决难题的 AI。光是这两件事,OpenAI 现有的算力都不够用。

一个技术细节:ChatGPT 里的图像生成功能不受影响。图像生成基于 GPT 架构,跟文本、语音同属一条技术路线;Sora 用的是扩散模型(diffusion model),是另一棵树上的东西。Sora 研究团队没有解散,而是转向了机器人领域的世界模拟研究。

Kantrowitz 又问:Google DeepMind 的 Demis Hassabis 认为最接近 AGI 的不是文本模型,而是图像生成器,因为它必须理解物体之间的交互和世界运作方式。OpenAI 放弃这条路线,会不会押错赛道?

Brockman 说这是真实的风险:”绝对是的。在这个领域你必须做选择,必须下注。” 但他话锋一转,说 OpenAI 之所以有信心,是因为”已经看到了通往 AGI 的路径”。他举了一个例子:一位物理学家把一个长期未解的问题交给 OpenAI 的模型,12 小时后拿到了解答,那位物理学家说这是他第一次觉得 AI 在”思考”。不过 Brockman 没有透露具体是哪位物理学家和哪个问题,这一案例无法独立验证。【注:公开信息中,数学家 Terence Tao 在 2026 年 3 月发表了一篇由 ChatGPT Pro 完成核心证明的论文,但那属于数学领域,与此处描述不直接对应。】

Super App:愿景宏大,但谁来回答微软问题?

砍掉了视频生成,OpenAI 把赌注押在了另一个方向。Kantrowitz 问 Super App 到底是什么。

Brockman 说:把 ChatGPT(聊天)、Codex(编程 Agent)和浏览器整合到一个应用里

电脑本该迁就人,而不是人迁就电脑。

他把 Super App 比作笔记本电脑:你的笔记本是私人用还是办公用?两者都是。Super App 也一样,既是个人助理也是工作工具。个人场景包括 ChatGPT 现有的一切用途,加上更深的记忆和上下文能力。ChatGPT 里已经有一个叫 Pulse 的功能(个性化内容推送),每天根据 AI 对你的了解推送内容。

但 Brockman 说用户看到的界面只是“冰山一角”。更重要的是底层技术的统一:过去两年 AI 的发展已从“只看模型”变成“看整个系统”,模型如何获取上下文、如何连接外部世界、能执行什么操作。这些东西 OpenAI 内部曾有多套实现,现在要收归一统,形成一个通用“AI 层”。一般情况下用户不需要专门的金融 AI 或法律 AI,Super App 足够通用。这个愿景如果实现,确实有可能重新定义 AI 产品的形态。

问题在于,微软 365 Copilot 已经在做几乎一模一样的事,把 AI 嵌入工作流、统一入口、连接各种工具。OpenAI 既是微软的模型供应商又要做微软的直接竞品,这个张力怎么化解?Kantrowitz 没问,Brockman 自然也没答。加上 Sora 刚刚证明了 OpenAI 做消费者产品的记录并不光彩,”几个月内交付”的时间线值得打个问号。

Codex:从程序员到所有人的路还有多远

Super App 的野心如果要落地,Codex 是关键拼图之一。它最初为软件工程师设计,但 OpenAI 内部已经出现了大量非工程师的自发使用。Brockman 说 Codex 底层就是两样东西:一个通用 Agent 框架(能调用工具)加上一个会写代码的 AI。接上电子表格、Word 文档,就能做知识工作。

具体案例:有人用 Codex 做视频剪辑,AI 自己写了一个 Adobe Premiere 插件来自动分章节;内部通信团队的人把 Codex 接上 Slack 和邮箱来整合反馈。这些案例如果属实,说明 AI 作为“万能胶水”连接不同工具的能力正在成型。但它们都是 Brockman 口述的内部使用场景,没有外部用户验证,也没说是否已公开可用。

Brockman 说,对非程序员来说,Codex 现在的可用性很低,设置中碰到报错,开发者知道怎么处理,普通用户直接懵掉。但他认为最难的部分已经做完了(造出真正聪明的 AI),剩下的是“容易得多”的部分:降低门槛。这个判断是否成立,取决于你怎么定义“难”。对做过消费者产品的人来说,“最后一公里”往往比前面九十九公里还难。

他还提到了 OpenClaw 创始人 Peter Steinberger 加入 OpenAI 的事,作为 Codex 向非程序员扩展的信号。Steinberger 是奥地利开发者,2025 年底开发了一个能自主管理邮箱、订餐、控制智能家居的开源 AI Agent,GitHub 星标达 19.6 万,周活用户 200 万。2026 年 2 月,Altman 宣布他加入 OpenAI 负责“下一代个人 Agent”,这个职位本身就说明 OpenAI 在认真对待“让 AI 走出程序员圈子”这件事。

追赶 Anthropic:承认落后,声称追上

Kantrowitz 问:Anthropic 已经有了 Claude 聊天、Claude Cowork、Claude Code,相当于先做出了自己的 Super App。OpenAI 是不是看到 Anthropic 做了才醒悟过来的?

Brockman 承认 OpenAI 过去在编程领域一直有最好的竞赛成绩,但投入不够的是“最后一公里的可用性”。模型在编程竞赛里很聪明,但它从没见过真实世界里乱糟糟的代码库。

大约去年年中,OpenAI 组建了专门团队来解决这个问题:研究真实世界代码库的各种“脏活”,构建训练环境来模拟那些被奇怪的方式打断的混乱场景。他说此时此刻已经追上了,在正面对比中用户更倾向于选择 OpenAI。但他也说前端体验还落后。

第三方数据呈现的画面并不相同。据 Synergy Research 等机构报道,Anthropic 的企业市场份额从 2023 年的低位升至 2025 年底的约 40%,而 OpenAI 从 50% 降至约 27%。Claude Code 在软件工程师群体中获得了极高的口碑。Brockman 说“已经追上”时到底在比什么,模型能力还是市场份额?他巧妙地把战场限定在了前者

更大的变化在内部。Brockman 说 OpenAI 过去把研究和部署当成”几乎分开的两件事”,现在要整合成一个流程,做研究时就想着产品怎么用。这或许是整段话里最有实质信息量的一句,它暗示 OpenAI 此前的组织架构确实有问题。

“我们赢了”是最可怕的时刻

从产品竞争到内部文化,Kantrowitz 换了个角度。他问,从 2022 年以来 OpenAI 一直领先,现在竞争激烈了,内部氛围有变化吗?

Brockman 讲了一个故事。ChatGPT 发布后的节日派对上,他感受到公司弥漫着“我们赢了”的气氛。那是他最害怕的一刻

不,我们是挑战者,一直都是。

这是科技公司高管的标准话术,估值 3000 亿美元的公司自称挑战者,和 Google 当年说自己是创业公司一样,信不信随你。但他后面补了一句更实在的话:好的时候别信别人说你多好,坏的时候也别信别人说你多差

Spud:感觉对了,但数据在哪?

文化之外,接下来是技术。关于下一代模型 Spud,Brockman 确认它是一个新的预训练基础模型,凝聚了“大约两年的研究成果”。

他先纠正了一个思维框架:重要的不是任何一个模型,而是一整个“进步引擎”(engine of progress)。模型开发分几步:先做预训练,在海量数据上训练基础模型,这是最昂贵的阶段;然后做后训练,通过强化学习让模型学会解决各种问题;最后调整行为和可用性。Spud 是新一轮预训练的成果,后续还会经过大量后训练优化。

至于 Spud 能做什么,Brockman 没有给出任何具体指标。他造了一个概念叫 “big model smell”(大模型的味道),当模型真正够强时,你能“闻到”那种质感。你问它问题,它不会答偏,不需要你反复解释。他说 Spud 将同时提升天花板(解决更开放、时间跨度更长的问题)和地板(日常任务的实用性)

没有 benchmark 分数、没有和竞品的对比、没有具体任务的演示数据。 这种纯感性描述可能出于保密考虑,行业里模型发布前保密是常态。但在竞争对手频繁亮出 benchmark 的当下,光靠“感觉”说服力有限。OpenAI 此前在 GPT-4 发布时也曾刻意回避部分 benchmark 细节,后来靠实际产品表现赢回了话语权。Spud 能不能走同样的路,要看发布后的表现。

2025 年 12 月:拐点叙事

Kantrowitz 问:2025 年 12 月发生了什么?编程 Agent 似乎从理论走向了实用。

Brockman 说那次模型发布让 AI 从“能完成你 20% 的任务”跃升到了“80%”。从“锦上添花”变成了“你必须围绕 AI 重新组织工作流程”。

他分享了一个持续多年的个人测试:让 AI 帮他建一个网站,这个网站他当年学编程时花了几个月才做出来。2025 年大部分时间里需要四五个小时、多轮对话。12 月,一次提示,一次完成。

慢慢地、慢慢地、慢慢地,然后一下子全变了。

这个表达出自海明威《太阳照常升起》,原文说的是“你是怎么破产的”。用来形容 AI 能力的非线性增长,倒也贴切。

另一个案例:一位和他密切合作的工程师做底层系统工程,在 GPT 5.2 时代完全用不了 AI。到了 5.3,给 AI 一份设计文档,它能实现代码、添加监控指标、运行性能分析器、自行优化,产出就是工程师想要的东西。

对于“GPT-5 发布后公众反应有些失望”,Brockman 说每次发布都有两群人:一群觉得天壤之别,另一群用的场景不是智能瓶颈,感受不到差异。真正的转变是人们对 AI 能力的心智模型更新得比技术本身慢

AGI:70-80%,但定义是个“氛围(Vibe)”

Nvidia CEO 黄仁勋最近说“我认为我们已经实现了 AGI”。Kantrowitz 问 Brockman 是否同意。

Brockman 说 AGI 对不同人有不同定义。当前的技术“非常参差不齐”(very jagged):在写代码这类任务上已经超越人类,但某些人类轻松完成的基本任务 AI 仍然做不好。画线画在哪里,“更像是一种氛围和感觉,而不是科学”

他给了一个自我评估:70-80% 到了

这个自相矛盾的表述其实折射了整个行业的困境。如果 AGI 的定义连科学都算不上,更像“氛围”,那 70-80% 是怎么算出来的?每个人都在量化一个他们承认无法精确定义的概念。Huang 的 AGI 定义极窄,他自己也承认 10 万个 Agent 也造不出 Nvidia;Lex Fridman 曾提出“能否创办和运营一家 10 亿美元公司”作为判据。Brockman 比 Huang 谨慎,但陷入了同样的定义困境。

他认为未来几年内 AGI 会实现,AI 将能完成几乎所有你在电脑上做的智力工作,但能力仍然是参差不齐的。

自动化 AI 研究员:让 AI 加速 AI

OpenAI 计划今年秋季推出一个自动化 AI 研究员。

Brockman 把当前阶段称为“起飞”(takeoff),这个词在 AI 安全领域有特定含义,指 AI 进入快速递归自我改进的阶段。他在这个语境下用它,至少说明他清楚这个词的分量。他的意思分两层:技术上,AI 越好就越能加速自身改进;应用上,芯片厂商加大投入、企业在各领域摸索 AI 应用,所有能量在积累。

具体到 AI 研究员,Brockman 的定义是:把 OpenAI 一个研究科学家的完整端到端工作搬到硅片上运行

但他强调这不意味着放手让 AI 自己跑。他打了个比方:就像带初级研究员,放太久不管他就会走上没用的路。高级研究员提供方向和审查,AI 负责执行

Agent 时代,人类的责任不能外包

Kantrowitz 引用了 Brockman 此前说过的话:当你用 AI Agent 做事时,“你变成了一支成千上万 Agent 舰队的 CEO,它们在执行你的目标和愿景,但你并不了解每件事具体怎么解决的。这种新的工作方式可能让你觉得失去了对问题的脉搏感。”

Brockman 说,这是好坏参半的,关键是“承认优势、减轻弱点”。AI Agent 给人巨大的杠杆,但归根到底有一个“负责任的当事人”。你的 Agent 搞砸了网站影响了用户,那不是 Agent 的错,是你的错。

Kantrowitz 问:但“失去脉搏感”和“负责任”怎么兼得?

Brockman 用装修做比方:请总包装修房子,有些细节你确实不用操心,因为你信任专业人士。但如果出了问题,你应该知道。“你不能盲目接受'我可以不管了'。你需要主动保持对问题的把控。”

1100 亿算力豪赌:收入中心还是赌场筹码?

从技术路线和产品愿景回到商业现实。

Brockman 把算力类比为雇销售人员:只要产品能卖出去、有可扩展的销售方式,销售人员越多收入就越高。算力不是成本中心,是收入中心。他讲了 ChatGPT 发布当天的内部对话:

“该买多少算力?”“全部。”“不不不,说真的,该买多少?”“不管我们怎么建,我知道都跟不上需求。”

他说从那以后每一年都被证明是对的。ChatGPT 目前周活用户超过 4 亿,消费者订阅收入持续增长,从需求端看这个判断确实成立。但挑战在于算力采购需要提前 18 到 24 个月锁定,必须精准预判未来。

收入结构方面,消费者订阅目前仍是最大来源,但企业知识工作领域增长极快。他不认为未来收入会简单分为“消费者”和“企业”两类,更像是用户拥有一个数字世界的入口,收入来自这个入口。

Kantrowitz 提到了 Dario Amodei 暗示某些竞争对手在“YOLO 式”烧钱,把“风险旋钮拧得太远”。

Brockman 的回应是:“我不同意。我们一直非常深思熟虑,是最早意识到这一趋势并提前布局的。”

“我不同意”之后他没有回应 Amodei 的核心论点:如果收入预判偏差一年,公司就可能破产。这个问题不是态度能回答的,需要数字。而数字并不完全站在 Brockman 这边:1100 亿融资中只有约 250 亿是确定的即时现金,其余要么是算力(Nvidia 的 300 亿大部分是 GPU)、要么有里程碑条件(Amazon 的 350 亿)。据 The Information 等媒体估计,OpenAI 到 2027 年的年度现金消耗将达 570 亿美元。融资数字虽大,实际跑道可能只有 18-24 个月。相比之下,Anthropic 同样承诺了 500 亿美元基础设施,但严格按收入增长节奏推进。

关于是否还需要做更大规模的预训练,Brockman 的立场明确:改进预训练会让所有下游步骤都更容易。但他也说过去 24 个月有一个重要转变:不再只追求原始能力,开始同时考虑推理效率,目标是找到”智能 × 成本”的最优解。

安全与公众信任

谈到 AI 风险,Brockman 提出了“韧性”(resilience)框架:很多参与者共同开发技术,同时建立社会基础设施来确保技术走向正确。他拿电力做类比:很多人生产电力,电力也有危险,但我们围绕它建立了安全标准、监管、检查员等一整套体系。

关于公众对 AI 的负面态度(YouGov 数据显示认为 AI 对社会负面影响的美国人是认为正面影响的三倍),Brockman 认为关键是向人们展示 AI 如何具体帮助他们。他举了一个医疗案例:一个家庭的孩子出现头痛等症状,医保拒绝了 MRI 检查,他们用 ChatGPT 研究症状后找到了说服保险公司的论据,拿到了 MRI,发现是脑瘤,及时治疗救了孩子一命。他说这类故事“每天都在发生”。

数据中心争议方面,Brockman 声称 OpenAI 在 Abilene 的设施(部署超过 10 万 GPU 的世界最大超算之一)全年用水量“相当于一个家庭”。这个说法令人存疑。作为参照,微软在 2024 年环境报告中披露其全球数据中心年用水量约 70 亿升,Google 的数字也在同一量级。即便 Abilene 设施采用了最先进的空气冷却方案,10 万块 GPU 的散热量也远超一个家庭用水所能处理的范围,且这一说法未经任何独立验证。在电力方面,他承诺 OpenAI 会“自己买单”,不推高当地居民电价,并举了北达科他州的例子说数据中心到来反而帮助降低了电费,这一说法同样未找到第三方验证。

政治捐款:2500 万美元的争议

关于他向 MAGA Inc 捐赠 2500 万美元的争议,Brockman 说他和妻子也给两党超级 PAC 捐了款。他自称是“单一议题捐赠者”,关心的是支持那些“真正拥抱 AI 技术的政治家”。但 MAGA Inc. 是一个广泛支持 Trump 议程的 PAC,不是专门的科技政策组织,“单一议题”的定位与捐款对象之间存在明显张力。Brockman 同时向 Leading the Future(一个促进 AI 发展的两党 PAC)捐了 2500 万美元。2500 万美元是 MAGA Inc. 2025 年下半年收到的最大单笔捐款,这笔捐款在 2026 年初引发了 #QuitGPT 抵制运动。

这场访谈真正透露了什么

拆解 Brockman 在这场访谈中构建的叙事,可以看到 OpenAI 正面对三层压力

第一层是产品压力。 砍 Sora 确实止了血,ChatGPT 仍然是全球用户量最大的 AI 产品,但下一个赌注 Super App 面临的竞争环境比 Sora 更残酷。微软 365 Copilot 在企业端、Anthropic 在开发者端、Google 在消费者端,OpenAI 要在三者的夹击中做一个“统一入口”。Brockman 对此的回应是“几个月内交付”,但 Sora 的前车之鉴说明,有技术不等于能做好产品。

第二层是表达困境。 整场访谈中,Brockman 在被问到最硬的问题时,反复用感性语言代替数据。Spud 的能力靠“闻”,AGI 的进度靠“氛围”,追上 Anthropic 的证据是未公开的内部对比。这反映了 OpenAI 当前的一个结构性矛盾:作为一家刚融完 1100 亿美元的公司,它需要向投资人证明技术领先,但又不能在产品发布前泄露具体指标。于是“感觉”成了唯一安全的表达方式。问题是,感觉不能替代事实,尤其当竞争对手在用 benchmark 和市场份额说话的时候。

第三层是规模风险。 1100 亿美元的融资数字听起来吓人,但拆开看并没有那么安全。Brockman 对 Amodei“YOLO”批评的反驳停留在“我不同意”,始终没有正面回应核心问题:如果需求增长比预期慢一年会怎样?他自己说算力采购需要提前 18-24 个月锁定,这恰好也是分析师估计的 OpenAI 实际跑道。

Brockman 给 AI 怀疑者的建议是“先试试再说”。但 1100 亿美元的赌注正确与否,不是个人试用体验能回答的问题。接下来 6 个月,Spud 是否兑现承诺、Super App 能否在微软和 Anthropic 的夹击中交付,将决定这套叙事究竟是远见还是话术。

原始视频:https://www.youtube.com/watch?v=J6vYvk7R190


See all posts