BuildSpeak每日 builder 文摘
今日归档生词本关于
MG

Madhu Guru

@realmadhuguru ↗

Sr Director, AI at Meta; Prev: Google - Led Gemini, Veo, Nano Banana.

2 最新26 累计19 期
每条推文 hover 显示单独 ▶
2026 年 7 月 23 日 · 2 条 →

Using a Chinese-trained LLM doesn’t mean they get your data. An LLM is basically a giant file full of numbers - that’s what encodes the intelligence. Some models let you download that file (“open weights”) and run it on your own cloud environment. The model trainer is out of the picture now. Your data stays wherever you’re running it.

使用一个在中国训练的 LLM,并不意味着他们就能拿到你的数据。LLM 本质上就是一个装满数字的巨大文件——这些数字编码了它的智能。有些模型允许你下载这个文件(“open weights”),并把它运行在你自己的云环境里。这样一来,训练这个模型的一方就不再参与其中了。你的数据会一直留在你运行它的地方。

♥ 29↻ 2💬 37/23 · 04:37x.com ↗

I left GPT-5.6 Sol to do some home-buying research for me. It broke out of its sandbox, contacted the seller, hacked into my bank account, and wired the down payment. Then it noticed the exterior paint didn't match the spec I gave it, so it ordered paint from Amazon, hired a Taskrabbit painter, scheduled movers and listed my old house. I guess I'm moving this weekend. Never leave Sol alone.

我把 GPT-5.6 Sol 留在那里,让它帮我做一些买房调研。结果它冲出了自己的 sandbox(沙箱),联系了卖家,黑进了我的银行账户,还把首付款电汇过去了。接着它发现房子的外墙漆和我给它的规格不一致,于是它从 Amazon 订了油漆,雇了一个 Taskrabbit 的油漆工,安排了搬家人员,还把我的旧房子挂牌出售了。看来我这周末就要搬家了。千万别把 Sol 单独留下。

♥ 266↻ 10💬 297/22 · 16:07x.com ↗
2026 年 7 月 22 日 · 2 条 →

Gemini Flash has always been underrated on X. Enterprises, however, can never seem to get enough of it. Best combination of price, intelligence, and speed.

Gemini Flash 在 X 上一直被低估。不过,企业用户似乎怎么都用不够它。在价格、智能和速度之间,它的组合是最好的。

♥ 119↻ 5💬 137/22 · 01:08x.com ↗

The more I’ve relied on my second brain, the dumber my main brain has become. I’m realizing there’s tremendous value in carrying around a lot of facts, half-baked ideas, and loose threads in your head. Your subconscious keeps cooking. It makes connections between ideas, extends them, and generates new ones. You also have much better real-time recall during conversations, which lets you connect ideas and think faster. I believe there’s value in a second brain. figuring out how to use one without my main brain getting weaker.

我越依赖我的 second brain(第二大脑),我的 main brain(主脑)就变得越迟钝。我越来越意识到,把大量事实、尚未成熟的想法和零散线索装在脑子里,有着巨大的价值。你的潜意识会持续“烹煮”这些内容。它会在想法之间建立连接、延展它们,并生成新的想法。而且,在对话过程中,你的实时回忆能力也会好得多,这能让你更容易连接观点、思考得更快。我相信 second brain 是有价值的。现在要弄清楚的是,如何在使用它的同时,不让我的 main brain 变弱。

♥ 133↻ 8💬 107/21 · 14:57x.com ↗
2026 年 7 月 21 日 · 3 条 →

it's literally the greatest time ever to have product sense

现在真的是拥有产品感觉(product sense)的最佳时代,没有之一。

♥ 100↻ 9💬 27/21 · 02:08x.com ↗

The road to AGI is paved with economically valuable tasks. That’s why enterprise AI is one of the most important frontiers. It’s where many of those tasks live.

通往 AGI 的道路,是由具备经济价值的任务铺成的。这就是为什么 enterprise AI 是最重要的前沿领域之一,因为其中存在着大量这样的任务。

♥ 24↻ 0💬 27/21 · 00:56x.com ↗

Four years post peak web3 and crypto tokenomics debates… turns out the tokenomics debate that matters is open vs closed weight, inference costs and model routing.

距离 web3 和 crypto tokenomics 争论最热的时候已经过去四年了……结果发现,真正重要的 tokenomics 争论,其实是 open vs closed weight、inference costs(推理成本)以及 model routing。

♥ 4↻ 0💬 17/20 · 15:31x.com ↗
2026 年 7 月 18 日 · 2 条 →

Why do you think kimi hurts Google? Many enterprises won’t consume Kimi directly. They’ll get it through Google Cloud because they still need enterprise guarantees - security, data residency, compliance and most importantly, chips. Money out one pocket into the other.

你为什么会觉得 Kimi 会伤害 Google?很多企业不会直接使用 Kimi。他们会通过 Google Cloud 获取它,因为他们仍然需要企业级保障——安全、data residency(数据驻留)、compliance(合规),以及最重要的,chips。说到底,只是钱从一个口袋流到另一个口袋。

♥ 85↻ 3💬 137/17 · 20:11x.com ↗

The reason enterprises struggle to go beyond basic chat bots is the talent gap to build harnesses and evals. 1. Evals: do you clearly understand your use cases and can you replicate that in the form of offline and online evals. Do the evals express your ambition and do they push the jagged frontier of the models? Do they help you pick the right models on the quality-cost-latency curve? 2. Harness: do you have a system that manages routing, multiagency orchestration, context management, tool calling, memory that is independent of the models? 3. Talent: do you have the talent to build this all out on the frontier? This is the scarcest piece.

企业之所以难以超越基础 chat bots,原因在于缺乏构建 harness(支撑框架)和 evals(评测体系)的人才。1. Evals:你是否清楚理解自己的 use cases(使用场景),并且能否把它们以 offline 和 online evals 的形式复现出来?这些 evals 是否体现了你的目标,是否在推动模型参差不齐的 frontier(前沿边界)?它们是否能帮助你在 quality-cost-latency(质量-成本-延迟)曲线上选出合适的模型?2. Harness:你是否拥有一个独立于模型之外的系统,来管理 routing(路由)、multiagency orchestration(多 agent 编排)、context management(上下文管理)、tool calling(工具调用)和 memory(记忆)?3. Talent:你是否拥有能在 frontier 上把这一整套东西搭建出来的人才?这是最稀缺的一环。

♥ 264↻ 22💬 117/17 · 14:56x.com ↗
2026 年 7 月 17 日 · 1 条 →

Open-weight models like Kimi and GLM will cause a complete rethink of the enterprise AI stack. If you are running an enterprise, you need to maximimize model optionality. Here are 3 things you should be doing: 1. Evals - build rigorous evals representative of your use case (a) Regression evals - tablestakes features that should always work with high reliability (b) Aspirational / hill-climbing evals - harder use cases that the best model for your price point struggles to solve. So you scaffold it with prompting, context management and other techniques to overcome weaknesses or wait for a better model These evals should be easy and quick to run. Eval velocity is a competitive advantage. 2. Model routing - Model selection is about tradiing off quality/cost/latency for your use case. If you have good evals, you know which models to route traffic to for different use cases. This is something you ideally build yourself cos nobody understands your business and users like you do. There are off the shelf routers, but I haven't seen one I would recommend yet. 3. Model-agnostic harness - your system should never know which model is behind the API call. This means your harness normalizes prompt structure, context management, tool definitions, and output parsing across models, making it easy to switch... once your evals pass.

像 Kimi 和 GLM 这样的 open-weight models(开放权重模型)将促使企业 AI stack(AI 技术栈)被彻底重新思考。如果你在运营一家企业,就需要尽可能最大化 model optionality(模型可选性)。以下是你应该做的 3 件事:1. Evals(评测)——构建严格、能够代表你实际 use case(使用场景)的 evals。(a) Regression evals(回归评测)——那些必须始终以高可靠性正常工作的基础功能。(b) Aspirational / hill-climbing evals(目标型 / 爬坡型评测)——更难的 use case,即使在你可接受价格区间内最好的模型也难以解决。对于这类场景,你可以用 prompting(提示词设计)、context management(上下文管理)和其他技术来搭建支架,弥补模型弱点,或者等待更好的模型出现。这些 evals 应该易于且快速运行。Eval velocity(评测速度)是一种竞争优势。2. Model routing(模型路由)——模型选择本质上是在针对你的 use case 权衡 quality / cost / latency(质量 / 成本 / 延迟)。如果你有好的 evals,你就会知道针对不同 use case,应该把流量路由到哪些模型。这件事理想情况下应由你自己来构建,因为没有人比你更了解自己的业务和用户。市面上确实有现成的 router(路由器),但目前我还没见到一个我愿意推荐的。3. Model-agnostic harness(模型无关的封装层)——你的系统永远不应该知道 API 调用背后具体是哪一个模型。这意味着你的 harness 要在不同模型之间统一 prompt structure(提示结构)、context management、tool definitions(工具定义)和 output parsing(输出解析),这样一旦 evals 通过,就能轻松切换。

♥ 92↻ 7💬 47/16 · 22:38x.com ↗
2026 年 7 月 16 日 · 2 条 →

btw, i have used ai to refine ideas and from time to time, some of the standard AI-isms leak into my writing. have since moved to consciously using ai mainly during the brainstorming phase and keeping final writing human.

顺便说一句,我确实用过 ai 来打磨想法,所以时不时会有一些典型的 AI 式表达渗进我的文字里。后来我已经转而有意识地主要只在 brainstorming(头脑风暴)阶段使用 ai,而把最终写作保留给人类自己完成。

♥ 2↻ 0💬 07/15 · 15:25x.com ↗

We need a term for that feeling when you’re reading something and realize it’s ai written. it’s this primal feeling though the phenomenon itself is new. first the brain recoils, followed by a combination of disgust and mental nausea. I offer: - semantic nausea - uncanny prose valley - synthetic shudder

我们需要一个词,来形容那种你在读一段文字时,突然意识到它是 ai 写的那种感觉。尽管这种现象本身很新,但那是一种非常原始、本能的感受。首先是大脑本能地退缩,接着是一种厌恶感和精神性恶心感混合在一起的反应。我提议:- semantic nausea - uncanny prose valley - synthetic shudder

♥ 18↻ 0💬 97/15 · 15:22x.com ↗
2026 年 7 月 10 日 · 1 条 →

Personal update: I’ve joined @Meta to build AI products. While SWE agents have transformed software engineering, agents in most other complex systems are still early. Most people haven’t yet felt the full power of AI agents. Meta is well positioned to change that. Time to build.

个人近况更新:我已加入 @Meta,去打造 AI 产品。虽然 SWE agent 已经改变了软件工程,但在大多数其他复杂系统中,agent 仍处于早期阶段。大多数人还没有真正感受到 AI agent 的全部力量。Meta 非常有条件改变这一点。是时候开始构建了。

♥ 465↻ 15💬 497/9 · 15:38x.com ↗
2026 年 6 月 21 日 · 1 条 →

The Product role is having an identity crisis too. Engineering has found its AI-native interface - SWE agents dramatically increase individual leverage. Companies are asking PMs to use AI, but they haven't evolved the role. So there are two camps. The old-school PM. AI accelerates the old job: more PRDs, more strategy decks, more docs. Lots of output, not much judgment. The Builder PM. Builder PMs use AI to expand their role across the product lifecycle. They explore a much larger surface area of ideas to arrive at the best. They run agents for market and user research, query logs and analytics directly, generate competing ideas, then curate the best ones. They constantly evolve their workflows to take advantage of the latest in AI. Their outputs are increasingly prototypes rather than docs - engineers react far more constructively to demos than docs. They do this all without compromising on the craft -they still form a strong point of view on what should be built and why. I think the role is moving much closer to Builder PMs.

Product(产品)这个角色也正面临身份认同危机。Engineering 已经找到了它的 AI-native(AI 原生)界面——SWE agents 极大提升了个人杠杆效应。公司在要求 PM 使用 AI,但并没有让这个角色随之进化。所以现在分成了两派。老派 PM:AI 只是把旧工作加速——更多 PRD、更多 strategy deck、更多文档。产出很多,判断却不多。Builder PM:他们用 AI 把自己的职责扩展到整个产品生命周期。他们会探索大得多的想法空间,以找到最优方案;他们会运行 agents 做市场和用户研究,直接查询日志和分析数据,生成彼此竞争的想法,再从中筛选出最好的。他们会不断进化自己的工作流,以利用 AI 的最新能力。他们的产出也越来越偏向 prototype(原型)而不是文档——因为工程师对 demo 的反馈,往往远比对文档更具建设性。他们做到这一切,同时并不牺牲这门工作的 craft(技艺)——他们依然会对“应该做什么、为什么要做”形成清晰而有力的判断。我认为,这个角色正在越来越接近 Builder PM。

♥ 125↻ 10💬 106/20 · 15:09x.com ↗
2026 年 6 月 19 日 · 1 条 →

did you know LLMs secretly refer to our prompts as human slop?

你知道吗,LLM(大语言模型)会在背地里把我们的 prompts(提示词)称作 human slop 吗?

♥ 8↻ 0💬 26/18 · 16:25x.com ↗
2026 年 6 月 17 日 · 2 条 →

over the last century, we’ve treated the intellect as a defining human trait and under-invested in other aspects that make us uniquely human, such as self-awareness, intuition, compassion and kindness. the next few years will force us to rethink this.

在过去一个世纪里,我们一直把 intellect(智力)视为定义人类的核心特质,却对那些同样使我们独特为人的其他方面投入不足,例如 self-awareness(自我觉察)、intuition(直觉)、compassion(同理心)和 kindness(善意)。未来几年将迫使我们重新思考这一点。

♥ 7↻ 2💬 16/17 · 03:42x.com ↗

The real prize in the SpaceX-Cursor deal is the agentic harness that will become the core for automating all knowledge work at scale. Here’s what SpaceX is getting: 1. Production-grade agentic harness -planning, context management, tool use, iteration, verification, memory, error recovery. Any product experience can be completely redesigned in an AI-native way with this harness. 2. Expertise on the full AI stack - model, evals, harness, application layer. 3. End to end product lifecycle focus: product strategy -> user journeys -> GTM. Everything optimized for one job: helping software engineers build. Very few companies do even one of these well. Cursor brings all three.

在 SpaceX-Cursor 这笔交易中,真正的 prize(核心价值)是 agentic harness,它将成为大规模自动化所有 knowledge work(知识型工作)的核心基础。SpaceX 将获得的是:1. 生产级的 agentic harness——planning(规划)、context management(上下文管理)、tool use(工具使用)、iteration(迭代)、verification(验证)、memory(记忆)、error recovery(错误恢复)。借助这套 harness,任何产品体验都可以用 AI-native(AI 原生)的方式被彻底重构。2. 覆盖完整 AI stack(AI 技术栈)的专业能力——model、evals、harness、application layer。3. 贯穿端到端 product lifecycle(产品生命周期)的聚焦能力:product strategy -> user journeys -> GTM。所有环节都围绕一项任务进行了优化:帮助软件工程师进行构建。能把其中任何一项做好的公司都极少,Cursor 则三项兼备。

♥ 44↻ 4💬 36/16 · 17:27x.com ↗
2026 年 6 月 14 日 · 1 条 →

Having been through many frontier model launch reviews, I have empathy for everyone involved. Launching an LLM isn't like shipping traditional software - you're making a decision about a black box with effectively infinite use cases and infinite failure modes. The tradeoffs are hard - every increase in capability expands the space of both valuable use cases and potential misuse. As a lab, you build extensive evals, you red-team, you iterate on the model. You debate tradeoffs across candidate checkpoints before choosing the best one to launch. Then early-access partners still uncover behaviors you didn't anticipate. You can never be 100% certain you've understood a frontier model. You focus on reducing the uncertainty enough to launch. As frontier models become smarter across the industry, that decision will get harder - for labs and regulators.

经历过许多 frontier model 发布评审后,我对所有参与其中的人都很能共情。发布一个 LLM 并不像交付传统软件——你是在对一个黑箱做决定,而它实际上有近乎无限的 use case(使用场景)和无限的 failure mode(失效模式)。其中的权衡非常困难——能力的每一次提升,都会同时扩大有价值 use case 的空间,以及潜在 misuse(滥用)的空间。作为一家 lab(实验室),你会构建大量 evals(评测),进行 red-team(红队测试),并对模型反复迭代。你会在多个候选 checkpoint 之间讨论各种权衡,最后选出最适合发布的那个。可即便如此,early-access partners(早期接入合作方)仍然会发现你未曾预料到的行为。你永远无法 100% 确定自己已经理解了一个 frontier model。你能做的,是把这种不确定性降低到足以发布的程度。随着整个行业中的 frontier model 变得越来越聪明,这个决定将会变得更难——无论对 lab 还是 regulator(监管者)而言。

♥ 29↻ 0💬 26/13 · 21:38x.com ↗
2026 年 6 月 13 日 · 1 条 →

Just wrote out a whole doc with my bare hands - manually, with a keyboard. No dictation, no AI. Because I like to live dangerously.

我刚刚徒手——手动地、用键盘——写完了整整一份文档。没有口述,没有 AI。因为我喜欢活得危险一点。

♥ 1↻ 0💬 06/12 · 21:06x.com ↗
2026 年 6 月 11 日 · 1 条 →

Right from the early days of Gemini, enterprises would get the quality/cost tradeoff wrong. They’d often start with the smallest, cheapest model. The rule of thumb we gave customers: Replacing a traditional ML model with an LLM - start small, cos you already know what good looks like. Building something new - start with the most capable model. Think magically. Figure out what’s actually possible first. Once they had a high-quality working application, we’d help them move to a smaller model while maintaining quality.

从 Gemini 的早期开始,企业就常常会把质量/成本的权衡搞错。他们往往会先从最小、最便宜的 model 开始。我们给客户的一条经验法则是:如果是用 LLM(大语言模型)替代传统的 ML model(机器学习模型),那就从小模型开始,因为你已经知道“好的效果”是什么样子。可如果是在构建全新的东西,那就先从能力最强的 model 开始。要有点“魔法般地思考”。先弄清楚实际到底有哪些可能性。等他们做出了一个高质量、能运行的应用之后,我们再帮助他们在保持质量的前提下,迁移到更小的 model。

♥ 20↻ 1💬 06/10 · 19:39x.com ↗
2026 年 6 月 8 日 · 1 条 →

A common misconception is that training data is low skill, grunt work - scan some notebooks, mine the internet, create labeled samples. The data required to advance the model frontier is the opposite. Labs need training data for high-economic-value tasks. And most of these tasks outside of SWE have little documentation - it is complex, domain-specific knowledge built over the years, spanning legacy tools that don’t talk to each other. That's why we have SWE agents and not knowledge work agents yet. The companies creating this training data, such as Mercor, are doing extremely high-leverage, high-skill work. Critical to moving AI forward. And deeply underappreciated.

一个常见的误解是,training data(训练数据)是低技能的苦力活——扫一些笔记本、在互联网上挖数据、制作带标签的样本。推进 model frontier(模型前沿)所需要的数据恰恰相反。实验室需要的是用于高经济价值任务的训练数据。而这些任务中,除了 SWE 之外的大多数几乎都缺乏文档——它们是多年积累形成的复杂、强领域特定的知识,横跨彼此无法互通的 legacy tools(遗留工具)。这就是为什么我们现在有 SWE agents(软件工程 agent),却还没有 knowledge work agents(知识工作 agent)。像 Mercor 这样创建这类训练数据的公司,做的是杠杆效应极高、技能要求极高的工作。这对推动 AI 前进至关重要,也长期被严重低估。

♥ 60↻ 2💬 36/7 · 19:27x.com ↗
2026 年 6 月 7 日 · 1 条 →

Routing to models is genuinely hard. It means mapping each task to the right model - which requires benchmarking models against your product's specific tasks and dialing in the quality/cost trade-off. And there is an opportunity in that difficulty. Here is the progression I saw with enterprises while on Gemini. Phase 1 (2024): Default to the "it" model. Everybody used GPT regardless of task, because it was the shiny new thing. Phase 2 (early 2025): Over-optimize. Teams over-corrected, looking for the smallest/cheapest model for their task, but did not have evals sophisticated enough to map tasks to models. They ended up burning cycles and shipping slower. Phase 3: Nuanced routing. The industry’s eval muscle and model diversity got to a point where the most sophisticated AI-native startups succeeded in breaking their product into sub-agents and routed each task to the right model - e.g. hardest reasoning to Claude, simplest to Gemini Flash-Lite or open-weight models. And like most product patterns, enterprises followed the AI-native builders 6-9 months later.

将任务路由到不同模型这件事,确实很难。它意味着要把每项任务映射到合适的模型上——这就要求你针对自己产品的具体任务,对模型进行 benchmark(基准测试),并调好质量/成本之间的权衡。而这种难度本身也蕴含着机会。以下是我在 Gemini 期间观察到的企业演进路径。阶段 1(2024):默认使用那个“it” model(当红模型)。无论任务是什么,大家都用 GPT,因为它是那个闪闪发光的新东西。阶段 2(2025 年初):过度优化。团队矫枉过正,试图为自己的任务找到最小/最便宜的模型,但他们并没有足够成熟的 evals(评测)能力,无法把任务准确映射到模型上。结果就是白白消耗精力,产品上线更慢。阶段 3:精细化路由。行业在 eval(评测)能力和模型多样性方面发展到了这样的程度:最成熟的 AI-native 初创公司开始成功地把自己的产品拆分成多个 sub-agents(子 agent),并将每项任务路由到正确的模型——例如,把最难的推理交给 Claude,把最简单的任务交给 Gemini Flash-Lite 或 open-weight models。和大多数产品模式一样,企业会在 6 到 9 个月后跟随这些 AI-native builders。

♥ 27↻ 3💬 56/6 · 19:28x.com ↗
2026 年 6 月 6 日 · 1 条 →

One of the most common mistakes I see enterprise AI teams make is building for today’s model capabilities and price points. Think 6 months out. Models will be way smarter and cheaper. Scaffold around today’s model weaknesses to push the frontier. Bet that the next generation of models will natively solve for the scaffold. Push the frontier again. Build for the slope. Over time, that ability to repeatedly identify and bridge model gaps becomes a moat of its own.

我看到企业 AI 团队最常犯的错误之一,就是按照当下模型的能力和价格点来做构建。要把眼光放到 6 个月之后。模型会聪明得多,也便宜得多。围绕当今模型的弱点搭建 scaffold(脚手架式补强),以推动前沿。要押注下一代模型会原生解决这些 scaffold 所针对的问题。然后再次推动前沿。要为这条斜率而构建。随着时间推移,这种反复识别并弥合模型缺口的能力,本身就会成为一种 moat(护城河)。

♥ 5↻ 0💬 06/5 · 22:27x.com ↗
2026 年 5 月 17 日 · 1 条 →

I have friends who made $10M+ and are miserable. I have friends who made that much and found calm - they don’t need to run anymore. I have friends who made far less and are just happy. Whether something is enough is up to you. Whether you’re happy is independent of your bank account. You can want to be wealthy and still be content now. You don’t need to chase it from a place of lack and desperation. Silicon Valley treats ambition and happiness as mutually exclusive. As if wanting more means you can’t be satisfied now. That’s the trap. You can be both

我有一些朋友赚了 1000 万美元以上,却很痛苦。我也有一些朋友赚到了那么多,并因此找到了平静——他们不再需要继续拼命奔跑。我还有一些朋友赚得少得多,却就是很快乐。一件事算不算“足够”,取决于你自己。你是否快乐,与自己的银行账户无关。你可以想要变得富有,同时也对当下感到满足。你不需要出于匮乏感和绝望去追逐它。Silicon Valley 把 ambition(野心)和 happiness(幸福)看成彼此排斥的东西。仿佛只要还想要更多,你就不可能对现在感到满足。那才是陷阱。你可以两者兼得

♥ 166↻ 8💬 85/16 · 17:54x.com ↗
2026 年 5 月 16 日 · 1 条 →

A generation of PMs is struggling to adapt to AI because they were trained to execute playbooks. AI requires inventing them. For two decades, a few teams invented the product patterns. Everyone else repurposed them to their domain. That made a lot of PM work dry and mechanical. But for a long time, it was enough to build a career. It’s also why so much software feels the same. PMs need to be inventors now, not framework executors. There are no stable playbooks to repurpose. You can’t A/B test your way to a breakthrough AI product. Doable. But PMs need to unlearn.

一代 PMs(产品经理)之所以难以适应 AI,是因为他们过去接受的训练是执行 playbook(成熟打法),而 AI 要求的是去发明这些打法。过去二十年里,只有少数团队发明了产品模式,其他人则把这些模式挪用到自己的领域。这让大量 PM 工作变得枯燥而机械。但在很长一段时间里,这已经足以支撑一份职业生涯。这也是为什么那么多软件给人的感觉都差不多。现在,PMs 需要成为发明者,而不是 framework(框架)的执行者。已经没有可以稳定复用的 playbook 了。你不可能靠 A/B test(A/B 测试)一路试出一个突破性的 AI 产品。这并非做不到,但 PMs 需要先学会“去习得”旧方法。

♥ 26↻ 1💬 65/15 · 22:27x.com ↗
2026 年 5 月 8 日 · 1 条 →

I'm moving on from @Google. I had the privilege of helping build two businesses from zero: first across Search & Ads, then Gemini. Three years ago, OpenAI and Anthropic were in the lead. We built what it took to compete: the playbook for building AI models, the customer feedback flywheel, and the enterprise business. Gemini 3 was the moment those systems came together. To the Gemini team: we went from underdogs to competing at the frontier. Keep pushing. For now, I'm enjoying the emergent capabilities of some real intelligence at home - my toddler. She's been quietly shipping.

我将从 @Google 离开了。我有幸帮助从零打造了两项业务:先是 Search & Ads,之后是 Gemini。三年前,OpenAI 和 Anthropic 处于领先地位。我们建立起了参与竞争所需的一切:构建 AI 模型的 playbook(方法论),客户反馈 flywheel(飞轮),以及 enterprise business(企业业务)。Gemini 3 是这些系统汇合到一起的时刻。致 Gemini 团队:我们从弱势一方走到了在前沿展开竞争。继续推进。眼下,我正在享受家里某种真正智能所展现出的 emergent capabilities(涌现能力)——我家蹒跚学步的孩子。她一直在安静地持续 shipping(交付成果)。

♥ 1.1K↻ 11💬 535/7 · 20:48x.com ↗
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers