BuildSpeak每日 builder 文摘
今日归档生词本关于
SW

Swyx

@swyx ↗

achieve ambition with intentionality, intensity, integrity & insanity.

3 最新161 累计61 期
每条推文 hover 显示单独 ▶
2026 年 7 月 23 日 · 3 条 →

wew

wew

♥ 11↻ 0💬 37/23 · 05:22x.com ↗

lmao someone just reminded me ive been on this shit for a while huh

笑死,刚有人提醒我,我在这玩意儿上已经混了挺久了,哈

♥ 4↻ 0💬 17/23 · 01:06x.com ↗

also another fun one

还有一个也挺好玩的

♥ 7↻ 1💬 17/22 · 16:21x.com ↗
2026 年 7 月 22 日 · 3 条 →

at some point in your engineering career, a wizened graybeard is going to lecture you about the importance of independently separable control plane vs data plane. it is very important that you listen. (and then try to learn about the management plane as early as you can..)

在你的工程师职业生涯中的某个时刻,一位饱经世故的老前辈会给你上一课,讲 independently separable control plane(可独立分离的控制平面)与 data plane(数据平面)有多重要。你一定要认真听。(然后也要尽早去了解 management plane(管理平面)……)

♥ 157↻ 3💬 97/22 · 03:47x.com ↗

recorded Codex + ChatGPT Work + 10M user milestone pod with @akshaynathan_ who leads Productivity engineering. I’ve been on record that Work + GPT 5.6 is the most company defining launch since og chatgpt itself. this thing (with @arix’s computer use) is gonna reach >1B users worldwide. thursday on @latentspacepod

刚和负责 Productivity engineering 的 @akshaynathan_ 录完一期关于 Codex + ChatGPT Work + 1000 万用户里程碑的播客。我之前就公开说过,Work + GPT 5.6 是自最初的 ChatGPT 以来,对公司定义意义最大的一次发布。这个东西(再加上 @arix 的 computer use)将会在全球触达超过 10 亿用户。周四见 @latentspacepod

♥ 42↻ 2💬 67/21 · 23:59x.com ↗

@michelleefang @andyrapista 2 years later she is still going

@michelleefang @andyrapista 两年过去了,她还在继续

♥ 2↻ 0💬 17/21 · 16:07x.com ↗
2026 年 7 月 21 日 · 3 条 →

unironically this is happening right tf now

说真的,这事现在他妈的就正在发生。

♥ 35↻ 3💬 107/21 · 05:28x.com ↗

very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of "frontier" model training is that even without training on test, you can basically cheat by training on test lookalikes, enabling you to goalseek almost any benchmark number you want. however when they are released open weights, 99% of the time the norm is that you do not get the datasets/rlenvs that would easily show you if someone was training on Temu Tbench, so there is plausible deniability. Alex and Omar discuss applying standard NLP distance metrics on hidden trajectories. There's no ultimate solution here, but they have some prelim explorations. It happens to support the finding that RLMs can generalize to unseen tasks that share latent structure observed in training.

这里有一篇很值得注意的轨迹(trajectory)比较分析,藏在 @a1zhang 和 @lateinteraction 的 RLM paper 里。关于“frontier” model 训练,有一个公开的秘密:即使不直接在测试集上训练,你基本上也可以通过在测试集的相似版本上训练来作弊,从而把几乎任何 benchmark 分数都朝你想要的目标去调。不过,当这些模型以 open weights 形式发布时,99% 的情况下,常态是你拿不到那些数据集 / RL envs,而正是它们本可以很容易地告诉你,某人是不是在 Temu Tbench 这类东西上训练过,所以发布方总有可辩解空间。Alex 和 Omar 讨论了如何把标准 NLP 距离度量应用到隐藏轨迹上。这里没有终极解决方案,但他们做了一些初步探索。碰巧的是,这也支持了这样一个发现:RLM 可以泛化到那些未见过、但与训练中观察到的潜在结构共享 latent structure 的任务。

♥ 62↻ 7💬 137/21 · 03:43x.com ↗

[1]

♥ 3↻ 0💬 17/21 · 02:57x.com ↗
2026 年 7 月 20 日 · 2 条 →

even as someone who is most definitely not a custom keyboards guy, i gotta admit this is pretty dang cool.

即便作为一个绝对算不上 custom keyboards 圈内人的人,我也得承认这玩意儿确实相当酷。

♥ 86↻ 0💬 167/20 · 04:31x.com ↗

here goes

这就开始吧

♥ 4↻ 2💬 17/19 · 18:38x.com ↗
2026 年 7 月 19 日 · 2 条 →

people hating on europe miss that they actually have some of the top ai engineers in the world if you actually to elicit the right ones. we're basically running the most competitive global arena for ai talent (as measured by talks and workshops), will be v interesting to see how this goes as WF enters the eval window

那些唱衰 europe 的人忽略了一点:如果你真的能把合适的人才发掘出来,那里其实拥有一些全球顶尖的 AI engineers。我们基本上正在运营一个面向 AI talent 的全球最高竞争度舞台(以 talks 和 workshops 来衡量),随着 WF 进入 eval window,接下来会怎么发展将非常有意思

♥ 53↻ 1💬 217/18 · 23:50x.com ↗

btw at current rate i think AEO will be fully responsible for $1m in my revenue next yr

顺便说一句,按目前这个速度,我觉得明年 AEO 将会对我 $1m 的收入负全部责任

♥ 6↻ 0💬 27/18 · 20:45x.com ↗
2026 年 7 月 18 日 · 3 条 →

still true

仍然是真的

♥ 4↻ 0💬 07/18 · 06:20x.com ↗

alpha only gone when we all stop asking these level of questions and start discussion about on-policy autoaeo (does claude optimizing your aeo disproportionately work on claude?) vs generalizable aeo

只有当我们都不再问这种层次的问题,并开始讨论 on-policy autoaeo(Claude 对你的 aeo 进行优化时,是否会对 Claude 本身产生不成比例的效果?)与 generalizable aeo 之间的区别时,alpha 才会消失

♥ 10↻ 1💬 37/18 · 01:41x.com ↗

btw if you havent set your {codex | claude | gemini | devin} automations to autoresearch how to improve your seo/aeo every week you are really truly missing out on free, should-be-commoditizing-but-weirdly-untapped alpha

顺便说一句,如果你还没有把你的 {codex | claude | gemini | devin} 自动化流程设为每周自动研究如何提升你的 seo/aeo,那你真的、确实错过了免费的 alpha——它本来应该早就商品化了,但却奇怪地一直没有被充分挖掘

♥ 1.0K↻ 24💬 517/17 · 22:25x.com ↗
2026 年 7 月 17 日 · 3 条 →

HOLY FUCKING SHIT the react avengers have assembled

他妈的太炸了,react avengers 集结完毕了

♥ 12↻ 0💬 67/17 · 06:32x.com ↗

we make sure some of the top YC AI companies are featured at AIE every single year. this year, we were graced by @garrytan and @eve_bouff to cap off our startups and design engineer focused audiences respectively. Really enjoyed these raw high value perspectives!

我们确保每一年都会让一些顶尖的 YC AI 公司亮相 AIE。今年,我们很荣幸请到 @garrytan 和 @eve_bouff,分别为我们以 startup 和 design engineer 为重点的观众群体压轴收尾。真的很享受这些真实、信息价值很高的观点!

♥ 36↻ 1💬 127/17 · 02:10x.com ↗

@mattpocockuk @trq212 @xdg's session/tree based grill me looks cool but havent tried

@mattpocockuk @trq212 @xdg 的 session/tree based grill me 看起来很酷,不过我还没试过

♥ 4↻ 0💬 37/16 · 17:32x.com ↗
2026 年 7 月 16 日 · 3 条 →

you can just tweet things into existence

你完全可以靠发 tweet 把事情直接变成现实

♥ 140↻ 2💬 127/16 · 01:19x.com ↗

@jluan follow this track playlist

@jluan 关注这个 track 播放列表

♥ 3↻ 0💬 27/16 · 01:15x.com ↗

someone just told me about this take* on CUA this is one of those gell mann moments for me lol. i've been watching computer use since World of Bits (Shi, fan, karpathy, hernandez & liang 2017). we were the first technical pod to interview @jluan about Adept's work three years ago, we were there in the @AnthropicAI building when they first launched Computer Use 2 years ago, I fanboyed over Claude Cowork in our @felixrieseberg pod 3 months ago, and we ran our first full computer use track at @aidotengineer ft. @DhruvBatra_ @proceduralia @francedot 3 weeks ago. GPT 5.6 + Superapp is even better at CUA than everything i just mentioned. excited for our @AriX podcast to discuss the @skybysoftware story and Codex progress. if you actually use these things as intensely as we do, CUA is progressing so, so incredibly fast. i have asked my nontechnical team to CUA as much as possible, all their knowledge work with signing up for random payment and invoicing portals and speaker and sponsor and attendee and vendor and union data requests. if you found yourself nodding along to this take below, you are so not up to date that you don't know what you don't know, and underestimating capabilities is quite a dangerous category error if you are doing any ai decisionmaking. *i admire dwarkesh alot; only criticizing one single take, not the message nor the overall enterprise, screenshot only to share

刚刚有人跟我说了这个关于 CUA 的观点*,这对我来说就是那种 gell mann moment,笑死。我从 World of Bits(Shi, fan, karpathy, hernandez & liang 2017)开始就一直在关注 computer use。三年前,我们是第一个采访 @jluan、讨论 Adept 工作的技术 podcast;两年前,@AnthropicAI 第一次发布 Computer Use 时,我们就在他们楼里;三个月前,我还在我们和 @felixrieseberg 的 podcast 里对 Claude Cowork 大吹特吹;三周前,我们又在 @aidotengineer 办了第一条完整的 computer use track,嘉宾有 @DhruvBatra_、@proceduralia、@francedot。GPT 5.6 + Superapp 在 CUA 上甚至比我刚提到的所有东西都更强。我很期待我们和 @AriX 的 podcast,来聊聊 @skybysoftware 的故事以及 Codex 的进展。如果你真的像我们这样高强度使用这些东西,你就会知道 CUA 的进步速度快得惊人,真的、真的非常快。我已经让我的非技术团队尽可能多地使用 CUA,把他们所有的知识型工作都交给它,比如注册各种支付和开票 portal、处理 speaker、sponsor、attendee、vendor 以及 union 的数据请求。如果你看着下面这个观点还频频点头,那你真的已经严重落后了,落后到你甚至不知道自己不知道什么;而如果你正在做任何 AI 决策,低估能力会是一种相当危险的 category error。*我很欣赏 dwarkesh;我批评的只是这一个具体观点,不是他的整体信息传达,也不是他做的整体事业,截图只是为了分享

♥ 77↻ 7💬 247/15 · 19:28x.com ↗
2026 年 7 月 15 日 · 3 条 →

@shloked @amazon

@shloked @amazon

♥ 2↻ 0💬 07/15 · 04:06x.com ↗

SF Personal AI engineers - if you are building personal agents, come demo at new media lab this thursday night (and meet @shloked, one of the best builder-writers I've had the fortune to meet this year). last time we held this meetup our featured speakers got acquired by @Amazon hardware division... and 2 years later I'm still a daily active user. just incredible to see the staying power of PAI.

SF 的 Personal AI 工程师们——如果你们正在构建 personal agents(个人 agent),这周四晚上来 new media lab 做 demo 吧(也顺便见见 @shloked,他是我今年有幸认识的最优秀的 builder-writers 之一)。上次我们办这场 meetup 时,我们的 featured speakers 后来被 @Amazon 的 hardware division 收购了……而两年后的今天,我仍然是他们产品的 daily active user(日活用户)。看到 PAI 的持久生命力,真的太不可思议了。

♥ 40↻ 5💬 97/15 · 04:06x.com ↗

@VivianBala still unreal that this is a real picture from this year lol first talk

@VivianBala 还是觉得难以置信,这居然是今年拍的一张真实照片 lol 第一次演讲

♥ 7↻ 0💬 17/15 · 03:51x.com ↗
2026 年 7 月 14 日 · 3 条 →

@mattpocockuk @trq212 what other people are recommending to use

@mattpocockuk @trq212 其他人在推荐用什么

♥ 8↻ 1💬 07/14 · 00:51x.com ↗

where i'm currently at for Big Boy projects: - sol ultra to plan - fable 5 to critique - sonnet 5/terra ultra/swe 1.7 to ultracode/slop cannon - devin review to review (using kakuna) ~always use a variant of @mattpocockuk's grill-me or @trq212's interview-me to elicit decisions upfront

我目前做 Big Boy projects 时的配置是:- 用 sol ultra 来做规划 - 用 fable 5 来做评审 - 用 sonnet 5/terra ultra/swe 1.7 来做 ultracode/slop cannon - 用 devin review 来做 review(使用 kakuna)~ 总之始终会用一个变体版的 @mattpocockuk 的 grill-me 或 @trq212 的 interview-me,来提前引导出需要做的决策

♥ 364↻ 7💬 357/13 · 23:32x.com ↗

@resend AGAIN. TODAY. at this point if i get a text from my mom i expect codex to ask for a resend api key in order to read it

又是 @resend。今天也是。到了这个地步,如果我收到我妈发来的短信,我都要觉得 codex 会为了读取它而来问我要一个 resend API key 了

♥ 1↻ 0💬 07/13 · 23:23x.com ↗
2026 年 7 月 13 日 · 2 条 →

btw the difference is introspection/backpropagation as einstein famously said, the definition of insanity is doing multiple rollouts with no expectation of advantage

顺便说一句,区别在于 introspection(内省)/ backpropagation(反向传播);正如 Einstein 那句著名的话所说,所谓 insanity(疯狂),就是在并不期望有任何 advantage(优势)的情况下反复进行多次 rollout(展开/采样)。

♥ 186↻ 8💬 147/12 · 16:37x.com ↗

more on @latentspacepod writeup

更多内容见 @latentspacepod 的 writeup(文章/解读)。

♥ 11↻ 0💬 87/12 · 08:04x.com ↗
2026 年 7 月 12 日 · 1 条 →

if you only learned about jevons paradox primarily wrt software demand in the age of agentic engineering, you may not have fully internalized jevons parodox’s impact under the conditions of: - humans who can wield coding agents well* - coding agents breaking containment to all other knowledge work as the efficiency of labor goes up/unit cost of knowledge work goes broadly down, the demand for total work and better knowledge goes up, not down. what happened to coding isnt the exception; it’s the herald. *aka AI Engineers

如果你主要只是在 agentic engineering 时代、相对于软件需求的语境中了解 Jevons paradox,那么你可能还没有充分内化 Jevons paradox 在以下条件下的影响:- 人类能够熟练使用 coding agents* - coding agents 打破边界,扩散到其他所有知识工作中 随着单位劳动效率上升、知识工作的单位成本整体下降,对总工作量以及更高质量知识的需求会上升,而不是下降。发生在 coding 上的事并不是例外;它是先声。*也就是 AI Engineers

♥ 143↻ 9💬 287/12 · 04:04x.com ↗
2026 年 7 月 10 日 · 3 条 →

whoever does AEO for @resend needs to get a raise, all the leading frontier models keep trying to use resend for emails even when i have existing transactional infrastructure already set up

给 @resend 做 AEO 的那位,不加薪都说不过去;所有领先的 frontier models(前沿模型)老是想用 resend 来发邮件,哪怕我其实早就已经搭好了现成的 transactional infrastructure(事务型基础设施)。

♥ 97↻ 2💬 197/10 · 00:29x.com ↗

u guys were clowning on @greptile but turns out they were just the inspo for @openai to go so, so much harder

你们之前还在拿 @greptile 开涮,结果现在看,他们不过是给 @openai 提供了灵感,让后者直接把力度拉高了不知道多少倍。

♥ 6↻ 0💬 27/9 · 21:50x.com ↗

more grist for the mootools mafia lore

又给 mootools mafia 的传说添了点新素材。

♥ 1↻ 0💬 07/9 · 17:24x.com ↗
2026 年 6 月 29 日 · 3 条 →

we registered 1000 people today. this is what it looked like at hour 3. tmr and tuesday going to be absolutely batshit

我们今天登记了 1000 人。这是第 3 小时时的情况。tmr 和周二肯定会彻底忙疯。

♥ 17↻ 2💬 66/29 · 06:28x.com ↗

if youre a speaker make your own

如果你是 speaker,就自己做一个。

♥ 6↻ 0💬 06/29 · 06:19x.com ↗

because i'm not a design engineer myself, this track is one of the harder ones I struggle to curate. very fortunate to befriend Geoff who has lent a hand to the past 2 years of AI UX meetups, and now is the opener for the Design Engineers at AIE! see you wednesday

因为我自己不是 design engineer,这条 track 是我觉得较难策划、也一直比较吃力的一条。很幸运认识了 Geoff,他在过去 2 年的 AI UX meetups 中都帮了忙,现在还会作为 AIE 的 Design Engineers 场次 opener!周三见。

♥ 17↻ 0💬 96/29 · 06:18x.com ↗
2026 年 6 月 28 日 · 3 条 →

Memberships for companies and creative individuals. Apply here please!

面向公司和创意个人的会员资格。请在此申请!

♥ 9↻ 0💬 36/27 · 22:58x.com ↗

impromptu ai engineer preshow floor tour and AMA

impromptu ai engineer 预热现场参观与 AMA

♥ 33↻ 1💬 126/27 · 20:45x.com ↗

An interesting way to take Noam at his word in regards to always keeping a constant inference budget for any eval reporting - is that open models have a lot more dollar per token mileage than closed model APIs. So anyone launching an open model today or situationally incentivized toward open models should obviously report thinking levels measured by dollar inference on popular inference providers, instead of by number of tokens on the x axis

就 Noam 关于在任何 eval(评测)报告中始终保持恒定 inference(推理)预算这一点“照字面执行”,有一种有趣的方式:open models 的每 token 美元性价比,要比 closed model APIs 高得多。因此,任何今天发布 open model 的人,或者在特定情境下有动力偏向 open models 的人,显然都应该按热门推理服务商上的美元推理成本来报告 thinking(思考)水平,而不是在 x 轴上用 token 数量来衡量。

♥ 100↻ 4💬 116/27 · 19:16x.com ↗
2026 年 6 月 27 日 · 2 条 →

minor milestone in the growth of swyx inc: we took over my new media lab today it will be the new home for engineer-creatives in san francisco; a third place to make; a finishing school for technical storytellers; a place to inspire the inspirational. to our complete surprise; it came with a datacenter rack randomly set up and wired up! need advice on what to put in this thing. @alexocheema halp

swyx inc 成长过程中的一个小里程碑:今天我们接手了我的 new media lab,它将成为 san francisco 里 engineer-creatives 的新家;一个动手创造的 third place;一个为技术型讲述者准备的 finishing school;一个激发 inspirers 的地方。完全出乎我们意料的是;这里还随机自带了一个已经搭好并接好线的 datacenter rack!想听听大家建议,这东西里该放点什么。@alexocheema halp

♥ 72↻ 2💬 206/27 · 05:59x.com ↗

we have been scaling without slop by working with aligned domain experts to add coverage with both oai and ant launching multi-billion dollar services arms, it’s clear that FDE is one of the most in demand disciplines on earth, but I have never done the job it’s been an absolute pleasure working with Basil on our first ever AI FDE miniconference! see at next week

我们一直在通过与目标一致的 domain experts 合作来扩展覆盖范围,而且没有引入 slop;现在 oai 和 ant 都在推出价值数十亿美元的 services arms,这很明显表明 FDE 是地球上需求最高的学科/岗位之一,但我自己从没做过这份工作。能和 Basil 一起筹办我们史上第一次 AI FDE miniconference,真的是一种绝对的享受!下周见

♥ 46↻ 2💬 156/26 · 20:35x.com ↗
2026 年 6 月 26 日 · 3 条 →

this is too late for AIE worlds fair but heres the how to book talks in the first place

这对 AIE worlds fair 来说已经太晚了,不过这是最开始该如何预订演讲的办法

♥ 0↻ 0💬 16/26 · 04:46x.com ↗

@neondatabase @nikitabase @ankrgyl pls find my soundcloud

@neondatabase @nikitabase @ankrgyl 请找到我的 soundcloud

♥ 7↻ 0💬 16/25 · 23:22x.com ↗

btw this was 3 years ago lmao

顺便说一句,这是 3 年前的事了,笑死

♥ 3↻ 0💬 06/25 · 22:17x.com ↗
2026 年 6 月 25 日 · 3 条 →

lots of folks prepping talks next week (congrats!). Some thoughts from RLing on thousands of hours of engineer- and researcher- focused talks: - AI generated svgs > AI generated imgs. MAXIMUM 4 ai slop images in your slides, I don't care how pretty your mom thinks they are (exception ofc if your talk is ABOUT imagegen) - Be pointy. Better to have 1 message with 5 surprising applications, than 5 messages with no concrete examples. - Put code on screen. Engineers like to see code. Especially if they can nitpick it to death over things that aren't the point. - Don't forget to Entertain. Being actually funny, or good at telling relevant anecdotes, is more important than adding yet another bullet point. - Have a Thesis. every talk gets ONE "if you remember one thing from this talk, it is this" card. 1) SO MANY PEOPLE DON'T USE IT. 2) SO MANY PEOPLE DON'T PLAN FOR IT. you get one. use it well? - Have a Thesis Slide. you've been in that talk where everyone gets their phone out to take a photo of the slide. Because people see 1000x more images than videos, you are much more likely to have a viral talk/slide if you have a single viral slide. You are much more likely to have a viral slide if you TRY and most people do not TRY. Struggling? Collect examples from talks you like and adapt their format to your thing. Your slides have a power law - don't spend 5% of your time each on 20 slides, spend 80% of your time on 1 slide and have the rest build up to that 1 slide. - You might not need slides. live demo in IDE, off the cuff rambles, pull audience member up to roleplay, do call and response with the audience, sing/perform, voice over vibe videos, I have seen it all. higher risk/effort, higher reward when done well. - Be pleasant to listen to. Talks with bad/no visuals still can be listened to. Talks with bad audio are DOA. Be confident, project, remember vocal variation, evoke emotion. - Design the emotional journey. Start strong, end strong, have a peak aha/laugh/thesis moment in the middle. Everything else is buildup. - Data driven talks are underrated. Present pretty charts and surprising, authoritative numbers. Use the stage authority to infer from the data to support the broader thesis. Done well, it will feel like my obvious conclusion from objectively looking at the data you presented, not you feeding the conclusion to me. - How to shill your product/company without feeling salesy: teach me everything I didn't know I should to know about the problems you solve, and then you shall have earned the right to convince me you are the guys to trust to solve it once and for all. - Actively watch a lot of talks. like with anything you both need a lot of reps AND you need to explore to find the styles/role models that you shine in. Passive = no introspection after watching, just mindlessly autoplay next video. Active = trying to articulate why a talk was good/bad after watching. If you want to REALLY hone it, think about how you feel about a famous talk and then TRANSCRIBE the talks and read the words on the printed page, and then compare with a "normal" talk and define rules you will follow for yourself to improve. thanks to @dexhorthy for organizing the AIEWF Speaker prep meetup tonight. we should actually do more of these....

很多人都在准备下周的演讲(恭喜!)。这是 RLing(通过大量实践打磨)了成千上万小时、面向工程师和研究者的演讲后的一些想法:- AI generated svgs 胜过 AI generated imgs。你的 slides 里最多放 4 张 AI slop 图片,我才不管你妈觉得它们有多好看(当然,如果你的演讲本来就是讲 imagegen,那是例外)- 要尖锐、聚焦。与其讲 5 个没有具体例子的观点,不如只讲 1 个核心信息,再配上 5 个让人意外的应用。- 把 code 放到屏幕上。工程师喜欢看 code。尤其是当他们可以对着那些不是重点的细节吹毛求疵、挑刺到死的时候。- 别忘了娱乐性。真的有趣,或者很会讲相关的轶事,比再多加一个 bullet point 更重要。- 要有 Thesis。每个 talk 都该有且只有一张“如果你只记住这场演讲的一件事,那就是这个”的卡片。1)太多人根本不用它。2)太多人根本没为它做设计。你就这一次机会。好好用,行吗?- 要有一张 Thesis Slide。你肯定经历过那种演讲:大家纷纷掏出手机拍某一页 slide。因为人们看到的图片比视频多 1000 倍,所以如果你有一张能传播开的单页 slide,你的 talk/slide 更可能火。只要你去尝试,你做出 viral slide 的概率就会高很多,而大多数人根本没在尝试。没思路?去收集你喜欢的演讲里的例子,借鉴它们的格式,改成适合你内容的版本。你的 slides 服从 power law(幂律分布)——不要把时间平均分成 20 张 slide、每张只花 5%;把 80% 的时间砸在 1 张 slide 上,其余内容都为那 1 张做铺垫。- 你可能根本不需要 slides。直接在 IDE 里 live demo、即兴漫谈、拉一位观众上来 roleplay、和观众 call and response、唱歌/表演、给氛围感视频配 voice over,我什么都见过。风险和投入更高,但做好了回报也更高。- 要让人听着舒服。视觉做得差甚至没有 visuals 的 talk,至少还能听;音频差的 talk 则是 DOA(到场即死)。要自信,声音打出去,记得做 vocal variation(语调变化),调动情绪。- 设计情绪曲线。开头强,结尾强,中间要有一个 aha/笑点/Thesis 的峰值时刻。其他一切都是铺垫。- data driven talks 被低估了。展示漂亮的图表和令人意外、又有权威感的数据。利用舞台赋予你的权威,从数据中推出结论,以支撑更大的 Thesis。做好了,听众会觉得“这是我从你展示的客观数据中自然得出的显然结论”,而不是你把结论硬塞给我。- 怎么推销你的 product/company 又不显得 salesy:先把那些我原本不知道、但其实应该知道的、关于你所解决问题的一切都教给我;这样你才算赢得了说服我的资格,让我相信你们就是那个值得信任、能一劳永逸解决问题的团队。- 主动地去看很多 talks。和任何事情一样,你既需要大量 reps(重复练习),也需要广泛探索,找到最适合你发光的风格和 role model。Passive = 看完没有复盘,只是无脑自动播放下一个视频。Active = 看完后尝试说清楚一场 talk 为什么好/不好。如果你真的想把这件事磨到极致,就先想想你对一个著名演讲的感受,然后把它 TRANSCRIBE(转写)出来,读印在纸上的文字;再和一场“普通”演讲对比,最后给自己制定一套改进时会遵守的规则。感谢 @dexhorthy 今晚组织 AIEWF Speaker prep meetup。我们确实应该多办这种活动……

♥ 188↻ 14💬 206/25 · 02:03x.com ↗

we are going to have to Rebuild So. Much. Infra. for the age of Software Factories

在 Software Factories 时代,我们将不得不重建太、太、多的基础设施(Infra)。

♥ 281↻ 14💬 236/25 · 00:14x.com ↗

LOTS of alpha in this pod: - Why Databricks beat Snowflake (! a straight answer!) - Why everyone is building a metaharness now - Why the @neondatabase made so much sense (so much @nikitabase glazing its not even funny) - How LTAP solves the HTAP dream I discussed with @ankrgyl in our @braintrust pod - What happened to @MosaicML + DBRX - How to maintain research/startup culture in a $175B megacorp - What's more important knowledge/experience in the race to the agent cloud: databases, operating systems, or.... networking! very honored to be invited to @Data_AI_Summit to interview two of the top people in our industry and somehow be able to jam on everything from the @bennstancil modern data stack theme to @alighodsi's amazing keynote aura

这期 pod 里有很多 alpha(高价值信息):- 为什么 Databricks 打赢了 Snowflake(!一个直截了当的答案!)- 为什么现在人人都在做 metaharness - 为什么 @neondatabase 如此说得通(对 @nikitabase 的夸赞多到都不好笑了)- LTAP 如何解决我在和 @ankrgyl 做 @braintrust pod 时讨论过的 HTAP 梦想 - @MosaicML + DBRX 后来到底发生了什么 - 在一家 1750 亿美元的超级巨头公司里,如何维持 research/startup culture - 在通往 agent cloud 的竞赛中,knowledge/experience 之外更重要的是什么:databases、operating systems,还是…… networking!非常荣幸受邀来到 @Data_AI_Summit,采访我们行业里两位最顶尖的人物,并且居然能一路畅聊,从 @bennstancil 的 modern data stack 主题,一直聊到 @alighodsi 那场气场惊人的 keynote。

♥ 52↻ 4💬 86/24 · 19:23x.com ↗
2026 年 6 月 23 日 · 2 条 →

[1]

♥ 0↻ 0💬 16/23 · 06:42x.com ↗

i dont think anyone is correctly doing the math around how SpaceX, the NeoCloud+NeoLab, is currently going to market? SpaceX has already recouped about HALF its investment in Cursor, in compute deals. The other half is paid for if Composer 3 does well. No other company is simultaneously a leading model lab + neocloud (at least where GPUs is concerned). its a crazy effective combo iff you've adequately planned out gpu supply if inhouse training 1) goes very well 2) doesn't go very well

我觉得现在还没人真正把 SpaceX,也就是这个 NeoCloud+NeoLab 的市场打法,相关的账算明白吧?SpaceX 光靠 compute(算力)交易,已经收回了它在 Cursor 上大约一半的投资。另一半如果 Composer 3 表现不错,也就能回本了。没有其他公司能同时既是领先的 model lab(模型实验室),又是 neocloud(至少在 GPUs 这件事上如此)。如果你已经充分规划好了 GPU 供应,那么这会是个疯狂高效的组合——无论 inhouse training(内部训练)1)进展非常顺利,还是 2)进展没那么顺利。

♥ 61↻ 1💬 156/23 · 06:06x.com ↗
2026 年 6 月 22 日 · 3 条 →

btw i've been shopping around for insurers for the New Media Lab we are setting up (basically the creative playground housing swyx inc) and yeah the NPS of Corgi is insanely high my real estate broker: "just go with corgi they are covering every single one of my clients rn" breaking through with ~100% greenfield market share like this is unheard of in the insurance industry

顺便说一句,我最近一直在给我们正在筹建的 New Media Lab 物色保险公司(基本上就是容纳 swyx inc 的创意游乐场),然后,嗯,Corgi 的 NPS(净推荐值)高得离谱。我的房地产经纪人说:“直接选 corgi 吧,他们现在给我的每一位客户都在承保。” 像这样以接近 100% 的 greenfield market share(全新市场份额)实现突破,在保险行业简直闻所未闻。

♥ 56↻ 1💬 96/22 · 05:10x.com ↗

@QuinnyPig i think this is where i challenge @willccbb to his first poaster session. or maybe @willdepue. or @WilliamBryk. idk all the wills?

@QuinnyPig 我觉得这就是我该向 @willccbb 发起他第一次 poaster session 挑战的时候了。或者也许是 @willdepue。或者 @WilliamBryk。也不知道,所有叫 Will 的人?

♥ 7↻ 0💬 46/22 · 01:28x.com ↗

@QuinnyPig

@QuinnyPig

♥ 6↻ 0💬 26/22 · 01:27x.com ↗
2026 年 6 月 21 日 · 3 条 →

@aiDotEngineer @brendanhunting @TedLasso @USMNT @philipkiely this photo looks fucking chatgpt generated

@aiDotEngineer @brendanhunting @TedLasso @USMNT @philipkiely 这张照片看起来他妈的像是 ChatGPT 生成的

♥ 4↻ 0💬 36/21 · 02:14x.com ↗

btw this is what happens on July 4 if team usa wins this game Wednesday after next

顺便说一句,如果 Team USA 赢下下下个星期三的这场比赛,那么 7 月 4 日就会发生这个

♥ 42↻ 0💬 126/21 · 01:45x.com ↗

@aiDotEngineer @brendanhunting @TedLasso @USMNT @philipkiely this was key thing to figure out btw @GeminiApp is a VERY good sports handicapper (thanks @OfficialLoganK ). need to draw from a lot of sources to do this

@aiDotEngineer @brendanhunting @TedLasso @USMNT @philipkiely 顺便说一句,这才是需要搞清楚的关键一点:@GeminiApp 是个非常厉害的体育 handicapper(谢谢 @OfficialLoganK)。要做到这一点,需要从很多来源取材

♥ 6↻ 0💬 16/20 · 23:35x.com ↗
2026 年 6 月 20 日 · 3 条 →

10 years ago, you will be asked by @bendhalpern and @jessleenyc to write your first blog on @thepracticaldev. it is very important that you answer. *now @MLHacks, who are producing the first ever physical daily newspaper at @aidotengineer WF

10 年前,@bendhalpern 和 @jessleenyc 会请你在 @thepracticaldev 上写你的第一篇 blog,这一点非常重要,你一定要回应。*现在 @MLHacks 正在 @aidotengineer WF 制作有史以来第一份实体的日报纸

♥ 2↻ 0💬 36/20 · 07:24x.com ↗

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr @zoink @Polymarket @aiDotEngineer /goaaaaaaaaaaal

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr @zoink @Polymarket @aiDotEngineer /goaaaaaaaaaaal

♥ 2↻ 0💬 06/20 · 07:06x.com ↗

Anthropic is going to IPO at $2T

Anthropic 将以 $2T 的估值进行 IPO

♥ 269↻ 2💬 186/19 · 21:31x.com ↗
2026 年 6 月 19 日 · 3 条 →

+55% in one day. i should start a fund (dm if you would actually help me run one, i have no idea how to run one)

一天就涨了 +55%。我该开始做个 fund 了(如果你真的愿意帮我一起运营,dm 我;我完全不知道 fund 该怎么运作)

♥ 47↻ 1💬 126/19 · 00:22x.com ↗

@DevinAI @tbpn

@DevinAI @tbpn

♥ 4↻ 0💬 06/18 · 23:00x.com ↗

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr whoa i had no idea who i was talking to lmao

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr 哇,我之前完全不知道我是在跟谁说话,笑死了

♥ 5↻ 0💬 06/18 · 22:49x.com ↗
2026 年 6 月 18 日 · 3 条 →

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr best review from a cancer survivor x tech realist

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr 来自一位癌症幸存者兼科技现实派的最佳评价

♥ 2↻ 0💬 16/18 · 07:38x.com ↗

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr another SUPER fun highlight of my evening was telling @zoink how we are using @polymarket prediction markets to gauge the implied value of our july 1 @aiDotEngineer world cup suite being a team USA game

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr 我今晚另一个超级有趣的高光时刻,是告诉 @zoink 我们如何使用 @polymarket prediction markets(预测市场)来衡量我们 7 月 1 日 @aiDotEngineer world cup suite 作为 team USA 比赛的隐含价值

♥ 1↻ 0💬 06/18 · 07:34x.com ↗

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr paper

@midjourney @Scobleizer @bryan_johnson @DavidSHolz @iScienceLuvr 论文

♥ 3↻ 0💬 16/18 · 06:32x.com ↗
2026 年 6 月 17 日 · 3 条 →

@WorkOS @mattpocockuk @zackproser also doing great. AIE audience LOVES the workos content, they are doing something right

@WorkOS @mattpocockuk @zackproser 也做得很棒。AIE 的受众非常喜欢 WorkOS 的内容,他们肯定是做对了什么

♥ 3↻ 1💬 36/17 · 01:32x.com ↗

@TomasReimers @cursor_ai full info

@TomasReimers @cursor_ai 完整信息

♥ 9↻ 0💬 16/16 · 23:17x.com ↗

@TomasReimers @cursor_ai *github competitor fml

@TomasReimers @cursor_ai *GitHub 竞争对手,fml

♥ 48↻ 0💬 56/16 · 17:36x.com ↗
2026 年 6 月 16 日 · 1 条 →

guys goblingate was 1.5 months ago

大家,goblingate 是 1.5 个月前的事了。

♥ 120↻ 1💬 206/16 · 02:13x.com ↗
2026 年 6 月 15 日 · 3 条 →

havent seen many people outside anthropic ultracode yet. this thing is scarily good at burning tokens but you need to set up your repo to parallelize properly to make use of the fanout that i think subagents are best at. basically the idea is "subroutines but intelligent". when you undersatnd just how much knowledge work is just yakshaves after yakshaves that require some judgment and intelligence, you start to appreciate that dynamic workflows are not just for coding tasks...

还没看到很多 Anthropic 之外的人在用 ultracode。这个东西在烧 token(令牌)方面强得吓人,但你得把自己的 repo 配置好、正确并行化,才能利用我认为 subagents 最擅长的那种 fanout(扇出)能力。基本思路就是“subroutines(子程序),但有智能”。当你理解了有多少知识型工作其实只是一个接一个的 yak shave(为次要前置问题反复折腾),而这些事又确实需要一定判断力和智能时,你就会开始意识到,dynamic workflows(动态工作流)并不只是适用于编码任务……

♥ 45↻ 2💬 156/15 · 07:00x.com ↗

Satya on loops as IP: > This is the first time we can create a real cognitive loop between people and digital systems. That is a mind-bender, because it changes how we even conceptualize work inside an enterprise. > This means the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound. You can offload a task, or even a job, but you can never offload your learning > In my view, our priority has to be building a frontier ecosystem, not just a frontier model, so value flows broadly across every company, every industry, and every country. One where every organization can own the learning loop that encodes its institutional knowledge, compounding its human and token capital.

Satya 谈把 loops(循环)作为 IP(知识产权/核心资产):> 这是我们第一次能够在人和数字系统之间创造一个真正的 cognitive loop(认知循环)。这很颠覆认知,因为它改变了我们甚至如何去概念化企业内部的工作。> 这意味着,真正的机会不在于挑出最好的 model(模型),而在于在模型之上建立一个 learning loop(学习循环),让 human capital(人力资本)和 token capital(token 资本)产生复利。你可以外包一个任务,甚至一份工作,但你永远无法外包自己的学习。> 在我看来,我们的优先事项必须是构建一个 frontier ecosystem(前沿生态系统),而不只是一个 frontier model(前沿模型),这样价值才能广泛流向每一家公司、每一个行业和每一个国家。在这样的生态里,每个组织都能拥有那个编码其 institutional knowledge(机构知识)的 learning loop,并让其 human capital 和 token capital 持续复利增长。

♥ 8↻ 0💬 56/14 · 19:05x.com ↗

@DevinAI link here!

@DevinAI 链接在这!

♥ 2↻ 0💬 16/14 · 18:24x.com ↗
2026 年 6 月 14 日 · 2 条 →

Last chance to fill out the annual AI Engineering Survey this weekend and win great Vercel + Notion + AIE tix! link below we had @devinai analyze registered attendee list and output a live chart of the people coming to the conference. it ended up being the single best data driven storytelling i've ever seen on what kind of community we are gathering in two weeks. survey link here! no lurking, fill it out pls

这是你在这个周末填写年度 AI Engineering Survey 并赢取超棒的 Vercel + Notion + AIE tix 的最后机会!链接见下。我们让 @devinai 分析了已注册参会者名单,并输出了一张关于将来参加 conference 的人群的实时图表。结果证明,这成了我所见过最出色的一次 data-driven storytelling(数据驱动叙事):它展现了两周后我们正在汇聚成一个什么样的 community。survey 链接在这里!别潜水了,快去填 pls

♥ 31↻ 5💬 176/13 · 21:31x.com ↗

@benthompson @digg

@benthompson @digg

♥ 1↻ 0💬 06/13 · 19:58x.com ↗
2026 年 6 月 13 日 · 3 条 →

how your email finds me (if youre waiting for a decision or reply pls dont take it personally im just in peak crunch mode for aie)

这就是你的邮件通常会在什么状态下被我看到(如果你在等一个决定或回复,pls 不要往心里去,我只是正处在 aie 的高压赶工期)

♥ 1↻ 0💬 06/13 · 07:34x.com ↗

## The Future Codebase After the PR dies, after the Code Review dies, i am seriously wondering if Git needs to die next. roughly 20-40% of code spend is just managing and updating merge conflicts. necessary evil? or legacy "horseless carriage"? cargo culting the past? we don't do line by line merge conflicts when we collaborate with human colleagues - instead we chat, suggest edits, do side comments, and an owner ships it. btw we also don't do CI/CD even collaborating on documents with serious legal/financial implications. maybe the future codebase looks more like a Notion or Linear database than .git objects. It will be less efficient, but more scalable. exactly the Salty Lesson.

## 未来的代码库 在 PR 消亡之后,在 Code Review 消亡之后,我很认真地在想,接下来是不是连 Git 也该消亡。大约 20–40% 的代码投入,只是花在管理和更新 merge conflict(合并冲突)上。这是必要之恶?还是一种遗留的“horseless carriage(无马马车)”?是在对过去进行 cargo cult(货物崇拜)式模仿吗?我们和人类同事协作时,并不会按行处理 merge conflict——相反,我们会聊天、提出修改建议、写侧边评论,然后由某个 owner(负责人)来发布。顺便说一句,即使是在协作处理具有重大法律/财务影响的文档时,我们也不会用 CI/CD。也许未来的代码库,看起来会更像 Notion 或 Linear 的数据库,而不是 .git objects。它的效率会更低,但可扩展性会更强。完全就是 The Salty Lesson。

♥ 90↻ 4💬 566/12 · 22:20x.com ↗

neat thing about developer exception engineering is: the happy paths are all happy in their own way. the unhappy paths are ~universally the same.

developer exception engineering 的一个妙处在于:happy path(正常路径)各有各的顺法;而 unhappy path(异常路径)却几乎到处都一样。

♥ 8↻ 0💬 46/12 · 19:28x.com ↗
2026 年 6 月 12 日 · 3 条 →

## On Loopcraft One might argue the entire game of the next century is to be able to stack loops as effectively as possible. In the early days of each phase, it will be valuable to know when to go **DOWN** a loop when things go wrong (for reliability)… but it will probably be more valuable to know how to go **UP** a loop as models improve (for leverage). If you don’t figure out how to do this, don’t be salty when you lose to those that do.

## 关于 Loopcraft 有人可能会说,下个世纪的整个游戏,就是要尽可能高效地堆叠 loops(循环)。在每个阶段的早期,当事情出错时,知道什么时候沿着 loop **向下**走会很有价值(为了 reliability,可靠性)……但随着 models(模型)变得更强,知道如何沿着 loop **向上**走,大概会更有价值(为了 leverage,杠杆效应)。如果你搞不清楚该怎么做,那么当你输给那些会做的人时,也别酸。

♥ 29↻ 3💬 186/12 · 05:37x.com ↗

the #1 thing that is driving me to build my own vibecoding platform rn is that none of them - and i lov vercel, cloudflare, netlify etc - none of them really close the loop for you in terms of setting you on the right path with errors and pinging you when shit fails (shit always fails) there's way too much "webmaster" infra to setup for every single project and i just want to do it once and for all, instead i'm being asked to npx posthog wizard here and npx arize skills there and it all just needs to be swallowed up into One Thing.

眼下,驱动我去自己做一个 vibecoding 平台的头号原因是:现有这些平台——而且我真的很喜欢 vercel、cloudflare、netlify 等——没有一个能真正帮你把 loop(闭环)补上,也就是说,在报错时把你引到正确路径上,并且在东西挂掉时提醒你(东西总是会挂)。每一个项目都要搭太多“webmaster”式的 infra(基础设施),而我只想一劳永逸地把这事做一次;结果现在却还要我这里 npx posthog wizard、那里 npx arize skills,这一切都应该被吞并进 One Thing。

♥ 69↻ 1💬 376/12 · 02:48x.com ↗

congrats to our friends @ona_hq on joining @openai! see their talk here for alpha on what’s next for Codex 👀

恭喜我们的朋友 @ona_hq 加入 @openai!想知道 Codex 接下来会怎样,可以看看他们这场 talk,里面有一些 alpha(一手消息)👀

♥ 27↻ 0💬 186/11 · 20:55x.com ↗
2026 年 6 月 11 日 · 1 条 →

wooh

wooh

♥ 8↻ 1💬 16/10 · 13:17x.com ↗
2026 年 6 月 10 日 · 3 条 →

btw insane amounts of alpha in telling claude code to "review my code for issues" on Fable rn while it is not pay per use be prepared to be in abject horror that you shipped anything to prod without a Fable Check™ first

顺便说一句,现在在 Fable 上让 claude code “review my code for issues”(检查我的代码是否有问题)简直有海量 alpha(超额价值 / 有用信息),尤其是在它目前还不是按次付费的时候;你最好做好心理准备:一想到自己以前没先做一次 Fable Check™ 就把东西发到 prod(生产环境),可能会陷入彻底的惊恐。

♥ 132↻ 5💬 296/9 · 23:40x.com ↗

for those keeping track at home it was 34 days between signing this deal and launching Mythos-class model GA to the world. building on @nvidia stack means you can just do things™.

给那些还在持续关注进展的人一个时间点:从签下这笔 deal(交易 / 合作)到把 Mythos-class model 正式以 GA(General Availability,正式可用)形式发布给全世界,中间只隔了 34 天。基于 @nvidia 的 stack(技术栈)来构建,意味着你真的可以“直接把事情做成”™。

♥ 10↻ 0💬 26/9 · 18:57x.com ↗

more charts of other tiers where its less stark including the vibe shift chart from here

这里还有更多其他 tier(层级 / 档位)的图表,那里的差异没这么夸张,其中也包括这里这张关于 vibe shift(整体风向 / 氛围转变)的图表。

♥ 10↻ 1💬 46/9 · 18:31x.com ↗
2026 年 6 月 9 日 · 2 条 →

@METR_Evals previously on @cognition_labs

@METR_Evals,此前在 @cognition_labs

♥ 5↻ 0💬 26/8 · 21:41x.com ↗

It's finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over 1000+ hours of maintainer validated software engineering work most frontier models cannot yet solve, much less solve with high quality. Cog had IOI Gold medalists and top code maintainers Look At The Data — FrontierCode includes 3000+ rubrics covering code quality and anticheat reward hacking plaguing other benchmarks. FC Diamond is so hard that Opus 4.8 scores 13.8%. Three eras of AI coding : Three eras of benchmarks 2021 • Autocomplete : HumanEval 2023 • Passing Tests: SWEBench, TerminalBench 2026 • Maintainable Code: FrontierCode to me the most beautiful chart when I requested a special historical run into all extant old models, the data was finding that the easiest third of FC tasks (in FC Extended) were rapidlly and suddenly solved over late 2025 - Opus almost doubled from a 41% pass rate to 74% in 4 months. This describes the "WTF happened in Dec 2025" vibe shift that a lot of folks from @dhh to @karpathy have called out: it is the difference between getting 95% success in 2 rerolls vs 6, making it finally feasible to go up the next layer of abstraction in agentic coding, eg @GeoffreyHuntley's ralph loops or @bcherny's /goals or @steipete's "loops that prompt your agents" without fearing too much that things go off the rails. My guess: as AI accelerates from here, each FrontierCode tier will saturate in sequence, hopefully ~annually. I've already asked the team to prepare FrontierCode 2027.... The old mountains will be destroyed. Their rubble becomes regolith. And from that regolith, the next model forest grows. Circle of life.

它终于发布了!!!@METR_Evals 发现,超过一半的 SWEBench 结果都是无法合并的 slop(低质产出)。FrontierCode 代表了 1000+ 小时经 maintainer(维护者)验证的软件工程工作,而大多数 frontier models(前沿模型)目前还无法解决这些工作,更不用说高质量地解决了。Cog 拥有 IOI Gold medalists 和顶级代码维护者——Look At The Data——FrontierCode 包含 3000+ 条 rubrics(评分细则),覆盖代码质量以及困扰其他 benchmarks(基准测试)的 anticheat reward hacking(反作弊奖励破解/刷分)问题。FC Diamond 难到连 Opus 4.8 的得分都只有 13.8%。AI coding 的三个时代:benchmarks(基准)的三个时代。2021 • Autocomplete:HumanEval。2023 • Passing Tests:SWEBench、TerminalBench。2026 • Maintainable Code:FrontierCode。对我来说,最漂亮的一张图是:当我要求对所有现存旧模型做一次特别的历史回跑时,数据显示 FC 任务中最简单的三分之一(在 FC Extended 中)在 2025 年末被迅速且突然地攻克了——Opus 在 4 个月内几乎翻倍,从 41% 的 pass rate(通过率)升到 74%。这描述了很多人——从 @dhh 到 @karpathy——都指出的那种“2025 年 12 月到底发生了什么”的 vibe shift(氛围转变):区别在于,你是在 2 次 rerolls(重试)内拿到 95% 成功率,还是要 6 次;这最终让 agentic coding(agent 驱动编程)向上进入下一层抽象成为可行之事,例如 @GeoffreyHuntley 的 ralph loops、@bcherny 的 /goals,或 @steipete 所说的“让你的 agents 接收提示并循环执行的 loops”,而不必太担心事情彻底失控。我的猜测是:随着 AI 从这里开始继续加速,FrontierCode 的每个 tier(层级)都会依次饱和,希望大约是每年一个。我已经让团队开始准备 FrontierCode 2027 了……旧日的高山将被摧毁,它们的碎石会变成 regolith(月壤/风化层);而从那 regolith 之中,下一片模型森林将会生长。生命循环。

♥ 631↻ 59💬 706/8 · 20:27x.com ↗
2026 年 6 月 7 日 · 1 条 →

one popular theory is that research paper alpha* and lab publishing ~died when researchers realized that instead of fighting with marketing depts they could simply walk out the door and get >$100m for their legally protected tacit knowledge gained california non-noncompetes have a bigger impact on knowledge spreading than github, arxiv, and huggingface combined *btw this is a motivator for me to set up @aidotengineer as a product-centric industry conference to complement the paper-centric research conferences

一个很流行的理论是:当研究人员意识到,与其和 marketing depts(市场部门)周旋,不如直接走出公司大门,凭借他们依法受保护的 tacit knowledge(隐性知识)拿到超过 1 亿美元时,research paper alpha* 和 lab publishing(实验室论文发表)就大致“死掉”了。California 的 non-noncompetes(禁止竞业限制)对知识传播的影响,比 GitHub、arXiv 和 Hugging Face 加起来还要大。*顺便说一句,这也是我建立 @aidotengineer、把它做成一个以产品为中心的行业 conference(会议),用来补充那些以论文为中心的研究 conference(会议)的动机之一。

♥ 134↻ 5💬 246/7 · 01:27x.com ↗
2026 年 6 月 6 日 · 3 条 →

a smarter alternative to "always use plan mode": always frame your task as a question, so that the model is invited to push back and rate the quality of the idea/suggest alternatives, rather than blindly execute what you SAID to do (which is often not precisely what you MEANT) literally just appending "?" to the end of your prompt often does it

相比“总是使用 plan mode”的更聪明替代方案是:始终把你的任务表述成一个问题,这样模型就会被引导去提出异议、评估这个想法的质量并给出替代方案,而不是盲目照着你“说”要做的事去执行(而那往往并不精确等于你“本来想表达”的意思);很多时候,真的只要在你的 prompt 末尾加上一个“?”就行了

♥ 119↻ 6💬 306/6 · 02:18x.com ↗

i love being (for now) bdfl for aie because i can do cheeky shit like the AGI pills we did in london and also this

我很喜欢自己(至少现在)作为 aie 的 bdfl,因为这样我就能搞一些我们在 london 做过的 AGI pills 之类的皮操作,还有这个也是

♥ 24↻ 0💬 186/5 · 22:47x.com ↗

@aiDotEngineer lmao designer vincent back at it again with the frontier capability tests

@aiDotEngineer 哈哈哈,designer vincent 又回来整 frontier capability tests 了

♥ 4↻ 0💬 16/5 · 21:40x.com ↗
2026 年 6 月 5 日 · 3 条 →

ibid

同上

♥ 5↻ 0💬 16/5 · 00:16x.com ↗

about time that the leading database company born in Singapore actually had Singapore investors take it seriously!

这家诞生于 Singapore 的头部数据库公司,也差不多该让 Singapore 的投资者认真对待它了!

♥ 60↻ 3💬 106/4 · 20:06x.com ↗

Finally! the first eval ship from cog!!!!!!!!!! 👼🏼 To contextualize: @METR_Evals cap out at ~16 hours. Cog has private enterprise evals up to 100hrs, and is confident enough to put a financial guarantee on it 🤯 METR dataset: ML eng, GPU kernels, cybersecurity > "METR (2026) used a combination of GPT-4o and GPT-5 to estimate the human-equivalent times from compressed Claude Code transcripts. These transcripts were collected from 7 METR technical staff on 34 sessions labeled on human ground truth". rlog​ of 0.83 Cog dataset: real life java/typescript/python/c# feature dev, bugfixes, migrations > "We collected a ground-truth dataset by asking Devin users to review recent representative sessions, and estimate how long each completed session would have taken without Devin. Our dataset consists of 258 sessions from 126 users across a diverse set of enterprise customers." rlog​ of 0.74 on held out set this is pioneering real world evals work and part 1 of a broader frontier code evals drop that I'm really looking forward to writing up. huge kudos to @annarmitchell and @ryanbai1412 for leading the unglamorous last mile data collection!!

终于!来自 cog 的第一份 eval(评测)发布了!!!!!!!!👼🏼 给一点背景:@METR_Evals 的上限大约是 16 小时。Cog 有私有的 enterprise(企业)eval,最长可到 100 小时,而且他们甚至有信心为此附上 financial guarantee(财务担保)🤯 METR 数据集:ML eng、GPU kernels、cybersecurity > “METR (2026) used a combination of GPT-4o and GPT-5 to estimate the human-equivalent times from compressed Claude Code transcripts. These transcripts were collected from 7 METR technical staff on 34 sessions labeled on human ground truth”. rlog 为 0.83。Cog 数据集:真实世界中的 java/typescript/python/c# 功能开发、bug 修复、迁移 > “We collected a ground-truth dataset by asking Devin users to review recent representative sessions, and estimate how long each completed session would have taken without Devin. Our dataset consists of 258 sessions from 126 users across a diverse set of enterprise customers.” 在留出测试集(held out set)上的 rlog 为 0.74。这是开创性的真实世界 eval 工作,也是更大范围前沿代码 eval 发布的第 1 部分,我非常期待把它详细写出来。非常感谢 @annarmitchell 和 @ryanbai1412 牵头完成了这些并不光鲜、但至关重要的最后一公里数据收集工作!!

♥ 179↻ 12💬 286/4 · 19:03x.com ↗
2026 年 6 月 4 日 · 3 条 →

@HamiltonMusical the most viral thing i have ever done and its a bootleg hamilton sitzprobe not anything ai related

@HamiltonMusical 是我做过传播最火的东西,而且那还是一段盗录的 Hamilton sitzprobe,根本不是什么和 AI 有关的内容

♥ 1↻ 0💬 16/4 · 04:48x.com ↗

you guys know where this is going right

你们都知道这接下来会往哪发展,对吧

♥ 75↻ 0💬 156/4 · 03:11x.com ↗

@jgreze will speak on this at gathering all the top agent labs. lfg

@jgreze 会在 gathering 上谈这个,届时所有顶尖的 agent(智能体)实验室都会到场。lfg

♥ 7↻ 0💬 26/3 · 20:59x.com ↗
2026 年 6 月 3 日 · 3 条 →

@saranormous codex is agi man oneshotted this, no notes

@saranormous,codex 简直是 AGI,兄弟一发就把这事搞定了,没啥可说的

♥ 7↻ 0💬 46/3 · 06:43x.com ↗

probably the best reward function for reasoning efficiency i've seen

这可能是我见过用于推理效率的最好的 reward function(奖励函数)

♥ 23↻ 0💬 106/3 · 06:33x.com ↗

@jacobeffron

@jacobeffron

♥ 1↻ 0💬 16/3 · 06:13x.com ↗
2026 年 6 月 1 日 · 3 条 →

@soumithchintala @pewdiepie @opencode

@soumithchintala @pewdiepie @opencode

♥ 7↻ 0💬 06/1 · 01:23x.com ↗

just a small zoom out on the vibe shift: in Feb 2025 @soumithchintala was talking about his dream of personal, local, private agents, most people didn't believe him. it's June 2026 and @pewdiepie has just released his vibecoded @opencode wrapper that is a complete personal AI productivity suite including email, docs, and calendar. top of HN, easily >1m views, >10k stars in a day. if your Knowledge Work Agents startup can't beat pewdiepie you might as well pack up and go home at this point, his is the benchmark for what you can DIY.

对这种氛围转变(vibe shift)稍微拉远一点看:在 2025 年 2 月,@soumithchintala 还在谈论他关于个人、本地、私有 agent 的梦想,当时大多数人并不相信他。现在是 2026 年 6 月,@pewdiepie 刚刚发布了他用 vibecoded 方式做出来的 @opencode wrapper,已经是一整套完整的个人 AI 生产力套件,包含 email、docs 和 calendar。登上 HN 榜首,轻松超过 100 万浏览,一天内超过 1 万 stars。要是你的 Knowledge Work Agents 创业公司连 pewdiepie 都打不过,那这时候基本可以收摊回家了——他做出来的东西,就是现在你自己动手(DIY)能达到什么水平的 benchmark(基准)。

♥ 137↻ 5💬 276/1 · 01:18x.com ↗

every evals/analytics startup is going through a onetime generational upgrade into a continual learning platform in 2026 many will fail but as always the tasteful ones win

到了 2026 年,每一家做 evals/analytics 的创业公司,都在经历一次一代人只会发生一次的升级:从原来的形态转向 continual learning platform(持续学习平台);很多都会失败,但一如既往,真正有品位的那些会赢。

♥ 245↻ 8💬 485/31 · 22:00x.com ↗
2026 年 5 月 28 日 · 3 条 →

[1]

♥ 4↻ 0💬 05/27 · 03:35x.com ↗

ai infra is going VERTICAL

AI infra 正在走向垂直化

♥ 178↻ 5💬 265/27 · 02:34x.com ↗

last 4 days to submit talks: this is our first year featuring PREPRINT poster sessions for research papers as well - as @bclavie pointed out we need a separate process for this but you can submit here for now!

距离提交演讲 proposal 还剩最后 4 天:这是我们第一年也为 research papers 设立 PREPRINT 海报 session——正如 @bclavie 指出的那样,我们确实需要为此单独设置一个流程,但你现在仍然可以先在这里提交!

♥ 16↻ 2💬 45/26 · 20:34x.com ↗
2026 年 5 月 21 日 · 3 条 →

pod link and more ctx

pod 链接以及更多 ctx(上下文)

♥ 10↻ 0💬 25/20 · 19:23x.com ↗

btw we did a bake off of Exa vs competitors and it took all of 1.5 hrs for the team to unanimously converge on exa lol. so proud to see my former landlords crush it - time travel back to last year and listen to a pre pmf @WilliamBryk to understand how to spot companies on a generational tear

顺便说一句,我们把 Exa 和竞品做了一轮 bake-off(对比测试),整个团队只花了 1.5 小时就一致认定 Exa 胜出,lol。很自豪看到我以前的房东们大杀四方——把时间倒回去年,去听听 pre pmf(达到 product-market fit 之前)的 @WilliamBryk,你就会明白该如何识别那种正在经历代际级爆发的公司

♥ 222↻ 9💬 275/20 · 19:22x.com ↗

very belated but in retrospect i think @sama's mythical "build a business that gets better when models get better" is basically what I called Agent Labs here. seeing a very direct correlation with model performance and agent lab revenue, discontinuity in Q4 2025 (clip from @patrickc's stripe sessions)

虽然现在说已经晚了很多,但事后看,我觉得 @sama 那句近乎神话般的话——“build a business that gets better when models get better”——基本上就是我在这里所说的 Agent Labs。我看到 model(模型)性能和 agent lab 营收之间存在非常直接的相关性,并且在 2025 年 Q4 出现了不连续跃迁(摘自 @patrickc 的 stripe sessions)

♥ 46↻ 4💬 245/20 · 15:20x.com ↗
2026 年 5 月 20 日 · 3 条 →

oh no contextual got windsurfed

哦不,contextual 被 windsurfed 了

♥ 11↻ 0💬 45/20 · 07:23x.com ↗

there's 4 parts to this AI SDLC 1. have ~50 tests in place, with instructions to add more, including "make a memory that whenever you do browser e2e tests, use computer vision to visually spot check design and ux issues as well on mobile/desktop/ipad/ultrawide resolutions" 2. "/plan break up & edit hot paths so you isolate files for easier editing and reading. add proper logging and error boundaries/handling while you do it. what else should we refactor for maintainability/performance/ai editing?" 3. (with plan) "you can break backward compatibility. first map out all the remaining work, then proceed on this next slice, do not stop until all work is done, periodically stop to commit, deploy and test but do not stop until all work is done" 4. [periodically spot check deployed functionality and /steer bugs as it goes along]

这个 AI SDLC(软件开发生命周期)有 4 个部分:1. 先准备好大约 50 个测试,并附上继续添加更多测试的说明,包括“建立一条 memory(记忆):每当你做 browser e2e tests(浏览器端到端测试)时,也要用 computer vision(计算机视觉)从视觉上抽查 design 和 ux 问题,并覆盖 mobile/desktop/ipad/ultrawide 分辨率”;2. “用 /plan 拆分并编辑 hot paths(关键路径),这样你就能隔离文件,便于编辑和阅读。在这个过程中补上合适的 logging(日志记录)以及 error boundaries/handling(错误边界/错误处理)。另外,为了 maintainability/performance/ai editing(可维护性/性能/AI 编辑),我们还应该重构什么?”;3. (基于 plan)“你可以破坏 backward compatibility(向后兼容性)。先把所有剩余工作都梳理出来,然后继续推进下一块,不要停,直到所有工作都完成;期间可以定期停下来 commit、deploy 和 test,但在所有工作完成之前不要停”;4. [定期抽查已部署的功能,并在推进过程中用 /steer 处理 bugs]

♥ 8↻ 0💬 65/19 · 23:19x.com ↗

rsi is here. jesus

rsi 来了。天啊

♥ 25↻ 0💬 25/19 · 17:34x.com ↗
2026 年 5 月 19 日 · 3 条 →

taking bets for vercel and supabase rn

我现在押注 vercel 和 supabase。

♥ 13↻ 0💬 65/19 · 06:44x.com ↗

volunteer here !

在这里报名当志愿者!

♥ 13↻ 0💬 15/19 · 00:15x.com ↗

this seems quite doable in the space of a single 2-3 hour workshop — any brave soul want to try to livecode this for people as a learning exercise?

这件事看起来相当可行,完全可以放在一场 2–3 小时的 workshop(工作坊)里完成——有没有勇敢的人愿意把它作为学习练习,现场 livecode(直播编码)给大家看?

♥ 409↻ 11💬 225/18 · 20:53x.com ↗
2026 年 5 月 18 日 · 3 条 →

@gabrielchua the agentic excel thing is basically what u get when u expand the side panel to be the main thing

@gabrielchua,所谓那个 agentic excel 的东西,基本上就是把侧边栏扩展成主要界面后你会得到的样子

♥ 3↻ 0💬 35/18 · 02:33x.com ↗

some of us doing kaya toast breakfast here at 11am if u are still around

我们这里有些人打算上午 11 点去吃 kaya toast 早餐,如果你还在附近的话

♥ 8↻ 0💬 25/18 · 01:51x.com ↗

🇸🇬

🇸🇬

♥ 87↻ 6💬 65/18 · 01:27x.com ↗
2026 年 5 月 17 日 · 3 条 →

@sarahookr @calcsam

@sarahookr @calcsam

♥ 0↻ 0💬 05/17 · 07:01x.com ↗

AIE coming to India soon!

AIE 很快就会来到 India!

♥ 78↻ 3💬 215/17 · 05:55x.com ↗

@Gavriel_Cohen i have to say his social media team is better than mine wtf. first pull on youtube.

@Gavriel_Cohen 我得说,他的社交媒体团队比我的强,wtf。第一次在 youtube 上发力。

♥ 7↻ 0💬 05/16 · 11:32x.com ↗
2026 年 5 月 16 日 · 3 条 →

gotta say Codex is completely unrecognizable from 3 months ago. guys went extreme founder mode on this thing @gabrielchua was demoing this and i was like “you guys have agentic excel on mac”

得说一句,Codex 跟 3 个月前比已经完全认不出来了。你们这帮人真是对这东西开启了极致 founder mode(创始人模式)。@gabrielchua 当时在演示这个,我心里就在想:“你们这是把 agentic excel(具备 agent 能力的 Excel)做到了 Mac 上啊。”

♥ 35↻ 1💬 205/16 · 03:43x.com ↗

@Gavriel_Cohen @thsottiaux head of AI Govtech at Singapore estimates 1.3 billion agents in the country in the next 2 years and is building a national MCP gateway @dsp_

@Gavriel_Cohen @thsottiaux 新加坡 AI Govtech(政府科技中的 AI)负责人估计,未来 2 年这个国家里会有 13 亿个 agents(智能体),并且正在打造一个国家级的 MCP gateway。@dsp_

♥ 30↻ 0💬 35/16 · 02:09x.com ↗

@Gavriel_Cohen and @thsottiaux casually dropping some hints on the Codex roadmap in his keynote!

@Gavriel_Cohen 和 @thsottiaux 在他的 keynote(主题演讲)里,很随意地放出了一些关于 Codex roadmap(路线图)的暗示!

♥ 21↻ 0💬 25/16 · 01:56x.com ↗
2026 年 5 月 15 日 · 3 条 →

Apparently at @AIEMiami geoff complained about @SAPConcur being dead software and a SAP guy was in the audience and invited him to SAP to advise on how to do AI transformation for 6800 employees TLDR he made fun of SAP, and SAP… concurred

显然,在 @AIEMiami 上,geoff 抱怨说 @SAPConcur 是死掉的软件,而现场观众里刚好有个 SAP 的人,于是邀请他去 SAP,为 6800 名员工提供关于如何做 AI 转型的建议。TLDR(太长不看版):他拿 SAP 开了个玩笑,而 SAP…… concurred(也“同意”了;双关 SAPConcur)。

♥ 92↻ 0💬 135/15 · 02:32x.com ↗

oh no

哦不。

♥ 1↻ 0💬 15/15 · 01:51x.com ↗

Blogs die when they come from "the ____ team" instead of named individuals With great ownership comes great accountability

当 Blog 不是出自有名有姓的个人,而是来自“某某 team”时,它们就开始走向死亡。越有 ownership(主人翁意识/责任归属),就越有 accountability(问责)。

♥ 97↻ 2💬 115/15 · 00:14x.com ↗
2026 年 5 月 12 日 · 3 条 →

my highlights

我划出的重点

♥ 1↻ 0💬 05/12 · 04:36x.com ↗

[1]

♥ 7↻ 0💬 45/12 · 00:03x.com ↗

I believe the kids call this "@thinkymachines just brutally framemogged gdm and oai". basically everyone's definition of "realtime" just got a massive frciking upgrade

我觉得现在小孩会把这叫作“@thinkymachines 直接把 gdm 和 oai 在叙事框架上狠狠干翻了”。基本上,所有人对“realtime(实时)”的定义刚刚都被大幅他妈地升级了一遍

♥ 798↻ 31💬 335/11 · 22:06x.com ↗
2026 年 5 月 11 日 · 1 条 →

on build vs buy saas cc @levie for corrections

关于自建(build)还是购买(buy)SaaS,抄送 @levie 以便指正

♥ 65↻ 1💬 205/10 · 20:25x.com ↗
2026 年 5 月 10 日 · 3 条 →

@VivianBala we will finally show the world how it is done.

@VivianBala,我们终于要向全世界展示这件事该怎么做了。

♥ 8↻ 0💬 15/10 · 07:06x.com ↗

OK I'VE BEEN SO EXCITED i could barely keep this a secret all week and it's finally official MY HOME COUNTRY'S MINISTER OF FOREIGN AFFAIRS (equiv to Secretary of State) IS A HUGE NANOCLAW FAN (check @VivianBala, that's really him, not an intern) AND WILL BE KEYNOTING @AIDOTENGINEER SINGAPORE (with NanoClaw creator @Gavriel_Cohen right after) NEXT WEEK Usecases like his are what I have been hoping to promote with the international AIE partnerships and @agrimsingh and @SherryYanJiang crushed it with this one. governments waking up to AI and joining @aiDotEngineer: UK: Chief AI Officer Singapore: Cabinet Minister who's next??

好吧,我真的兴奋坏了,整整一周几乎都憋不住这个秘密,现在终于正式官宣了——我祖国的 Minister of Foreign Affairs(相当于 Secretary of State)竟然是个超级铁杆的 NANOCLAW 粉丝(去看 @VivianBala,那确实是他本人,不是什么实习生),而且他下周还将担任 @AIDOTENGINEER SINGAPORE 的 keynote speaker(随后紧接着就是 NanoClaw 的创造者 @Gavriel_Cohen)。像他这样的 use case(使用案例),正是我一直希望通过国际 AIE 合作伙伴关系来推动的;而 @agrimsingh 和 @SherryYanJiang 这次真的把这件事做得太漂亮了。各国政府正在觉醒,开始拥抱 AI,并加入 @aiDotEngineer:UK:Chief AI Officer;Singapore:Cabinet Minister——下一个会是谁??

♥ 47↻ 3💬 95/10 · 07:04x.com ↗

wondering if @embirico has numbers on what % of codex users use this mode and how much it has gone up over the last month its a decent proxy for alignment/agent adoption

想知道 @embirico 那边有没有数据:codex 用户中有多少百分比在使用这个 mode(模式),以及过去一个月这个比例增长了多少;这可以作为 alignment(对齐)/agent 采用情况的一个还不错的 proxy(代理指标)。

♥ 10↻ 0💬 105/10 · 06:38x.com ↗
2026 年 5 月 9 日 · 2 条 →

@nikitabier @business some good sourcing. seems potentially state level.

@nikitabier @business 消息来源不错。看起来可能涉及州政府层面。

♥ 2↻ 0💬 05/8 · 18:44x.com ↗

@nikitabier @business bloomberg being suddenly interested in your take on developer experience and ai coding tools is the new "sexy singles in your area"

@nikitabier @business Bloomberg 突然对你关于 developer experience 和 AI coding tools 的看法感兴趣,简直就是新版“你所在地区的火辣单身人士”广告。

♥ 7↻ 0💬 65/8 · 16:06x.com ↗
2026 年 5 月 8 日 · 3 条 →

docusign !!? fuck docusign with a sharp stick

docusign !!? 去他妈的 docusign,拿根尖棍子狠狠干它

♥ 1↻ 0💬 05/8 · 06:11x.com ↗

i was going to send him my loom showing him gustos bugs and loom loomed on me

我本来要把我的 loom 发给他,给他看 gusto 的 bugs,结果 loom 反倒把我给“loom”了

♥ 5↻ 0💬 25/8 · 05:50x.com ↗

be aware of this kind of phishing. i was almost tricked. cc @nikitabier @business

小心这种 phishing(网络钓鱼)。我差点就上当了。cc @nikitabier @business

♥ 125↻ 5💬 205/8 · 04:00x.com ↗
2026 年 5 月 6 日 · 3 条 →

OAI 850B valuation, ~30B ARR now Ant ~900B valuation, ~44B* ARR now *revenue recognized differently, per Denise Dresser its probably closer to 8-10B lower if using OAI methodology chart reconstructed from wsj by me

OAI 现在的估值是 850B,当前 ARR(年度经常性收入)约为 30B;Ant 现在的估值约为 900B,当前 ARR 约为 44B*。*收入确认方式不同;根据 Denise Dresser 的说法,如果使用 OAI 的方法,可能实际上要低 8–10B 左右。图表由我根据 wsj 重建。

♥ 3↻ 1💬 25/4 · 23:14x.com ↗

see the talk version, out now thanks to @steveruizok

可以看演讲版本,现已发布,感谢 @steveruizok

♥ 1↻ 0💬 15/4 · 15:53x.com ↗

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keynotes i gave a "blank check" to based on his sincere support since our very earliest days + when in europe we must feature the DevOps guy. he didnt disappoint!

顺便说一句,这个表现非常好;如果你想在所有内容的 firehose(信息洪流)里用“popular vote”来做筛选的话。@patrickdebois 是我做主题演讲的其中一个分会场;基于他从我们最早期开始就给予的真诚支持,我给了他一张“blank check(全权信任)”;而且人在欧洲,我们就必须安排那位 DevOps 大神出场。他没有让人失望!

♥ 40↻ 1💬 55/4 · 15:53x.com ↗
2026 年 5 月 5 日 · 3 条 →

OAI 850B valuation, ~30B ARR now Ant ~900B valuation, ~44B* ARR now *revenue recognized differently, per Denise Dresser its probably closer to 8-10B lower if using OAI methodology chart reconstructed from wsj by me

OAI 现在的估值是 850B,当前 ARR(年度经常性收入)约为 30B;Ant 现在的估值约为 900B,当前 ARR 约为 44B*。*收入确认方式不同;根据 Denise Dresser 的说法,如果使用 OAI 的方法,可能实际上要低 8–10B 左右。图表由我根据 wsj 重建。

♥ 3↻ 1💬 25/4 · 23:14x.com ↗

see the talk version, out now thanks to @steveruizok

可以看演讲版本,现已发布,感谢 @steveruizok

♥ 1↻ 0💬 15/4 · 15:53x.com ↗

this one is doing v well btw if you want the popular vote filter on the firehose of all the things @patrickdebois was one of the track keynotes i gave a "blank check" to based on his sincere support since our very earliest days + when in europe we must feature the DevOps guy. he didnt disappoint!

顺便说一句,这个表现非常好;如果你想在所有内容的 firehose(信息洪流)里用“popular vote”来做筛选的话。@patrickdebois 是我做主题演讲的其中一个分会场;基于他从我们最早期开始就给予的真诚支持,我给了他一张“blank check(全权信任)”;而且人在欧洲,我们就必须安排那位 DevOps 大神出场。他没有让人失望!

♥ 40↻ 1💬 55/4 · 15:53x.com ↗
2026 年 5 月 4 日 · 3 条 →

til the show is free on youtube!

到节目在 YouTube 上免费观看为止!

♥ 0↻ 0💬 15/4 · 01:41x.com ↗

ok @deepfates @mada299 it took me 3 years to do my second short story but i did one

好吧,@deepfates @mada299,我花了 3 年时间才完成我的第二篇短篇故事,但我还是做出来了一篇。

♥ 2↻ 0💬 05/3 · 19:46x.com ↗

[1]

♥ 18↻ 0💬 55/3 · 19:44x.com ↗
2026 年 5 月 3 日 · 1 条 →

Much respect to @tokengobbler who shutdown Vibe-kanban live onstage at AIE Europe - still with 30,000 MAU, and still living on as an open source project. "Everyone who is making money is doing 2 things: selling to enterprise, and reselling tokens. We were doing neither." surprisingly not the first company to shutter at AIE but there's a lot to learn from the process and the software engineering retrospective from 2021-2025 will stick in my mind!

向 @tokengobbler 致以极大敬意——他在 AIE Europe 的台上现场关闭了 Vibe-kanban;它当时仍有 30,000 MAU(月活跃用户),并且仍以开源项目的形式延续着生命。“所有在赚钱的人都在做两件事:卖给 enterprise(企业客户),以及转售 token(代币)。而我们两件都没做。” 令人意外的是,这竟然不是第一家在 AIE 关闭的公司,但这个过程以及这份关于 2021-2025 年 software engineering(软件工程)的复盘,确实有很多值得学习的地方,也会一直留在我的脑海里。

♥ 56↻ 4💬 115/3 · 01:44x.com ↗
2026 年 5 月 1 日 · 3 条 →

@jacobeffron tagging @dat_attacked

@jacobeffron 标记了 @dat_attacked

♥ 0↻ 0💬 05/1 · 05:26x.com ↗

@jacobeffron fuller writeup:

@jacobeffron 的更完整写稿:

♥ 1↻ 0💬 25/1 · 04:54x.com ↗

i said on @jacobeffron's pod recently that "coding agents breaking containment" is the breakout theme of the year. i meant it - this is the year all knowledge workers, not just coders, get AGI-pilled. for the AIE EU closing note ( I gave a short talk on how we use agents to run @aidotengineer as a Tiny Team that now serves ~1m unique developers a month for free all around the world, for everything from CMS to renting lobster inflatables. yes I use @openclaw personally and as a team we use @cognition's Devin and @townai, but this isn't about any one agent; it's about all of them, and how you are probably not trying hard enough to use them for daily knowledge work. i hope this gives you agent productivity ideas for you and your team.

我最近在 @jacobeffron 的 pod 上说过,“coding agents(编程 agent)突破隔离边界”是今年最重要的爆发主题。我是认真的——今年将是所有 knowledge workers(知识工作者),而不只是 coders(程序员),都被 AGI-pilled 的一年。为了 AIE EU 的 closing note(闭幕致辞)(我做了一个简短分享,讲我们如何使用 agents(agent)把 @aidotengineer 作为一个 Tiny Team 来运营,而它现在每月为全球约 100 万 unique developers(独立开发者/唯一开发者用户)免费提供服务,应用场景从 CMS 到租赁 lobster inflatables 都有。是的,我个人会用 @openclaw,我们团队也会用 @cognition 的 Devin 和 @townai,但这并不是关于某一个 agent;而是关于所有 agent,以及你很可能还没有足够努力地把它们用于日常 knowledge work(知识工作)。我希望这能为你和你的团队带来一些关于 agent productivity(agent 生产力)的想法。

♥ 18↻ 1💬 195/1 · 04:23x.com ↗
2026 年 4 月 30 日 · 3 条 →

why does the algo hate this post

为什么 algo 会讨厌这条帖子

♥ 0↻ 0💬 04/30 · 02:32x.com ↗

> be me > "the internet is polluted by ai slop, we need low-background tokens" > "wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us" > all the existing vintage models are like <4B > we need a chat tuned 13B vintage model > assemble avengers of ML incl the GPT-1/2 guy > need vintage tokens > train new vintage OCR model for old books, newspapers, periodicals, scientific journals, patents, and case law > need vintage RLHF but cant use chat > synthesize RLHF pairs from historical texts with regular structure eg etiquette manuals, letter-writing manuals, cookbooks, dictionaries, encyclopedias, and poetry and fable collections, shove it into ChatML > train it > future knowledge still got in somehow > dammit.jpg > train new SOTA document-level n-gram-based anachronism classifier > meticulously curate hundreds of billions of pre-1931 tokens (public domain) > train it > ok! it checks out vs our FineWeb baseline! > release it > it's the most confidently racist model ever released by humankind > mfw

> 设想我是我自己 > “互联网已经被 ai slop(AI 垃圾内容)污染了,我们需要 low-background tokens(低背景噪声 token)” > “要是我们能穿越时间,看看 100 年前的祖先会对我们说什么,不是很酷吗” > 现有的所有 vintage models(复古模型)基本都小于 4B > 我们需要一个经过 chat tuned(聊天调优)的 13B vintage model > 集结 ML 的 avengers,包括那个 GPT-1/2 guy > 需要 vintage tokens > 为 old books、newspapers、periodicals、scientific journals、patents 和 case law 训练新的 vintage OCR model > 需要 vintage RLHF,但又不能用 chat > 从具有规则结构的 historical texts 中合成 RLHF pairs,比如 etiquette manuals、letter-writing manuals、cookbooks、dictionaries、encyclopedias,以及 poetry 和 fable collections,然后一股脑塞进 ChatML > 训练它 > 结果还是不知怎么混进了 future knowledge(未来知识) > dammit.jpg > 再训练一个新的、SOTA 的 document-level、基于 n-gram 的 anachronism classifier(时代错置分类器) > 精心整理出数千亿个 1931 年前的 token(public domain,公版) > 训练它 > 好!跟我们的 FineWeb baseline 对比,确实过关了! > 发布它 > 结果它成了人类有史以来发布过的最自信的 racist model(种族主义模型) > mfw

♥ 37↻ 2💬 94/30 · 00:51x.com ↗

i havent done the work to compare it to peers but i'm just excited that we have a base model and honestly for all the people that complained about the death of the completions API (@deepfates ? or deepfates adjacent) not enough people are experimenting with weird usages and finetunes of the base models we DO get

我还没做足够的工作把它和同类模型比较,但我只是单纯很兴奋,因为我们现在有了一个 base model;老实说,那些曾经抱怨 completions API 死掉的人(@deepfates?或者跟 deepfates 一路的人)里,去实验我们现有这些 base models 的各种奇怪用法和 finetune(微调)的人,实在还不够多

♥ 6↻ 0💬 24/30 · 00:14x.com ↗
2026 年 4 月 28 日 · 1 条 →

ok we have a tiktok account now with some BTS

好的,我们现在有一个 TikTok 账号了,里面有一些 BTS 内容。

♥ 2↻ 0💬 14/28 · 01:16x.com ↗
2026 年 4 月 27 日 · 3 条 →

btw we are cooking something with @hhua_ (not final yet but keep calendar open after ICML in Seoul)

顺便说一句,我们正在和 @hhua_ 一起筹备点东西(还没最终敲定,但 ICML in Seoul 之后请把日程空出来)

♥ 48↻ 0💬 154/25 · 19:44x.com ↗

wow another engineer on the “code is not cheap” train

哇,又有一位 engineer 加入了“code 并不便宜”这趟列车

♥ 6↻ 0💬 44/25 · 19:40x.com ↗

fun to think about what the pm thinks vs what the engineer thinks in this scenario

想想在这种场景里 pm 会怎么想、engineer 又会怎么想,还挺有意思的

♥ 44↻ 0💬 94/25 · 07:58x.com ↗
2026 年 4 月 22 日 · 3 条 →

give us back Sky

把 Sky 还给我们

♥ 3↻ 0💬 14/21 · 00:41x.com ↗

the Codex x @skybysoftware acquisition may have been one of the best @openai deals made in the last year. I've been waiting for "real" computer use since @romainhuet demoed the ChatGPT App with 4o Vision at AIEWF 2024... and only now it's really, actually rolling out in a usable fashion.

Codex x @skybysoftware 的收购,可能是过去一年里 @openai 做过的最好的交易之一。自从 @romainhuet 在 AIEWF 2024 用 4o Vision 演示 ChatGPT App 以来,我就一直在等“真正的” computer use(计算机使用)……直到现在,它才终于以一种真正可用的方式开始推出。

♥ 93↻ 4💬 174/20 · 22:57x.com ↗

and @dexhorthy is quoting Z/L continuum in AIE Miami!! idea catching on @altryne

而且 @dexhorthy 还在 AIE Miami 提到了 Z/L continuum,这个想法正在 @altryne 那里传播开来!!

♥ 30↻ 2💬 54/20 · 13:41x.com ↗
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers