BuildSpeak每日 builder 文摘
今日归档生词本关于
🐦 X · 动态Aaron Levie @levie· 2026 年 7 月 15 日· 307 词 · 约 2 分钟

Aaron Levie · @levie

SPACE 播放 / 暂停·←→ 上一句 / 下一句
One of the many properties that code has that makes it highly amenable to agents is that you can more or less quickly test it. You can either go see if the application works manually, or you can actually run a test on what you built. Most other areas of work don’t have this benefit. You only get the testing when the final product hits the real world in some capacity - a stock trade is executed, a contract is negotiated, a sales pitch is delivered, and so on. There’s probably going to be a whole new set of opportunities for how we begin to test the rest of work in this way. Ultimately it will mean more agents being layered into workflows. It also means we need much better evals on most of our workflows. Most work today in enterprises doesn’t have an associated eval to know if something broke or improved with a model, prompt, or system change. The enterprises that are able to eval their knowledge work the best also stand to gain the most from AI. Will become a critical aspect of agent adoption over time.
代码之所以有很多特性让它对 agent 特别友好,其中一点是你基本上可以相对快速地测试它。你既可以手动去看应用是否能正常工作,也可以直接对你构建的内容运行测试。其他大多数工作领域都没有这种优势。你往往只有在最终产品以某种形式进入现实世界时,才会得到“测试”——比如一笔股票交易被执行、一份合同被谈成、一次销售陈述被完成,等等。很可能会出现一整套新的机会,让我们开始以这种方式测试其余类型的工作。归根结底,这意味着会有更多 agent 被叠加进工作流中。这也意味着,我们需要对大多数工作流做得更好的 evals(评估)。如今企业中的大多数工作,并没有配套的 eval 来判断某个 model、prompt 或 system 的改动,究竟让事情出错了还是改进了。那些最能评估自身知识型工作的企业,也最有可能从 AI 中获得最大收益。随着时间推移,这将成为 agent 采用过程中的一个关键方面。
♥ 149↻ 19💬 40x.com ↗
Thoughtful proposal for a standards body for AI. This is distinct from a regulatory agency, and would certainly allow for much faster improvement of standards and collaboration with the industry. What you definitely don’t want is for AI progress to start to move at the speed of that the government classically operates at. If that happens then we can essentially guarantee progress begins to stall, and worse, America likely just loses the AI race. This framework threads the needle mostly. The only challenge is still that even those in industry don’t agree on the same safety risks in AI. But if you can get alignment there, this looks better than most other proposals.
这是一个关于为 AI 设立标准机构的深思熟虑的提案。它不同于监管机构,而且肯定能让标准的改进以及与行业的协作推进得快得多。你绝对不希望 AI 的进步开始以政府传统上的运作速度前进。如果那样的事情发生,我们几乎可以肯定进展会开始停滞,更糟的是,America 很可能就会输掉这场 AI 竞赛。这个框架在很大程度上拿捏住了平衡。唯一的挑战仍然是,即便是行业内部的人,对 AI 中哪些安全风险最重要也并没有共识。但如果你能在这一点上达成 alignment(对齐),那么这看起来比大多数其他提案都更好。
♥ 112↻ 16💬 14x.com ↗
原文 ↗https://x.com/levie
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers