Great post on some of the dynamics to think through for the future competitive advantage in world when AI models are shared amongst firms and packing so much for the intelligence of that industry. This is going to become a core question for companies and the economy broadly over the next decade and beyond. If AI is trained on the best datasets in every single industry - like law, finance, healthcare, or life sciences - then how do you compete and differentiate in the future? This is a great open question that I don’t think is perfectly knowable right now because of how fast AI progress is happening. But ultimately it stands to reason that if intelligence is abundant and broadly available to anyone in a field, then the companies that effectively use it the best and against a set of data and knowledge that grows in value over time, will be in a strong position. There’s a huge reinforcing loop between the intelligence from models, a company’s own data, the connection of that data and AI in their workflows, and how employees ultimately interact with that system to create value. There’s no obvious point where this will become uniform across all companies in a vertical because each company will approach this in a different way, just as they already do with their talent and workflows. If anything, there will be compounding returns to those that do this best that accelerate their advantage over time. Overall, super interesting question to see how this plays out over time.
这篇文章很好地讨论了一个值得深入思考的动态:在一个 AI model 被各家公司共享、并且承载了该行业大量 intelligence(智能能力)的世界里,未来的竞争优势会来自哪里。这将在未来十年乃至更长时间里,成为企业乃至整个经济体的一个核心问题。如果 AI 是基于每一个行业里最优质的数据集训练出来的——比如 law、finance、healthcare 或 life sciences——那么未来你要如何竞争、如何实现差异化?这是一个非常好的开放性问题,而考虑到 AI 进展的速度之快,我认为现在还不可能把它完全看清。但归根结底,可以合理推断的是:如果 intelligence 变得充裕,并且某个领域里的任何人都能广泛获得它,那么那些最有效地使用它、并且将其与一套会随时间不断增值的数据和知识结合起来的公司,将处于非常有利的位置。model 提供的 intelligence、公司自身的数据、这些数据与 AI 在工作流中的连接方式,以及员工最终如何与这个系统交互来创造价值,这几者之间存在一个巨大的强化循环。并没有一个明显的时间点会让同一垂直领域中的所有公司都变得趋同,因为每家公司都会以不同方式处理这件事,就像它们今天在人才和工作流上的做法本就不同一样。真要说的话,那些把这件事做得最好的人,反而会获得复利式回报,随着时间推移进一步加速其优势积累。总的来说,这是一个非常有意思的问题,值得观察它会如何随时间演变。
GPT-5.6 is now out. We've been evaluating the model family on the Box AI Complex Work eval, which tests the model with the Box AI Agent on a variety of extremely hard tasks using enterprise document sets. Sol is a big step up from GPT-5.5, especially on complex data-oriented tasks that require deep reasoning and analysis, and the gains concentrate exactly where enterprise work is hardest. Here are a few examples that we saw across our tests: * Financial Services (76% vs 71%): On a multi-year projection, Sol anchored to the correct opening balance sheet date rather than assuming a clean January 1 start, then carried revenue, earnings, and interest through to the right figures year over year where one early wrong assumption compounds through every downstream cell. * Healthcare (58% vs 46%): On a critical-care case review, Sol identified the correct diagnosis and intervention and avoided the dangerous misstep of ordering imaging before the time-critical procedure, a trap GPT-5.5 walked into. * Public Sector (74% vs 63%): Handed a class's raw gradebook and a new grading directive, Sol mapped each assignment to the right weight bucket, treating homework as zero-weight practice per the directive. It recomputed every student's grade to within a tenth of a percent, where GPT-5.5 drifted partway through. * Life Sciences (60% vs 51%): Across four separate compound datasets, Sol intersected the ranked target lists exactly (case-sensitive, no shortcuts) to find the biological targets common to all four, catching the shared targets GPT-5.5 missed. Sol reasons from the source definitions and checks the documents rather than taking them at face value and it's most reliable exactly where the numbers drive real decisions. This will be huge for enterprise agents using unstructured enterprise data. GPT-5.6 will be available to customers shortly within the Box AI Studio for building custom agents with.
GPT-5.6 现已发布。我们一直在用 Box AI Complex Work eval 评估这一 model family;该评测会结合 Box AI Agent,使用企业文档集对 model 在各种极其困难任务上的表现进行测试。Sol 相比 GPT-5.5 是一次明显升级,尤其是在那些需要深度推理和分析的复杂、数据导向型任务上;而这些提升恰恰集中在企业工作最困难的地方。以下是我们在测试中看到的一些例子:* Financial Services(76% vs 71%):在一个跨多年的预测任务中,Sol 锚定了正确的期初资产负债表日期,而不是假设从 1 月 1 日这个“干净”的起点开始;随后它将 revenue、earnings 和 interest 逐年正确传导到对应数值上,而在这类任务中,一个早期错误假设会在后续每一个单元格中不断累积。* Healthcare(58% vs 46%):在一个重症病例复核任务中,Sol 识别出了正确的诊断和干预方案,并避开了一个危险误区——在时间高度敏感的操作之前先安排 imaging,而 GPT-5.5 正是掉进了这个陷阱。* Public Sector(74% vs 63%):在拿到一个班级的原始成绩册和一项新的评分指令后,Sol 将每项作业映射到了正确的权重类别,并按照指令将 homework 视为权重为零的练习。它将每个学生的成绩重新计算到了百分之一百分点以内,而 GPT-5.5 在中途开始出现偏差。* Life Sciences(60% vs 51%):在四组彼此独立的化合物数据集中,Sol 精确求出了排序后目标列表的交集(区分大小写,不走捷径),从而找到了四组数据共同的 biological targets,而这些共享目标是 GPT-5.5 漏掉的。Sol 是根据源定义进行推理,并会核查文档,而不是想当然地接受表面信息;并且它恰恰在那些数字会驱动真实决策的地方最为可靠。这对使用非结构化企业数据的企业 agent 来说将是巨大的提升。GPT-5.6 很快就会在 Box AI Studio 中向客户提供,用于构建自定义 agent。