Based on internal evals: ▪️ Kimi K3 is top-tier at cybersecurity There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ. ▪️ Sol is a leap ahead in cyber capability At a significantly higher cost, but quite remarkable still. ▪️ Fable refuses everything We couldn’t get it to complete the run at all. What’s interesting is that Sol in comparison was much more open to helping with defensive cyber hardening TL;DR: frontier, open-weight cybersecurity capability is here. Try it on for defensive purposes.
根据内部 eval(评测):▪️ Kimi K3 在 cybersecurity(网络安全)方面属于第一梯队。X 上有人在议论 Moonshot 是 benchmark-overfit(对基准测试过拟合)。但这些是 stealth evals(隐藏评测)。这个模型有真材实料的 raw IQ。▪️ Sol 在网络安全能力上领先了一个量级。成本也显著更高,但依然相当惊人。▪️ Fable 则什么都拒绝。我们完全没法让它跑完整个流程。有意思的是,相比之下,Sol 更愿意帮助做防御性的网络安全加固。TL;DR:前沿级、open-weight(开放权重)的网络安全能力已经来了。出于防御目的,你可以试试看。
The term "AGI" has aged very poorly. AI can't possibly me more different from human intelligence. It's *far better* than human intelligence, for most economically-relevant tasks. But that doesn't make AIs better than humanity. There are wonderful things humans can do that these superintelligences (the better term) can't. For one, we care-for and care-about other humans in a powerful and inherent way, that we then teach machines to emulate. We must keep that as the top priority in everything we do. We have to be as pro-human-life as it gets. AIs can replace tasks you do, but they can't replace the proverbial you. This is why they horribly suck at writing! Even when given a corpus of your writing and 200 Skills and 50 subagents running in gRaPh LoOps, they produce robotic prose… because they are robots. I know one definitive way that people can become irrelevant. They stop being themselves: lose their identity, delegate all their writing, their unique thoughts, the creative ways in which they can use the machines. Opting out of 'weighing in'. I think we're safe though. There's nothing more distasteful right now than AI replies in social media or (🟢 T H I S T H I N G ) in landing pages. Quality and humanity will prevail.
“AGI” 这个词已经老化得很糟糕了。AI 不可能和 human intelligence(人类智能)更不同了。对于大多数具有经济相关性的任务来说,它*远远强于*人类智能。但这并不意味着 AIs 比 humanity(人类整体)更好。人类能做到许多美妙的事,而这些 superintelligences(我认为更准确的说法)做不到。首先,我们会以一种强大而内在的方式 care-for 和 care-about 其他人类,而机器只是后来被我们教着去模仿这一点。我们必须把这件事作为我们做任何事情时的最高优先级。我们必须尽可能坚定地站在人类生命这一边。AIs 可以替代你做的任务,但它们不能替代那个抽象意义上的“你”。这就是为什么它们写作会糟糕得离谱!哪怕给它们一整套你的写作 corpus(语料),再配上 200 个 Skills 和 50 个在 gRaPh LoOps 中运行的 subagents(子 agent),它们产出的依然是机械化的文字……因为它们本来就是机器人。我知道一种人真的会变得无关紧要的明确方式:那就是他们不再做自己——失去自己的身份,把所有写作都委托出去,把自己独特的想法,以及自己能以创造性方式使用机器的能力,都一并交出去。选择不再“发表看法”。不过我觉得我们还是安全的。眼下没有什么比社交媒体上的 AI 回复,或者落地页上的(🟢 T H I S T H I N G)更令人反感了。质量和人性终将胜出。