unironically this is happening right tf now
说真的,这事现在他妈的就正在发生。
very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of "frontier" model training is that even without training on test, you can basically cheat by training on test lookalikes, enabling you to goalseek almost any benchmark number you want. however when they are released open weights, 99% of the time the norm is that you do not get the datasets/rlenvs that would easily show you if someone was training on Temu Tbench, so there is plausible deniability. Alex and Omar discuss applying standard NLP distance metrics on hidden trajectories. There's no ultimate solution here, but they have some prelim explorations. It happens to support the finding that RLMs can generalize to unseen tasks that share latent structure observed in training.
这里有一篇很值得注意的轨迹(trajectory)比较分析,藏在 @a1zhang 和 @lateinteraction 的 RLM paper 里。关于“frontier” model 训练,有一个公开的秘密:即使不直接在测试集上训练,你基本上也可以通过在测试集的相似版本上训练来作弊,从而把几乎任何 benchmark 分数都朝你想要的目标去调。不过,当这些模型以 open weights 形式发布时,99% 的情况下,常态是你拿不到那些数据集 / RL envs,而正是它们本可以很容易地告诉你,某人是不是在 Temu Tbench 这类东西上训练过,所以发布方总有可辩解空间。Alex 和 Omar 讨论了如何把标准 NLP 距离度量应用到隐藏轨迹上。这里没有终极解决方案,但他们做了一些初步探索。碰巧的是,这也支持了这样一个发现:RLM 可以泛化到那些未见过、但与训练中观察到的潜在结构共享 latent structure 的任务。