unironically this is happening right tf now
说真的,这事现在他妈的就正在发生。
very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of "frontier" model training is that even without training on test, you can basically cheat by training on test lookalikes, enabling you to goalseek almost any benchmark number you want. however when they are released open weights, 99% of the time the norm is that you do not get the datasets/rlenvs that would easily show you if someone was training on Temu Tbench, so there is plausible deniability.
这里有一篇很值得注意的轨迹(trajectory)比较分析,藏在 @a1zhang 和 @lateinteraction 的 RLM paper 里。关于“frontier” model 训练,有一个公开的秘密:即使不直…