Everyone should develop their "personal eval set" for AI models: a few tasks that are actually relevant to your day-to-day work/life The industry benchmarks help but they might not reflect what will make it actually useful to you You find the model's capability boundary by poking at it & bumping into it for fun
每个人都应该为 AI models 建立自己的“personal eval set”:也就是几项真正与你日常工作/生活相关的任务。行业基准测试有帮助,但它们未必能反映出什么才会让它对你真正有用。你要通过不断试它、撞它的边界、带着一点玩乐心态去摸索,才能找到这个 model 的能力边界。
The biggest barrier for enterprise AI adoption: the people who understand AI don't understand the business, and the people who understand the business don't understand AI
企业采用 AI 的最大障碍在于:懂 AI 的人不懂业务,而懂业务的人不懂 AI。