一龍馬/AI 情報站讀懂消息背後的脈絡
星期五
搜尋

François Chollet 表示,ARC-AGI-3 在 3 月推出時,前沿模型得分低於 1%,但六個月內已提升至 100%,他據此主張這套基準成功捕捉了代理式能力的快速進展

中文摘要

他也反駁基準無法由人類完成的批評,稱只要使用比未篩選人類受測者基準更少的操作次數,人類便能拿到滿分。貼文未提供各模型、測試設定或第三方驗證,相關結論仍是基準作者的說法。

一龍馬判讀

若測量方式與成績可重現,這代表代理系統在短期內跨越了極大的能力差距;但基準設計者同時也是解讀者,仍須檢查是否有測試污染、過度最佳化或評分設定改動。

原文節錄

François Chollet · @fchollet

Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult,…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc. We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark. As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers). And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us.

收錄日期
2026-09-04
來源
Nitter RSS(公開貼文)
抓取時間
2026/09/04 06:14(台北)