一龍馬/AI 情報站讀懂消息背後的脈絡
星期五
搜尋

François Chollet 主張,AI 評測必須與模型共同演進:新基準先以新問題引導研究方向並提供回饋,之後再隨模型進步調整,持續鎖定 AI 與人類智慧之間尚未補上的差距

中文摘要

他表示團隊在今年發布 ARC-AGI-3 後已開始開發 ARC-AGI-4,預定於 2027 年第一季推出。貼文尚未揭露 ARC-AGI-4 的題型、評分方法或防止資料污染的設計。

一龍馬判讀

明確時程讓模型開發者與評測研究者可提前規畫下一輪驗證,但若細節公布過早,也可能增加針對基準最佳化的風險;最終效力仍取決於它能否測到尚未被模型掌握的人類能力。

原文節錄

François Chollet · @fchollet

R to @fchollet: Benchmarking AI systems is a continual process that co-evolves with the models.…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

R to @fchollet: Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between AI and human intelligence. We are still working on ARC-AGI-4, which we started developing after releasing ARC-AGI-3 earlier this year. It is coming Q1 2027. We think it's going to be really special.

收錄日期
2026-09-04
來源
Nitter RSS(公開貼文)
抓取時間
2026/09/04 06:14(台北)