一龍馬/AI 情報站讀懂消息背後的脈絡
星期六
搜尋

DeepLearning.AI 介紹名為 AREX 的代理式研究模型與框架,主張逐項驗證需求、保留已驗證內容,只補查缺口

中文摘要

貼文列出 BrowseComp 82.5%、WideSearch-en F1 82.0,以及框架帶來最高 10 分提升等數字。基準設定、微調 4B 模型的比較基礎都只見於貼文,缺乏論文佐證。

一龍馬判讀

做深度研究助理與評測的團隊可追蹤其驗證式檢索作法,但引用分數前應查閱原始評測條件。

原文節錄

DeepLearning.AI · @DeepLearningAI

checks its answers requirement by requirement…

取得部分原文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

A new agentic research model and harness called AREX checks its answers requirement by requirement, keeps what's verified, and researches only the gaps. 📰✅ 📊 82.5% BrowseComp, 82.0 F1 WideSearch-en 📈 Harness alone: up to +10 pts 🧠 Fine tuned 4B beat untuned 35B on 5 of 6 benchmarks https://hubs.la/Q04ztJKG0 #DeepLearningAI #AIAgents #LLMs

收錄日期
2026-10-10
來源
Nitter RSS(公開貼文)
抓取時間
2026/10/10 05:37(台北)