一龍馬/AI 情報站讀懂消息背後的脈絡
星期四
搜尋

AgileRL 宣布與 NVIDIA AI 合作,支援 Nemotron 模型後訓練,並稱 Nemotron 開源模型家族可在 Arena 平台上做 fine-tuning

中文摘要

AgileRL 說,NVIDIA 提供早期存取的 30B3A parameter MoE Nemotron 3.5 Lightning,在 Arena 上經訓練後,於一個長程推理挑戰與一個真實客戶支援工作負載上超過 Claude Sonnet 5。貼文還稱,一次使用四張 H100、六小時的訓練,就足以掌握該推理任務,且上下文長度超過 50,000 tokens;但完整結果需看其公告,這裡未列出評分表。

一龍馬判讀

這把 Nemotron 3.5 Lightning 推向「企業用自有資料與自有基礎設施訓練專門 agent」的路線,特別針對銀行、保險、航太與政府等重視資料主權的組織。採用者要注意,貼文同時是平台銷售訊息,所稱超越 Claude Sonnet 5 的任務範圍有限,不能外推到所有客服或推理場景。

原文節錄

NVIDIA AI · @NVIDIAAI

AgileRL (@AgileRL_Inc) We are proud to announce a collaboration between AgileRL and @NVIDIAAI to support the post-training of Nemotron models.…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

AgileRL (@AgileRL_Inc) We are proud to announce a collaboration between AgileRL and @NVIDIAAI to support the post-training of Nemotron models. NVIDIA's Nemotron open-source model family is now available for fine-tuning on Arena, our platform for creating AI agents specialized at any task. NVIDIA gave us early access to their latest 30B3A parameter MoE model, Nemotron 3.5 Lightning. On Arena, we trained it to outperform Claude Sonnet 5 on two different tasks: a long-horizon reasoning challenge and a real customer support workload. Nemotron proved exceptionally easy to post-train: a single six-hour run on four H100 GPUs was enough to master the reasoning task, at context lengths beyond 50,000 tokens. The full results are in the announcement, linked below. With this collaboration, Arena customers get access to NVIDIA Nemotron models as soon as they are released, with the training and evaluation setup already built around them. Post-train Nemotron on your own data, environment and edge cases, to build an agent that masters your task. Across industries including banking, insurance, aerospace and government, production traffic is moving off frontier model APIs and onto infrastructure these companies control. Businesses want agents with genuine expertise in their specific task, trained on their own data. Security, control and data sovereignty demand that model weights remain on their own infrastructure. And they want the fixed cost of hardware they own, rather than per-token pricing that compounds with every request. This collaboration provides that path. Begin with a dataset, an RL environment, or simply a description of the job the agent has to do, and our team will build the rest with you. Arena handles the training and deploys the finished agent in one click, with the weights yours to keep. — https://nitter.net/AgileRL_Inc/status/2087169292982714631#m

收錄日期
2026-08-13
來源
Nitter RSS(公開貼文)
抓取時間
2026/08/13 06:12(台北)