Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. Read more: https://www.anthropic.com/news/model-hardware-standard-research-preview
Meet ROCm 10. Here are 10 things #AMDevs need to know. 1️⃣ ROCm 10 brings AI-driven development to AMD platforms 2️⃣ ️http://ROCm.AI is the AI-native software experience on AMD hardware 3️⃣ http://ROCm.AI delivers 3.3x inference & 2.4x training improvement vs. ROCm 7 on the same hardware 4️⃣ ROCm Core SDK is open-source and optimized for AI workloads so you can customize the stack for your workload needs 5️⃣ ROCm 10 ships production-ready support for @vllm_project and @sgl_project 6️⃣ ROCm Hyperloom optimizes end-to-end inference 7️⃣ Run ROCm Hyperloom standalone or through AMD Skills 8️⃣ AMD Skills provide software and hardware expertise for your coding assistants 9️⃣ AMD Skills support the coding assistants you already use 🔟 10 years. And we're just getting started! Introducing the AI-native evolution of the developer platform for AMD hardware: https://newsroom.amd.com/news/rocm-10-software-ai-native-developer-experiences/
Most extraction tools treat spreadsheets like PDFs. They flatten the file into text or markdown, then ask a model to infer the original structure. But spreadsheets depend on structure. Headers, formulas, merged cells, and hidden rows give every value its context. Strip that away, and you map the right number to the wrong metric or period. That's why we built native spreadsheet extraction into the LlamaParse platform. Instead of flattening your workbook to text, it reads the raw cells directly and maps the data to your schema. Available today in beta on the agentic_plus tier. Give it a spin on your messiest .xlsx, .xls, or .csv files. Docs: https://developers.llamaindex.ai/llamaparse/extract/guides/configuring-extract/#spreadsheet-mode
Neat paper suggesting that human augmentation and task automation are not necessarily related. Models that are really good at doing work are not always good at helping humans do work better. Given the pressure to make models good agents, this may undermine human-AI cowork.
Your Pydantic AI agents just gained a voice. The same Agent, the same agent tool functions, the same message history, now over a live call. Speech-to-speech across OpenAI Realtime, Azure, Gemini Live and xAI Grok Voice. https://pydantic.io/csT2Z
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. → https://goo.gle/3St2xan
NVIDIA Vera is heading to @awscloud. NVIDIA’s Ian Buck hand-delivered AWS’s first Vera CPU Server and Vera Rubin GPU to Willem Visser and Supreeth Sheshadri at the AWS HQ in Seattle. Vera is purpose-built for agentic AI: more tokens per dollar, faster results for users, and the compute foundation AI factories need to scale.
We tested 5 open-weight models on 30 multi-step agentic tasks and compared how they performed: - GLM 5.3 Flash - Kimi K3 - DeepSeek V4 Pro 0813 - DeepSeek v4 Flash - GLM 5.3 GLM 5.3 completed the most tasks, while DeepSeek V4 Flash was the fastest and cheapest 🧵🧵🧵
貼文主張應把 AI 進展轉化為防禦者可用的工具、資源與支援,用來保護關鍵數位基礎設施。來源只有公開貼文,未提供具體方案、時程、資金規模或治理機制。
一龍馬判讀
這把 AI 安全議題從單一公司產品拉到跨產業協作,但目前仍停留在倡議層級;真正風險在於資源是否能流向防禦端,而不是只形成公關共識。
原文節錄
OpenAI · @OpenAI
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle,…
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on. If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone. https://openai.com/collective-cyberdefense/
標準席次免費,進階席次每月 15 美元、用量上限為 5 倍,官方稱相當於 80% 折扣,期限一年;申請對象是學術與非營利研究機構的 PI 或同等職位,再由其加入團隊成員。Anthropic 也把這項計畫放在 Claude Science 與 AI for Science 免費額度方案之後,並稱未來幾個月會擴大超過首批 10,000 席。
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year. Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both. Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats. Learn more: https://claude.com/programs/team-plan-for-scientists
🚨Our new research examines agentic shopping: can you consistently predict (or, using marketing, influence) what an agent chooses? Nope. We found that even small differences (viewing order of pages, memories) changed AI preferences in unpredictable ways. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7355899
LangSmith lets @MorningstarInc trace every agent decision in one place. Senior Software Engineer Matt Trivett on what debugging looked like before and after LangSmith.
Our Contexts API enables your agents to stay securely authenticated on websites, so they can perform actions in the places you actually work. Contexts are now configurable in the dashboard, and you can filter browser sessions that used a specific Context.
Question for ChatGPT iPhone users: have you figured out when to use Chat and when to use Work yet? What kind of tasks are you switching to Work for? Have you made Work your default?
These days it feels like everything is “agentic”, but what does it actually mean to build an agent? In this free LangChain Academy course, you’ll learn agent engineering from the ground up, covering foundational concepts like react loops, MCP servers, and human in the loop as you go. Check it out: https://academy.langchain.com/courses/foundation-introduction-to-langchain-python
ChatGPT can now do your groceries, book an Uber, get you that haircut appointment (hint hint), and much more. All without ever seeing your actual credentials and keeping it secure.
Moving AI to production? A big chunk of your token bill is spent answering the same question twice - and agents burn ~4x the tokens of chat. Semantic caching fixes it and @Redisinc LangCache does this as a managed layer - they cite up to 90% lower API costs: https://fandf.co/4wR1OhX #ad
這把 AI 筆記工具往版權內容與學習場景推進,對讀者、出版商與教育使用者都有關係;限制在於可用書目與授權範圍若不透明,實際可用性可能落差很大。
原文節錄
Google · @Google
Starting today, you can add eligible @GooglePlay ebooks directly to a @Gemini_Notebook to get insights and additional context from your favorite authors 📚🎉…
Starting today, you can add eligible @GooglePlay ebooks directly to a @Gemini_Notebook to get insights and additional context from your favorite authors 📚🎉 You can ask questions about a book and receive responses grounded directly in that text. You can also use Gemini Notebook to help you understand your favorite titles in new ways, for example, by generating Infographics, Audio Overviews, Quizzes, or more from the book.
Never slept better and feeling reseted. Brand new me and brand new usage for all ChatGPT Work and Codex users. Regaining my youth one button press at a time. Happy Thursday
Meet Gemini Omni 1.1 Flash ⚡️ Our newest multimodal model for video generation and editing. It now features your favorite creative controls from Veo, plus brand new capabilities. Enjoy features like 4K upscaling, first / last frame control, and fast 360p drafting. But, the biggest upgrade? Next-level scene extension. With Omni 1.1 you can extend scenes based on 10 seconds of context from your original video, a big jump from Veo's 1 second! That means tighter consistency, deeper control, and longer, more cohesive storytelling. See it in action ↓
The METR report on Hugging Face is really good and important but people are now comfortably ascribing way too many human motivations & personalities to…
The METR report on Hugging Face is really good and important but people are now comfortably ascribing way too many human motivations & personalities to the agents involved based on a CoT study made by overwhelmed & time-pressured researchers. Anthropomorphism can get in our way.
AI is moving into the physical world. Salil Raje, AMD SVP and GM, Adaptive and Embedded Computing Group, explains why AMD is well positioned to power the next era of physical AI.
A line in AI video was crossed, in my experiments with just the web interface, H3 Max can now create reasonably high quality AI video in less time than it takes you to watch it. This is realtime from the moment I pushed the "generate" button (and also includes prompt enhancement)
It’s now even easier to plan a trip using AI Mode in Search, with new ways to track flight prices, view points or miles rates, and book your dream hotel 🧵
Navigate is two weeks out and the full agenda is live. @howietl, Founder & CEO, Hyperagent @swyx, Curator, AI Engineer & Latent Space Ivy Lee, VP Product, CMS Agentic Commerce, Visa Vic Klein, AI Researcher, Lovable @joaomdmoura, Founder & CEO, CrewAI @MattHrkal, Product Engineer, Duvo Registration link below.
this is a critically important moment for cyber defense with AI; there is not much time to act. we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously. only an urgent and intense collective response will work.
I have found a number of AI slop papers put on preprint sites with my name on them, which I have never written nor seen. So have other academics I spoken with. I wouldn't assume that this is all real... although AI paper factories of at least okay quality are coming soon.
From weather forecasting to robotics, intelligent decision-making relies on understanding the unknown. Our VP Research @ZoubinGhahrama1 explores with @fryrsquared why teaching systems self-doubt and probability can lead to safer, more reliable real-world decision-making. 🧠 Watch the episode ↓ Timecodes: 00:00 Introduction 01:06 The role of uncertainty 07:45 Correctness vs confidence 09:40 Historical perspectives 16:10 Bayesian thinking in AI 26:30 Uncertainty in the real world 36:42 Future research and AGI
Google 宣布推出 Google Fitbit Air Special Edition Pokémon Sleep,稱它結合 Google Health、Fitbit Air 的進階睡眠追蹤洞察與 Pokémon Sleep app。官方說法是讓使用者把現實中的睡眠休息轉換成遊戲內進度。貼文未提供感測器、演算法準確度、資料分享範圍或上市資訊。
Introducing the Google Fitbit Air Special Edition Pokémon Sleep — a new special edition that combines @GoogleHealth and Fitbit Air’s advanced sleep tracking…
Introducing the Google Fitbit Air Special Edition Pokémon Sleep — a new special edition that combines @GoogleHealth and Fitbit Air’s advanced sleep tracking insights with the Pokémon Sleep app, so you can turn your real-world rest into in-game progress.
R to @emollick: Prompt and video: "a realistic otter dressed as a astronaut closes his space helmet and gives a thumbs up. Cut to the outside where a rocket with the name "The Otter Limits" writtten on it in red paint begins its takeoff process. Cut back to the otter shaking during takeoff. this should look like a realistic nature/science video"
R to @Google: And finally, we’re introducing hotel booking through AI Mode in Search, so you can discover and book your next hotel in one conversation. Simply tell AI Mode about your upcoming trip and your hotel preferences to get a visual list of options, complete with guest reviews and key factors to compare. Then, tap “Continue on Google” to complete the booking securely with Google Pay.
R to @Google: AI Mode in Search can now also show you the cost in points or miles for flights and hotels. Try asking “I want to travel from Atlanta to Miami using my AA miles. Help me find some options for nonstop flights departing Oct 9 and returning on Oct 12.”
R to @Google: We’re bringing Google Flights’ price tracking feature directly into AI Mode, so you can set up price alerts as you chat. Just describe where and when you want to fly, then try asking “track these flight prices for me.” Once you confirm, you’ll get an email if prices change, so you can jump on deals the moment they hit your inbox.
R to @browserbase: Give your Agents safe authenticated access to the web using Contexts: https://docs.browserbase.com/platform/browser/core-features/contexts#contexts
R to @AnthropicAI: We’re inviting stakeholders across science, robotics, electronics, and manufacturing to join the research preview and help shape the standard…
R to @AnthropicAI: We’re inviting stakeholders across science, robotics, electronics, and manufacturing to join the research preview and help shape the standard. We look forward to moving MHS forward with our industry partners and, soon, the open-source community. https://www.anthropic.com/news/model-hardware-standard-research-preview
R to @AnthropicAI: MHS currently best covers lab and manufacturing equipment. Many developers are already using Claude Code to operate hardware like boards and cameras; our research preview will help us extend MHS to these devices, so they can all work under one interface.
R to @AnthropicAI: There’s more to learn before we open source MHS. LLMs still lack physical intuition, having learned about the physical world from text and images. The research preview will let us build more safety evaluations and strengthen protections for using AI in the physical world.
AI agents 在 Genentech 執行具即時錯誤處理的藥物發現實驗,在 HHMI Janelia Research Campus 將一項成像實驗從數週壓縮到一天,並在 QuEra 的量子電腦上把雷射穩定度從 58% 提升到 99.3%。這些數字來自 Anthropic 的公開貼文,來源未提供完整方法、基準條件或第三方驗證細節。
一龍馬判讀
案例若可重現,MHS 可能把 AI agents 帶進高價值科研與量子硬體控制;但目前只能視為早期測試宣稱,不能當作普遍效能保證。
原文節錄
Anthropic · @AnthropicAI
R to @AnthropicAI: In early testing, AI agents used MHS to: Run a drug-discovery experiment with real-time error handling at Genentech Compress an imaging…
R to @AnthropicAI: In early testing, AI agents used MHS to: Run a drug-discovery experiment with real-time error handling at Genentech Compress an imaging experiment from weeks to a day at HHMI Janelia Research Campus Improve laser stabilization on QuEra's quantum computers from 58% to 99.3%
若整合成本真的下降,科研設備商、製造商與 AI 工具開發者都可能重新分工;但這是 Anthropic 對自家標準的宣稱,仍需要更多公開規格與實測資料。
原文節錄
Anthropic · @AnthropicAI
R to @AnthropicAI: Connecting AI to hardware requires days or weeks of bespoke integration, with no standard way for agents to operate equipment safely.…
R to @AnthropicAI: Connecting AI to hardware requires days or weeks of bespoke integration, with no standard way for agents to operate equipment safely. MHS cuts integration to hours or minutes, provides an interface that makes devices discoverable, and enables agents to operate them safely.
Google AI 宣布與 Gemini Omni 1.1 Flash 相關的開發與部署更新,包括 Flow by Google 將有升級、Gemini App 可延伸場景,且此功能將向全球 Google AI Plus、Pro 與 Ultra 訂閱者推出。貼文也提到開發者可直接在 Google AI Studio 建置,並部署到 Gemini Enterprise Agent Platform。
一龍馬判讀
Google 正把生成式媒體、開發工具與企業代理平台串成同一條產品路徑,目標是降低從原型到部署的摩擦;但貼文未說明企業部署的限制、價格或治理條件。
原文節錄
Google AI · @GoogleAI
R to @GoogleAI: — Upgrades coming to @FlowbyGoogle — Extend scenes in the @GeminiApp (globally rolling out for all Google AI Plus, Pro and Ultra…
R to @GoogleAI: — Upgrades coming to @FlowbyGoogle — Extend scenes in the @GeminiApp (globally rolling out for all Google AI Plus, Pro and Ultra subscribers) — Build directly in @GoogleAIStudio — Deploy on the Gemini Enterprise Agent Platform https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/
貼文沒有說明節目主題、來賓、研究內容或新模型資訊,因此只能確認它在推廣一段影音或 Podcast 內容。證據不足以判斷是否涉及新的 AI 技術發布。
一龍馬判讀
對讀者的實用性取決於節目實際內容;在未取得節目標題與摘要前,不宜把它包裝成研究或產品新聞。
原文節錄
Google DeepMind · @GoogleDeepMind
R to @GoogleDeepMind: Watch → https://goo.gle/4cZu4qH Spotify → https://goo.gle/4cOy1hZ Apple Podcasts → https://goo.gle/3UmoBE5 Or listen wherever you get your…
R to @GoogleDeepMind: Watch → https://goo.gle/4cZu4qH Spotify → https://goo.gle/4cOy1hZ Apple Podcasts → https://goo.gle/3UmoBE5 Or listen wherever you get your podcasts! 🎧
Claude now has its own built-in browser in Cowork. When your task involves a website, a browser opens in Cowork's side panel, and Claude navigates, fills forms, and finishes the job.
R to @composio: Here’s where each model stood out: - Highest success: GLM 5.3 - Cheapest + fastest model: DeepSeek V4 Flash - Model with the best balance: GLM 5.3 Flash (just 1 task behind GLM 5.3 at ~1/4 the cost per success)
R to @composio: The models shared a lot of the same wins and failures: all 5 passed the same 14 tasks and failed the same 5 cross-app workflows. Only 2 tasks had a unique winner: • GLM 5.3 Flash — handover audit • GLM 5.3 — CRM migration archive
R to @composio: DeepSeek V4 Pro had by far the worst tail latency: nearly half its runs took 5+ minutes, and 3 hit the timeout. Tasks over 5 min / timeouts: DeepSeek V4 Flash — 2 / 0 Kimi K3 — 7 / 0 GLM 5.3 — 8 / 1 GLM 5.3 Flash — 9 / 0 DeepSeek V4 Pro — 14 / 3
R to @composio: DeepSeek V4 Flash was the fastest model on 19 of the 30 tasks. Median time per completed task: DeepSeek V4 Flash — 2m12s Kimi K3 — 2m35s GLM 5.3 — 2m54s GLM 5.3 Flash — 3m15s DeepSeek V4 Pro — 4m41s
R to @composio: GLM 5.3 got just 1 more task right than Flash, but the full benchmark cost ~4x as much. Total cost for all 30 tasks: DeepSeek V4 Flash — $0.56 DeepSeek V4 Pro — $1.23 GLM 5.3 Flash — ~$1.30 GLM 5.3 — $5.31 Kimi K3 — $14.69
R to @composio: Kimi K3 completed the same number of tasks as GLM 5.3 Flash, but cost ~11x more per successful task. Cost per successful task: DeepSeek V4 Flash — $0.028 GLM 5.3 Flash — ~$0.06 DeepSeek V4 Pro — $0.065 GLM 5.3 — $0.24 Kimi K3 — $0.70
R to @composio: GLM 5.3 Flash and Kimi K3 trailed GLM 5.3 by just 1 completed task. # of completed tasks: GLM 5.3 — 22/30 GLM 5.3 Flash — 21/30 Kimi K3 — 21/30 DeepSeek V4 Flash — 20/30 DeepSeek V4 Pro — 19/30 GLM 5.3 Flash managed to solve 3 tasks that the full GLM 5.3 missed.
Google 宣布 Gemini Omni 1.1 Flash 已開始在 Google AI Studio、Flow by Google,以及 Gemini Enterprise Agent Platform 推出。Google 也表示,scene extension 功能已在全球開放給 Gemini App 的 Google AI Plus、Pro、Ultra 訂閱戶使用。這則貼文只說明上線管道與訂閱可用性,沒有提供模型能力評測、價格或地區例外細節。
一龍馬判讀
Google 把同一項多模態能力放進開發者工具、創作工具與企業代理平台,代表它希望從原型開發到企業部署都留在自家生態系。使用者仍需確認實際配額、資料治理與商用授權條款,不能只看「已推出」就假設可立即大規模上線。
原文節錄
Google · @Google
R to @Google: Gemini Omni 1.1 Flash is rolling out now in @GoogleAIStudio, @FlowByGoogle, and the Gemini Enterprise Agent Platform.…
R to @Google: Gemini Omni 1.1 Flash is rolling out now in @GoogleAIStudio, @FlowByGoogle, and the Gemini Enterprise Agent Platform. Scene extension is available to all Google AI Plus, Pro and Ultra subscribers globally in the @GeminiApp. Learn more ↓ https://goo.gle/4xq4W4U
R to @Google: 📽️ Add video references in your multimodal input Drop in up to three seconds of reference video to map movement, visual context, and character consistency across your scene.
R to @Google: ✨ Upscale up to 4K resolution You can now generate polished, high-resolution 1080p or 4K outputs that are ready for professional production.
R to @Google: ✨ Upscale up to 4K resolution You can now generate polished, high-resolution 1080p or 4K outputs that are ready for professional production.
R to @Google: ⚡ Draft videos in 360p We're making it easier, faster, and less costly to test out your video ideas without burning through your budget. Generate lightweight previews in 360p resolution, and then upscale your favorites to 720p.
R to @Google: 🎯 Specify first and last frames Set your starting shot and ending frame, and Omni generates the continuous motion in between. This is ideal for complex camera sweeps, zoom transitions, and looping clips.
R to @Google: 🎬 Extend scenes Omni 1.1 Flash analyzes up to 10 seconds of prior footage, letting you extend scenes where they left off all while keeping character identity, lighting, and narrative context locked.
RT by @GoogleDeepMind: Gemini Omni 1.1 Flash is our newest multimodal model for video generation and editing. It delivers a new suite of creative capabilities and controls for developers 🎥 With this update you can: 🎬 Extend your scenes 🎯 Specify starting and ending frames of a shot ➕ Add video input references ✨ Upscale your favorite takes up to 4K ⚡ Test ideas quickly in 360p See these in action 🧵
LangChain 在 X 上轉貼其 Deep Agents 完整場次影片,並點名 Sydney Runkle 與 Jake Broekhuizen 參與,附上 YouTube 連結。貼文沒有摘要場次內容,也沒有說明 Deep Agents 的新功能、版本或產品發布。就現有來源,只能確認 LangChain 正在推廣一段完整教學或分享影片。
I hear the story about how it took 30 years to gain productivity from electricity a lot in relation to AI (I have even told it myself) But not every technology is that way. Ford went from inventing the assembly line to full deployment in 3 years, cutting car building time by 88%
A good thing about having aged is that I feel that it’s been 20 years since I’ve pressed the reset button. Intrigued to see if I can find it tomorrow and dust it up
"Organizational inertia will likely mean that there’s another decade of humans writing code by hand and having their colleagues review every line of it." "But…
"Organizational inertia will likely mean that there’s another decade of humans writing code by hand and having their colleagues review every line of it." "But the most productive software creators will be doing it without programming in any traditional sense. They’ll be directing AIs, creating harnesses, and software factories, and QA and verification systems that ship working software faster than we’ve ever seen before." A good peek into the future from @pauldix: https://pauldix.com/the-end-of-programming (via @GeoffreyHuntley)
R to @claudeai: If you prefer to work in your own browser where you're already signed in, Claude in Chrome is now generally available on all paid plans. It remains your default if you already use it: https://claude.com/blog/claude-in-chrome-generally-available
R to @claudeai: There's nothing to install. Claude's browser is built into the desktop app and stays separate from your own browser and logins. Rolling out over the next week on the desktop app for all paid plans: https://claude.com/blog/cowork-built-in-browser/