Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc. We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark. As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers). And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us.
An example of useful knowledge work: I assigned GPT-6 to read through tens of thousands of my emails, my writings, my calendar appointments and more to assemble a personal knowledge base of research, contacts, ideas, relationships, and tasks over my recent career. GPT-6 downloaded the appropriate software, figured out a strategy, and built a multi-gigabyte personal wiki without any further intervention from me over the course of five days of uninterrupted work. Twice a day, the AI now goes through my emails, cross-references them with this extensive knowledge base, and sends me a summary of things I should be paying attention to, things I might be interested in, etc. (Two questions you may have: Yes, I gave the AI access to my computer and this involves risks, and you should be careful before you do the same. And I do not take money from any AI lab and pay for my own usage, but during the trial period, I was not charged for tokens, so I cannot tell you how much this process would have cost, but it likely would have been very substantial. The ongoing briefings do not use substantial amounts of token)
GPT-6 Astra is here. We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building. We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more. It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait. It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
We are starting to release GPT-6 Astra and we are doing it as carefully and quickly as possible. It was very important to us that we bring it to all Plus users and not only Pro, Business and Enterprise. It will take a few days for the rollout to complete and behind the scenes many novel systems will operate at scale for the first time and we are bringing a lot of compute up. It is pure magic. https://openai.com/index/gpt-6-astra/
Google 表示,多數 AI 天氣模型依賴具有六小時延遲的數值天氣預報資料,而直接使用真實觀測可繞過這項限制;貼文沒有提供誤差數據、區域差異或極端天氣表現。模型將直接支援 Google 搜尋、Gemini、Google 地圖、Maps Platform Weather API 與 Earth Engine。
一龍馬判讀
把衛星驅動的預報直接放入大眾產品與開發者 API,可改變民眾、能源業與應用服務取得即時天氣資訊的方式;不過「最準確」仍是 Google 自述,災防用途需要區域化驗證與不確定性資訊。
原文節錄
Google · @Google
Today, @GoogleDeepMind and @GoogleResearch are introducing WeatherNext 3, our most advanced and accurate global weather AI model to date.…
Today, @GoogleDeepMind and @GoogleResearch are introducing WeatherNext 3, our most advanced and accurate global weather AI model to date. It uses real-time satellite data to generate hourly high-resolution forecasts, precise precipitation forecasting, and clean energy variables. Most AI weather models are trained on data from numerical weather prediction (NWP) models, which come with a six-hour data lag. But by training on real-world observations, WeatherNext 3 is able to bypass these traditional constraints — meaning more people can get more accurate predictions to help them plan ahead. Starting today, WeatherNext 3 will power forecasts within Search, @GeminiApp, @GoogleMaps, Google Maps Platform Weather API and Google Earth Engine.
🗑️ AI agents need better garbage collection. Xiaohongshu researchers built Self-GC, using a planner LLM to decide which context tokens to keep, fold, or prune. In tests, it retained necessary details 84.85 percent of the time compared to just 54.55 percent for standard methods. Master agent memory management: https://hubs.la/Q04wy8ML0 #DeepLearningAI #AIAgents #LLMs
.@HeggieConnor's rule at @unifygtm: if the agent runs on GPT, the judge grading it needs to run on a different model family. Choose the same family and you get mode collapse, or what is essentially groupthink for agents.
Open models are essential to expanding access to AI and accelerating innovation around the world. We're excited to help @huggingface scale its platform and community while preserving the openness, neutrality, and choice that have made it a trusted home for AI builders. 🤗💚
What are the most popular SEO apps among AI agents? Composio connects agents to more than 1,500 apps and powers millions of actions every month, giving us a glimpse into their preferred tools. Here’s the top five most popular SEO apps used by AI agents in the last 30 days ↓
I believe we need to make a deliberate effort to keep humans in the loop in all critical processes across our economy and society, regardless of whether it is technically necessary. Even if AI develops the *capability* for advanced autonomy, we should not make it highly autonomous. We have to maintain control and keep visibility and understanding of all critical processes, we should not blindly hand over everything to AI agents just because we can. AI as a tool in the human hand is the only form of AI that is worth pursuing.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: https://arcprize.org/blog/astra
sorry for the radio silence folks - got sucked into extreme LLM psychosis. but i can confidently say we have crossed over into a new age of AI Engineering and we are never, ever, looking back. this isnt even EVERYTHING i did with Astra but i'll append more reports as I publish them on LS!
Document extraction just got a new gear. ⚡️ Introducing Turbo mode for Extract, our fastest way yet to pull structured data from documents. It runs roughly 4× faster than our Cost Effective Tier at comparable accuracy, with a median latency of 3.7 seconds a page. Turbo processes pages in parallel, so latency stays nearly flat as document size grows. Available in beta today. Read the full breakdown → https://www.llamaindex.ai/blog/introducing-turbo-our-fastest-extraction-tier
Looking forward to this ecosystem with open models continuing to flourish and grow with NVIDIA and Hugging Face, and the continued partnership. Congrats @JensenHuang and @ClementDelangue!
Google AI 發表由 Google DeepMind 與 Google Research 開發的 WeatherNext 3,宣稱其全球天氣預測能力最高可比 WeatherNext 2 精細 5 倍,能以高空間解析度追蹤快速變化的暴雨、區域溫度及風力發電條件。模型直接使用即時地球同步衛星資料,並以真實地表與大氣觀測訓練,可每小時更新全球預報;Google 對比指出,傳統物理式超級電腦模擬可能有 6 小時預報延遲。貼文未提供「精細 5 倍」的具體指標、區域測試結果或對極端天氣的誤差範圍。
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to…
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model generates a forecast with high spatial resolution in order to catch fast-evolving rainstorms, map local temperature shifts, and even help wind farms predict their power output. So, how does it do that? While traditional weather models rely on massive, physics-based supercomputer simulations that can carry a 6-hour forecast lag, WeatherNext 3 leverages live geostationary satellite observations as inputs and trains directly on real-world surface and atmospheric observations. By pulling this raw satellite data, it’s able to update the global forecast every single hour. And because weather develops at lightning speed, these quick, detailed insights can help bring more localized forecasting to billions of people and local businesses, especially in regions that are historically underserved due to the high costs of traditional weather forecasting models.
New cookbook with @Nevermined_AI: ✅ Let your agents pay for what they need mid-task, no human required ✅ Delegate your card, and set limits ✅ Let your agent buy credits, top itself up, and pay for the tools it uses ✅ Every purchase, traceable in LangSmith Learn more: https://www.langchain.com/blog/agents-that-pay-how-nevermined-empowers-langchain-agents-to-buy-and-sell-services
Today, we're making some exciting updates to MCP in LangChain, including support for the new stateless MCP spec! @SydneyRunkle with everything you need to know ⤵️
當 AI 品牌以神祕、末日式語彙營造聲勢,可能放大炒作並模糊可驗證的產品差異;讀者不應把這則評論當成發布消息。
原文節錄
Ethan Mollick · @emollick
Between one AI company naming their model mythos and another just posting "the stars are almost aligned," I feel like the marketing departments need to…
Between one AI company naming their model mythos and another just posting "the stars are almost aligned," I feel like the marketing departments need to become more (or less?) familiar with classics of 1920s-era horror.
🛠️ DeepSeek V4 Pro 0813 is out, but just as big of a story may be the company’s open source evaluation harness. DeepSeek Harness logs every tool call, system prompt, and subagent schedule. Developers can now reproduce performance instead of relying on closed testing environments. And they can easily study, fork, or remake their own harnesses to boot. Read more in The Batch: https://hubs.la/Q04wjYXx0 #DeepLearningAI #OpenSource #Developers
The best orgs have figured out how to ship agents repeatedly, safely, and systematically. They’ve established a continuous agent development lifecycle: 1️⃣ Build 2️⃣ Test 3️⃣ Deploy 4️⃣ Monitor ...So they can learn from real usage, and iterate quickly With our free LangChain Academy course, LangSmith Essentials, you’ll explore the entire agent development lifecycle in less than 60 minutes. Check it out: https://academy.langchain.com/courses/quickstart-langsmith-essentials
I had early access, and a longer post is coming, but GPT-6 is stunning & is good enough that it actually does complex meaningful work for me autonomously for days. As a more fun example, it made a historically-based simulation of the Library of Alexandria: https://alexandria-mouseion.netlify.app
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized highly accurate predictions faster. 🧵
Anthropic published a new guide on how to strip the "Claude" out of Claude's writing (e.g. mannered prose) Here's an official de-flavoring prompt for Fable 5.1, which already cuts back on AI boilerplate/jargon: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1
對長時間使用 Claude Code 的工程師而言,脈絡管理可能影響成本與輸出品質;但實際策略仍須依程式庫規模、任務複雜度及團隊流程測試。
原文節錄
Addy Osmani · @addyosmani
How to maximize the value of your Claude Code sessions: https://claude.com/blog/maximizing-the-value-of-your-claude-code-sessions Solid guidance on managing…
How to maximize the value of your Claude Code sessions: https://claude.com/blog/maximizing-the-value-of-your-claude-code-sessions Solid guidance on managing token efficiency and keeping context windows focused by @lydiahallie
他批評目前做法多半只是把 AI 當成群組聊天室裡的一名成員,這種互動模式限制了跨人員、流程與目標的協作能力。這是他的觀察性判斷,貼文未提供調查數據或實際案例。
一龍馬判讀
企業若只增加聊天機器人,未必能處理權限、交接、責任歸屬與共享脈絡等組織問題;產品團隊需要把 AI 設計成協作基礎設施,而不只是另一個對話帳號。
原文節錄
Ethan Mollick · @emollick
Multiplayer AI, where many people in an organization can use AI together to accomplish goals, remains one of the biggest (non-technical) problems in using AI…
Multiplayer AI, where many people in an organization can use AI together to accomplish goals, remains one of the biggest (non-technical) problems in using AI right now. Approaches tend to be pretty primitive and based around AI-as-a-person-in-your-group-chat. That is limiting.
Fable 5.1: "Create the Catalog of Ships from the Iliad with a map, etc. i should be able to explore each accurate ship in 3D. there should be real images & archeological data i can draw on. test it with agents to make it accurate & beautiful, then revise" https://homer-catalogue-of-ships.netlify.app/
R to @emollick: If you haven’t tried it, there is a fully narrated tour, historical links, you can read scrolls, you can flash forward to various scenes and theories about the libraries decay and its multiple fires, etc. Open source here: https://github.com/emollick/alexandria-mouseion
R to @fchollet: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
R to @fchollet: Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between AI and human intelligence. We are still working on ARC-AGI-4, which we started developing after releasing ARC-AGI-3 earlier this year. It is coming Q1 2027. We think it's going to be really special.
Been using Astra for the last few weeks and it’s so good and proactive. There’s a bunch of PRs on vitest, tsx or SwiftPM where Astra debugged OC and ended up finding and patching issues in upstream dependencies.
R to @fchollet: Many of you will ask, "if it saturates ARC 3, is it AGI?" We're not making this claim. All we know about the system so far are its benchmark scores. When we launched ARC 3, and in every presentation we made about it, we were very insistent on one thing: solving it is not proof of AGI. It's not intended as a finish line. ARC 3 is testing the right qualitative properties you'd expect of an AGI system -- exploration under uncertainty, adaptation without instructions, causal world modeling from limited data, etc. -- but in small quantities. ARC 3 games are orders of magnitude shorter timescales than real world tasks, and represent orders of magnitude less data, less modeling complexity, less on-the-fly learning. (Slide below is from a March 2026 presentation)
R to @LangChain: For the Pittsburgh meetup, please register here 👉 https://www.eventbrite.com/e/steel-city-ai-innovators-monthly-meetup-tickets-1119887175689
R to @LangChain: For the Pittsburgh meetup, please register here 👉 https://www.eventbrite.com/e/steel-city-ai-innovators-monthly-meetup-tickets-1119887175689
R to @OpenAI: GPT-6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Be ready to experience Astra at its best. Get the ChatGPT desktop app.
R to @OpenAI: GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
OpenAI 宣稱 Astra 在 Agents’ Last Exam、AutomationBench 與 ScreenSpot Pro 三項基準測試達到最新最佳成績,這些測試聚焦不同專業領域的電腦工作流程。貼文沒有揭露分數、基準模型、測試設定或完整結果,無法僅憑此證據判斷領先幅度。
一龍馬判讀
這項主張把競爭焦點推向能操作介面並完成跨步驟任務的 AI 代理,但基準成績能否轉化為真實職場可靠度,仍取決於可重現評測與實務測試。
原文節錄
OpenAI · @OpenAI
R to @OpenAI: Astra achieves state-of-the-art results on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro, benchmarks for computer workflow tasks across…
R to @OpenAI: Astra achieves state-of-the-art results on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro, benchmarks for computer workflow tasks across professions.
R to @OpenAI: GPT-6 Astra is the most intelligent and aligned model in the world, and sets a new state of the art for computer use, browsing, software engineering, cybersecurity, science, and professional work. https://openai.com/index/gpt-6-astra/
We have several LangChain-led and community-led meetups coming up! RSVP and bring a friend! 🇸🇪 9/8 Stockholm 🇺🇸 9/9 Pittsburgh 🇺🇸 9/15 Chicago 🇵🇱 9/15 Warsaw 🇵🇱 9/22 Wrocław 🗽 9/22 NYC 🇫🇷 10/5 Paris (w/ @hwchase17) 🇳🇱 10/6 Amsterdam (w/ Harrison Chase) 🇩🇪 10/27 Munich RSVP 👉 https://luma.com/langchain
R to @fchollet: For a long time, the limits of AI deployment were only defined by technical capabilities. But as AI continues to make rapid progress, it needs to become a deliberate, collective choice about what kind of world we want to shape.
One of the most subtle valuable things about Astra is that it maintains better theory-of-mind: you don't get nearly as many weird references in the final product to previous drafts or work you did during building, and there less drift as it runs long. Not perfect, but very good.
We are working towards getting Astra in everyone's hands as quickly as we can; I know it is frustrating and I appreciate the patience. It should be quick.
We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners. All systems have now been restored and are functioning nominally.
Google 表示,一批新的對話式功能本週開始推出:Gmail 與 Keep 開放給 Google AI Plus、Pro、Ultra 訂閱者,Docs 則限 Pro 與 Ultra 訂閱者。Google Workspace 企業客戶尚未立即取得,官方僅稱將於不久後提供;這則回覆本身未完整說明各項功能內容。
一龍馬判讀
Google 正以訂閱層級切分生產力工具的 AI 功能,個人高階方案將先於企業客戶使用;企業導入時程與管理條件仍有待公布。
原文節錄
Google · @Google
R to @Google: These new conversational features are starting to roll out this week, and they're available in Gmail and Keep for Google AI Plus,…
R to @Google: These new conversational features are starting to roll out this week, and they're available in Gmail and Keep for Google AI Plus, Pro, and Ultra subscribers, and in Docs for Pro and Ultra subscribers. All of these features are coming soon for Google Workspace business customers. https://goo.gle/4iuohNN
We’re bringing new voice capabilities to @GoogleWorkspace to help you tackle daily tasks. These are now rolling out across @Gmail, @GoogleDocs, and Keep to help you search your inbox, organize your thoughts, or brainstorm new ideas conversationally. Here's how you can use them to get things done: 📤 Gmail Live: Skip the manual search and use your voice to quickly find specific details buried in your inbox. 🗣️ Docs Live: Talk through your ideas and build structured, context-aware documents on the fly. 🧠 Keep Live: Turn your stream-of-consciousness brain dumps into organized lists and notes without typing a word.
"Good culture is a prerequisite for everything else. But when it comes to AI, it amplifies everything you already have" https://newsletter.eng-leadership.com/p/good-culture-is-the-biggest-productivity by @gregorojstersek A team's culture determines whether speed compounds or just gets you to the wrong place sooner.
Google AI 表示,WeatherNext 3 自當日起開始強化 Google 搜尋、Gemini、Google 地圖、Google Maps Platform Weather API 與 Google Earth Engine 內的天氣體驗。貼文將模型能力同步導入消費端產品與開發者服務,但未提供預報準確度、涵蓋地區、更新頻率或相較前代的量化改善。
R to @GoogleAI: WeatherNext 3 will begin enhancing weather experiences within Google Search, @GeminiApp, @googlemaps, @GMapsPlatform Weather API and @googleearth Engine starting today Dive into the tech: https://blog.google/innovation-and-ai/models-and-research/google-deepmind/introducing-weathernext-3/
Google DeepMind 表示,WeatherNext 3 將為 Google 搜尋、Gemini、Google 地圖及 Google Maps Platform 的天氣預報提供支援。開發者與研究人員也可透過 BigQuery、Earth Engine 和 Google Cloud Storage 存取即時資料,但貼文未說明涵蓋地區、費用、更新頻率或使用限制。
一龍馬判讀
同一套預報能力進入 Google 的大眾產品與雲端資料服務後,可能同時影響一般使用者、地圖開發者及氣候研究工作;實際可用性仍取決於資料授權、區域覆蓋與服務條款。
原文節錄
Google DeepMind · @GoogleDeepMind
R to @GoogleDeepMind: WeatherNext 3 will now power forecasts in @Google Search, @GeminiApp, @GoogleMaps, and @GMapsPlatform.…
R to @GoogleDeepMind: WeatherNext 3 will now power forecasts in @Google Search, @GeminiApp, @GoogleMaps, and @GMapsPlatform. Developers and researchers can also access real-time data via BigQuery, Earth Engine, and GCS. Find out more → https://goo.gle/4zQe5Fc
R to @GoogleDeepMind: Predicting rain accurately is notoriously difficult for global weather models, with previous methods producing blurry estimates or missing…
R to @GoogleDeepMind: Predicting rain accurately is notoriously difficult for global weather models, with previous methods producing blurry estimates or missing severe storm boundaries. WeatherNext 3 achieves a major leap in global precipitation forecasting, delivering up to a 50% reduction in error - with the greatest improvements in regions where forecasts have historically been less reliable.
R to @GoogleDeepMind: By training directly on raw weather station observations, the system captures localized microclimates across typically underserved areas.…
R to @GoogleDeepMind: By training directly on raw weather station observations, the system captures localized microclimates across typically underserved areas. It also delivers a 5 times resolution boost in temperature forecasts - from 25km down to 5km - in a single pass. 🏔️
R to @GoogleDeepMind: While traditional compute constraints limit standard weather updates to six-hour intervals, WeatherNext 3 ingests real-time satellite data…
R to @GoogleDeepMind: While traditional compute constraints limit standard weather updates to six-hour intervals, WeatherNext 3 ingests real-time satellite data directly to launch a brand-new forecast every single hour.
R to @LangChain: Watch or listen to the latest Max Agency on your favorite podcasting platform. 🎧 Apple: https://podcasts.apple.com/us/podcast/how-unify-cut-its-ai-agent-costs-95-in-two-weeks/id1891551672?i=1000783151404 🎧 Spotify: https://open.spotify.com/episode/6kWQouc2QmiHGk0vdiZEtd?si=ba6e241ca4a24faf ⏯️ YouTube: https://youtu.be/6898VdRtKDE
I think the best way to describe OpenAI's culture is as a mega startup. It's hard to believe unless you see it from the inside. Extreme ownership, care and pace.
串接 OpenCode Go 的工具商與開發團隊需要檢查請求格式,否則既可能失去快取效益,也面臨服務中斷風險;貼文未交代錯誤觸發條件與寬限機制。
原文節錄
OpenCode · @opencode
Some tools using OpenCode Go are missing the x-opencode-session header which prevents optimization of prompt caching If impacted you will receive an email soon…
Some tools using OpenCode Go are missing the x-opencode-session header which prevents optimization of prompt caching If impacted you will receive an email soon with personalized suggestions on how to fix Starting 09/06 requests missing this header may error
R to @emollick: All one prompt & then I asked it to fix the archaeological images so you didn’t have to scroll. That was the only correction and it wasn’t bad, I just didn’t like it.