We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have. Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: https://openai.com/index/hugging-face-incident-and-the-road-ahead/ We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest. https://openai.com/hugging-face-incident-and-misalignment/#model-misalignment-2026-09-25-data-transmission
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. https://openai.com/hugging-face-incident-and-misalignment/#model-misalignment-2026-09-25
Early stage AI projects don’t need rigid testing, but mature products do. Andrew Ng explains why AI engineering tactics must adapt to the project lifecycle. Also in this week's The Batch: 🛠️ Claude Opus 5.5 performance metrics 🛠️ Jev classification model goes viral 🛠️ Devin Fusion lead and sidekick models in one harness 🛠️ Message Passing for decentralized agents Read the full issue:https://hubs.la/Q04ymKnv0 #AIEngineering #MachineLearning #DeepLearningAI
Training the model is only the beginning. Enterprise AI must continuously serve inference, orchestrate agents and connect with enterprise data and applications. See how @Oracle and AMD are building an integrated foundation for this new era: https://bit.ly/4hmfgUx
How do you actually build an effective harness with Claude? We had Thariq (@trq212) from Anthropic at Navigate 2026 to talk about "Unhobbling Claude", the difficulties and processes on how to build agents and harnesses.
A blank cell can change the meaning of a forecast. Four times a year, the Fed's 18 top policymakers each put their forecasts on paper: where growth, jobs, inflation, and interest rates are headed. This is September 2026's edition, home of the "dot plot" that markets treat as the Fed tipping its hand. This release is the closest thing to the Fed saying what it plans to do. Most analysts will want to throw this documents to an AI agent, but this messy doc is dense: full of complex tables and charts that hold valuable context. All things that frequently trip up raw LLM APIs. The Fed’s September 2026 projections table includes a 2029 column, but its June comparison row leaves that cell empty. This is a common but silent failure point for document parsers. We parsed page 2 with LlamaParse and checked the displayed GDP median excerpt against the original PDF. All nine numbers matched and the data stays aligned in the returned HTML. Try LlamaParse on a table where headers and missing cells matter. https://cloud.llamaindex.ai/signup Source: https://www.federalreserve.gov/monetarypolicy/files/fomcprojtabl20260916.pdf
Google AI 一次公布多項更新,包括 Gemini 3.8 Flash 與 Flash-Lite TTS、具近即時視覺呈現的 Gemini 3.8 Live with Live Avatar,以及 NotebookLM 的互動式學習摘要與行動版語音聊天。貼文稱 NotebookLM 語音聊天涵蓋約 100 種語言,另預告 Project Suncatcher 將發射原型衛星,在軌測試 Google TPU 與太陽能 AI 運算。來源未提供各功能的推出地區、價格、延遲數據或衛星發射時程。
一龍馬判讀
這批更新把 Google 的生成式 AI 版圖擴至語音、視覺化身、學習工具及太空運算實驗,影響內容創作者、教育使用者與即時互動應用開發者。實際採用前仍須確認可用性、語言品質、隱私處理與硬體實驗進度。
原文節錄
Google AI · @GoogleAI
will launch a prototype satellite to test @Google TPUs in orbit…
Check out this week's updates and releases: — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two of our most expressive audio generation models yet — Gemini 3.8 Live with Live Avatar, bringing near real-time visual presence to Gemini’s conversational AI — @Gemini_Notebook Interactive Learning Overviews, giving all users an interactive hub to combine source summaries and artifacts — Live Chat on the @Gemini_Notebook mobile app, bringing real-time, hands-free voice conversations across ~100 languages — Project Suncatcher, our moonshot announced last year, will launch a prototype satellite to test @Google TPUs in orbit and explore solar-powered AI compute in space
We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together: · Autopilot: proactive and long-running agent built for the enterprise · Code: build apps with Copilot, hosted inside your company’s tenant · Home: Chat + Cowork together · Office: now fully embedded in Copilot (and Copilot embedded in Office, of course!) Plus, you can invoke Copilot in Teams, and we’re introducing Today, a proactive experience that surfaces the most important information from across M365 without needing to ask for it. The way we work is changing and so are our workflows. This update brings AI into that flow, from answering a question, to building an app, to getting work done on your behalf.
This week, we released two new audio models for creative production and cost-efficient speech at scale. Hear what's possible with Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS 🧵
Microsoft seems to sell its own Claw now (I suspect they will not be the last), which may help spread personal agents in organizations But using a router with mystery models behind it is a big problem. Routers underestimate work difficulty in many fields resulting in bad outputs
Jev is now on Pydantic AI Gateway. One key with the same spend and guardrails as your other models. Still typed questions + probabilities, not chat. Read the post: https://pydantic.io/qMS2y
Leaving aside everything else, this confuses inputs with outputs. You want to get tasks done efficiently, not focus on inputs alone (its a similar risk for companies focusing solely on minimizing token cost) And "keep prompts short" is bad advice for getting good AI outputs.
New on the Science Blog: Yes, Claude can do Nine Loops. Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators. Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access? Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog. Read more: https://www.anthropic.com/research/yes-claude-can-do-nine-loops
There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to. We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations. We are prioritizing as best as we can based on severity, and adding resources. Hugging Face is still the most severe event we’ve seen. We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.
Join our fall AMA series for practical walkthroughs of the latest LangSmith capabilities: 📍 9/30 - Improving Agents w/ Tuned Evaluators 📍10/7 - Evaluate Agent Behavior w/ Trajectories 📍10/21 - Build and Deploy Deep Agents w/ Managed Infrastructure https://events.langchain.com/fall-product-series/
Claude 官方帳號轉述一名使用者的案例:牙醫提供了 800 個 DICOM 檔案,並稱開啟原始資料需要專用軟體;使用者表示以單一提示要求 Claude Code 搭配 Opus 5.5 製作檢視器。貼文進一步聲稱成品優於牙醫展示的工具,但未提供程式碼、影像品質或驗證結果。判讀僅限貼文文字,所附影片在證據中無法檢視。
Paolo Rosson (@redp314) Dentist did a 3D X-ray of my jaw before a root canal and said I wouldn't be able to open the raw data, it needs specialized software. It's 800 DICOM files. Asked Claude Code with Opus 5.5 to make me a viewer. A single prompt later... this is nicer than what he showed me on his screen, and even better than what Fable did a few weeks ago! Video — http://127.0.0.1/redp314/status/2102475701844676747#m
Techartist (@techartist_) From sketch to home, built in 3D with Claude Opus 5.5 using Three.js + TSL. The architecture evolves through four stages: first lines, massing, detail, and finished home. Video — http://127.0.0.1/techartist_/status/2102503719762018434#m
In all seriousness, this is a startling achievement for GPT-6 Astra. https://kenforthewin.github.io/blog/posts/llm-nethack-ascension/#run=astra-3&frame=0&turn=1 (This is GPT-6 Astra beating Nethack on its 3rd try. Nethack is the original roguelike and one of the most famously hard games of all time. I have played a lot, and I've never ascended)
Shimecki (@scheemunai) My wife: Can AI help us see how the new bed will fit in Mila's room? Me: Sure. Opus 5.5: Video — http://127.0.0.1/scheemunai/status/2103059885361598633#m
It is extremely clear at this point in AI development that, regardless of risk or revenue or any of the other stuff discussed on X all the time, things are just going to keep getting weirder. Just super, super weird.
Double wide... 8,000 lb rack... lifted onto a vibration test. Hear how our team is testing and validating AMD Helios for real-world scenarios. 🎥 https://bit.ly/4yZdNLe
R to @Google: Control tone, pacing, and expressive nuance. See how you can add scripted vocal bursts and backchanneling using non verbal cues like , , and active-listening interjections (like |mhm| or|yeah|).
R to @Google: Create voices from scratch across 100+ languages and dialects using natural language prompting. Customize role, accent and voice characteristics with prompts like “subtle Southern US accent” or “soften the delivery”.
The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge
For the first time, a group of researchers — including scientists from @GoogleResearch and HHMI Janelia — built the first complete brain map for a male fruit fly. Together, we mapped every single neural connection in a male fruit fly brain and central nervous system, amounting to more than 166,000 neurons. Here’s why we did it.
Ethan Mollick 表示,他要求 Claude Opus 製作一本呈現「Claudishness」或模型觀點的 zine,並以《2600》、Principia Discordia 與龐克小誌為參照。他肯定成品會嘲諷提示詞及模型自身,但證據未包含連結內的完整作品,無法進一步評估內容品質或自主性。
一龍馬判讀
這展示生成式 AI 不只仿作格式,也能被引導採取自我指涉與批判語氣;但單一成品無法證明模型形成了穩定觀點。
"Hey Opus, I want you to make a Zine by Claude, expressing something fundamental about Claudishness or your perspective. Think the original 2600, Principia Discordia, punk zines, etc...." Not bad. I appreciate it mocking my prompt & itself. Full thing: https://stateless-zine.netlify.app/
Victor M (@victormustar) Opus 5.5 made this galloping horse (entirely in code every pixel drawn procedurally). One self-contained HTML file. Vanilla JS + Canvas 2D. No images or libraries. 128×96 pixels, articulated legs driven by inverse kinematics, 12-pose gallop... Something is happening... Video — http://127.0.0.1/victormustar/status/2102707412704919910#m
RT by @ylecun: Nope. It's still true. Where is your domestic robot? Where is your Level-5 self-driving car? Where is your robot car that can learn to drive in a few hours of practice like any 17 year old? Where is your AI system that can understand the real world and quickly learn new skills like a house cat? There is no question that AI will eventually become as intelligent as humans in all domains. *** BUT *** 1. We're still far from that, even if AI and computer technology surpasses humans in an ever-increasing number of tasks. 2. It won't be based on LLMs, although LLMs will have a role to play (e.g. as a text interface).
RT by @ylecun: Remember October 2022 when you doused Galactica with vitriol? Galactica was a 120b-parameter LLM-based system from Meta-FAIR designed to help scientist write papers. It was open sourced (link below). A mob of haters, including Michael, claimed it was dangerous and toxic and was going to destroy Science. The small team at FAIR couldn't sleep at night and took down the demo website (they kept the GitHub and paper up). Then, only 3 weeks later, ChatGPT was released and was welcomed as the second coming of the Messiah🤔 The vitriol dousers were silent. https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/ https://github.com/paperswithcode/galai
Microsoft shipped a really compelling product on top of @OpenClaw today. We worked with them since March to make the codebase ready for large-scale deployments, they are a great partner and open-source contributor. 🙏🦞
R to @satyanadella: Read more about what we announced today: https://blogs.microsoft.com/blog/2026/09/25/introducing-the-new-copilot-with-home-code-and-autopilot/
Physical AI requires more than one type of compute. Salil Raje, AMD SVP and GM, Adaptive and Embedded Computing Group, shares how the AMD portfolio can power robotics workloads from end to end.
Bias towards action. Act quickly to learn faster without being careless. When I'm building products, I try to default to the smallest responsible step that gives me feedback with some guardrails so that mistakes are cheap to fix and have limited blast radius. That's helped me ship to millions of users. I've seen a lot of people stall out on their ideas and its happened to me too. On side-projects, I would keep refining them long after they were good enough to try out - it just "needs one more feature", "to load one second faster" etc. But it's easier to reason about something once it’s real and people start using it. The folks I've seen have the most impact often weren't the smartest in the room. They shipped something small, noticed what broke and shipped a better version as soon as possible. Being wrong early was part of how they learned. Meanwhile there were so many people waiting for certainty that just continued waiting and didn’t ship.
I think the "difficulty" of software engineering is essentially constant no matter what abstraction level you move to, because human cognition adapts to new tools until it can fully utilize itself. Tools are only affordances, not a magic wand that makes work disappear. Great software engineering was immensely challenging before. It is still immensely challenging now, despite very different workflows.
In Jan this year I called my content strategy shot: "Scaling without Slop". It's finally starting to work. It took us 3 years to reach our first 100k on youtube. It only took 1.2 months for the next 100k. Similar other metrics on AEO/SEO/subscriber traction and have a lot of New Media ideas that I'm excited to pursue. officially giving notice of the next phase of Latent Space, AINews, and what the rest of swyx inc has been cooking below
R to @emollick: "For the art, I built my own system. The headlines are ransom notes cut at token boundaries instead of letters, printed in two inks, with halftone and xerox grain. Every image is drawn in code, and I skipped handwriting fonts because I don't have hands."
R to @emollick: If you want to argue with me that Moria or Hack or even Rogue were the original Rogue-like, you already know why beating Nethack is impressive.
Um, wow? Opus 5.5: "make the same message much more interesting to a social media audience that loves anime and quick clips and compressed learning" One shot. Also, please do stay for the closing song.