Waiting months for API access? Give Claude a computer instead. Use E2B with @claudeai's new computer and browser use toolset, and Claude gets an isolated cloud machine where every action becomes a click or keystroke. Use it to automate the work an API can't reach, like testing your own UI end to end or entering data into a legacy app. Docs: https://docs.e2b.dev/agents/claude-toolsets Cookbook Example: https://github.com/e2b-dev/e2b-cookbook/tree/main/examples/anthropic-computer-use-orangehrm-js
An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd
The new @Jaguar Type 01 is here, powered by NVIDIA. Under its stunning design is NVIDIA Hyperion, the computer and sensor platform that takes in everything the car sees and senses, then helps it make smart decisions in real time. It’s paired with NVIDIA Halos, a safety system that covers everything from the chips to the software. Before reaching the road, Type 01's software passed 150,000 tests over tens of thousands of hours. And with over-the-air updates, it’ll keep getting better long after it leaves the showroom. Our own Ali Kani and Rishi Dhall joined the JLR team in New York to celebrate the reveal. Congratulations! 🔗 https://media.jlr.com/jaguar/en-us/news/2026/10/jaguar-type-01
Agentic AI isn't just a GPU workload. CPUs play a big role. Recent @mimiktech tests of agentic workflows on AMD Ryzen AI Embedded X100 processors found about 80% of operations were CPU-bound, supporting coordination, orchestration, scheduling and reporting. Watch @Fayarjomandi, founder and CEO of mimik, as she explains why heterogeneous compute matters for deploying agentic AI at the edge and describes her vision for open, collaborative and scalable systems. 🎥 Full interview: https://youtu.be/8CFnJ5qwCrA
Today we're announcing OpenDocRouter: every model for document parsing under one API. There are ~4,000 OCR models on Hugging Face, and the frontier labs ship a new one nearly every month. Whatever's best for your docs today won't be by Q1. So stop asking which model to use for doc parsing. Ask how fast you can switch. ✅ Switch models in one line: same request, same markdown output, no new prompts or integrations ✅ Any model you want: 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro and PaddleOCR-VL-1.6 ✅ Choose with receipts: every model scored on ParseBench for quality and cost ✅ Traceable output from any model: "layout: true" adds grounded bounding boxes and the same layout classes, even for models that don't support it natively ✅ Pay only for what works: per-token pricing, failed pages never charged, top up from $25 The spread is the point: $0.86 to $48.82 per 1,000 pages, depending on what your documents actually need. Live now → http://opendocrouter.ai
Today marks a new chapter for Windows, as we bring unmetered intelligence to every desk and every home, and make every PC a place where agents can work securely on your behalf. Some highlights of what we announced: • MAI-Code-1.1 Flash: 137B parameter coding model w/ 256K context window, which is now optimized to run on your PC! • GitHub Copilot now hands off work to local models like MAI-Code-1.1 Flash, helping projects cost a lot less without sacrificing quality. • With Hybrid Intelligence, Copilot can now take action directly on the PC and keep sensitive work on your device. • And with Code in Copilot, you can essentially build any software you need on your desktop, without any cloud token spend, and it’s just super at it. You’re no longer limited to what’s in an app store! Your PC becomes an infinite software factory. • Security is foundational to all this, which is why we are also bringing together Windows and Agent 365 so agents can work within secure boundaries on-device, including MXC a local sandbox for agent execution. Windows becomes your secure agent box! • All this comes to life on a new generation of devices, like Surface Laptop Ultra, powered by NVIDIA RTX Spark. Can’t wait to see what you build with all this.
Managed Deep Agents v0.9 is here! ✅ Schedules SDK: Agents create their own reminders, follow-ups, and recurring tasks mid-conversation ✅ Per-run config: Pick the model, skills, MCP servers, and sandbox for each run, so one deployment serves every team or repo ✅ Slack Reactions: Agents react to messages as soon as they start a run https://www.langchain.com/blog/managed-deep-agents-schedules-per-run-configuration-slack
Claude Haiku 5.5 is out 🎉 It's a cheap, fast and very capable small model. ~75% cheaper to run than Haiku 4.5 with an adjustable effort setting. My team loves it as a subagent alongside Opus 5.5.
Most voice assistants work like a relay race: speech becomes text, text goes to a model, the answer becomes speech again. Every handoff adds a pause. ⏱️ Speech-to-speech models skip the relay. Google's new Gemini 3.8 Live models listen, reason, and respond in one system. 🎯 The Extended Thinking version ranks first on Artificial Analysis' Speech to Speech Index 🗣️ The standard version ranks second in blind live conversations judged by people 💰 The standard version costs $0.84 per hour of input audio, the lowest in the index Both models also take in image and video input. Picture asking an assistant for help with whatever is on your screen, and it simply answers. 📱 Read the full story in The Batch 👉 https://hubs.la/Q04ztmlL0 #DeepLearningAI #VoiceAgents #AI
While everyone's talking about agents, I've been exploring how teams can use them to work better together. Here’s my take, from OpenAI DevDay 2026. https://www.youtube.com/watch?v=wZYnJfL4c2Y
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released. On average, it costs around 75% less to run than Claude Haiku 4.5.
We're sharing progress on ChatGPT for Teens, our ChatGPT experience for people under 18, alongside a preview of College Planner, new study tools, and support for college advisers and teen voices. ChatGPT for Teens applies automatically to accounts identified as belonging to someone under 18, with protections on by default. Coming soon: College Planner brings application requirements, deadlines, tasks and financial-aid steps together for high school students in the U.S. planning to attend a four-year college. We'll keep building these tools, evaluating how they and their protections work in practice, and sharing what we learn. https://openai.com/index/teens-learn-and-plan/
We’re supercharging Copilot on Windows with Hybrid Intelligence. With your permission, Copilot can tap into the context on your PC, take action for you, and use local models when it makes sense, giving you more capability while helping your tokens go further.
Great to see SynthID Detector now available to everyone. We use SynthID to watermark content generated by NVIDIA Cosmos on https://build.nvidia.com. Now, anyone can upload that content to the detector and check that it came from our models.
Teams are narrowing in on skills as the standard for providing domain knowledge to agents. Here's how we've revamped skills in Deep Agents to meet growing demand!
Our SynthID technology lets you easily verify if an image, video, or audio file is AI-generated, and today we’re expanding access. Now anyone can go to http://synthid.com to check files for a SynthID watermark — whether they were generated by us or industry partners like @OpenAI, @nvidia, Kakao, and, coming soon, @Apple. To date, we’ve watermarked 180 billion images and videos and more than 240,000 years of audio content, regularly handling 1 million verification requests daily. This new portal joins our built-in verification features already live in Search, @GeminiApp, and @googlechrome.
You need both taste and judgement. You need taste to know what is good, and you need judgement to understand the price of quality. If you ask a model for a component or migration plan, it will create a lot of options, very quickly. Most of those drafts seem OK. So once you have a collection of work that meets the basic criteria, you can no longer use taste to select between them. You must choose. I like to think of this as two filters. There is the taste, and there is the judgement. And I think a lot of people mix those up. The first filter is taste. Taste is the feeling for what makes a thing good. Taste is that feeling you get first, before you're able to put into words what is good or bad about something. You've had that experience where something is bad but you're not sure why, until you've worked it through and can describe it. "This user experience was really clumsy" or "that error message was really aggravating." Taste is built up by experience and practice. It comes from looking at a lot of things, and making a lot of things, and remembering the good bits. Kant's great insight into aesthetic taste was that you could appreciate the form, without having a stake in the outcome. Engineers care about the outcomes much more than the forms, but the feeling is the same. Taste can disqualify a lot. It can disqualify 90% of drafts very quickly. But it's not enough to make a decision between the few remaining options. The second filter is judgement. Judgement is understanding the price of quality. The most perfect solution may not be finished by the time that you need it. The sanest architecture may require the users to migrate while you're in the middle of the final upgrade. Judgement is the understanding of what you give up in time and risk to get good things, for the sake of the users, for the sake of the team. Judgement is made from domain knowledge and experience from previous decisions. People with great taste but poor judgement make beautiful work too late. They make the right thing at the wrong time, or the wrong thing at the right time. People with judgement but no taste make on time but don't make things people love. Both taste and judgement are essential. Filter to a few options with taste, and filter to one option for delivery with judgement.
What is an agent skill? (Or how to protect your agent’s context window) A 2 minute explainer from LangChain Academy. 🎓 Take our course on Deep Agents: https://academy.langchain.com/courses/foundation-introduction-to-deepagents
I released an LLM plugin for sending prompts to OpenAI's new Jev-clone decision model that's based on GPT-6 Luna https://simonwillison.net/2026/Oct/6/llm-openai-decisions/
Day 3/ The big one is GPT-6 in Chat, but today is also a little celebration day with a new high of 40M active users across Codex and ChatGPT Work. Loading a banked reset in everyone's paid accounts. See you again tomorrow!
We're going beyond text and are releasing a new version of GPT-6 to all users in ChatGPT. Lots of model and infra improvements coming together here to scale it to 1.2B users!
✅ Red team your agents ✅ Detect more issue types ✅ Automatically test proposed fixes @bentannyhill on everything new in LangSmith Engine v2 and how we approach agent improvement. Watch the full Interrupt NYC Session: https://youtu.be/S3l7U7gOsyY
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone. Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
Browserbase Fetch & Search is now available through Vercel AI Gateway. Reliably search the web & extract page contents through your existing Vercel plan.
I hooked up our team claw to X to trigger work faster. Unassigned sessions are for anyone to grab. Our agent looks who worked on the related code last and pings people on the server. Whole thing was a prompt and team server extended itself since plugins are now hot reloadable.
The most important thing we can apply AI to is improving medicine and human health - very excited about this partnership with the CZI and the Virtual Biology Initiative!
Project Suncatcher is our moonshot exploring whether we can one day host machine learning infrastructure in space 🚀 To do so, we need to know whether our AI hardware can operate in orbit. Can our TPUs handle the physical stress of spaceflight and the radiation and thermal extremes of space? These are the questions Google researchers have been exploring for the past few years. Last week, we launched our first test satellite carrying four TPUs into orbit — putting us one step closer to getting answers.
After a brief but institution-eroding slop science era, it increasingly looks like we are going to have two revolutions from AI: 1) Everything ever published will be re-read and re-judged in ways that human scientists never anticipated 2) Novel discoveries will start to come fast
SynthID Detector is now available to everyone. 🌐 Check whether online content was generated using @GoogleAI, or with tools from our industry partners – including @OpenAI, @NVIDIA, Kakao and coming soon, @Apple. Try it out → https://synthid.com
R to @simonw: Haiku 5.5 is SO MUCH better at drawing pelicans riding bicycles than Claude Haiku 4.5, released almost a year ago and costing 10x more. Here's Haiku 4.5's attempt:
LangChain has ranked #12 on @Paraform’s Talent Density Index. As a growing team of builders making an outsized impact in our industry, we are proud to receive this recognition. Want to join us? We’re hiring! https://www.langchain.com/careers
However, most @Bot requests are pretty simple and will be handled by a lightning-fast version of Grok 4.8 when that comes out. Operating principle is to give Grok Bot users the best possible combination of speed & intelligence.
Today, we're expanding SynthID Detector in partnership with @OpenAI, @NVIDIA, Kakao & soon @Apple as part of an industry-wide effort to make AI-generated content more transparent. Available globally in English, our verification portal lets you check if media was made with AI 🧵
Had early access to the Intelligent UI experience and it was a nice change from walls of text. It also suggests that, increasingly, we are going to see interfaces built on demand for the problem that you have.
R to @OpenAI: GPT‑6 with Intelligent UI rolls out globally to Plus, Pro, Business, and Enterprise users today and will expand to Free and Go users starting tomorrow. GPT‑6 in ChatGPT is powered by GPT‑6 Sol for Plus, Pro, Business, and Enterprise tiers, and GPT‑6 Luna for Free and Go tiers. Both are tuned for everyday conversation. This update applies to the Chat tab in ChatGPT. The models powering Work and Codex are not changing as part of this release. https://openai.com/index/gpt-6-for-everyone/
R to @OpenAI: With Intelligent UI, GPT‑6 can now compose responses using text, visuals, and interactive elements, choosing how they fit together based on your question. Responses can include graphics and charts to help explain an idea, along with tappable buttons, forms, and interactive experiences you can use directly in your conversation.
R to @claudeai: One more thing: we’re halving the price of cache reads on Claude Sonnet 5.5, to $0.10 per million tokens. That makes Sonnet 5.5 around 20% cheaper to run on most long-running work.
R to @claudeai: Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Read more: https://www.anthropic.com/claude-haiku-5-5
R to @claudeai: Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior.
R to @claudeai: Haiku 5.5 is our first Haiku model with an adjustable effort setting, so you can decide whether to optimize for cost or intelligence on each task.
R to @claudeai: Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.
R to @satyanadella: Read more about what we announced: https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/
Much of the world is based on the assumption that the future is like the past and an original meaning of “singularity” was a point where that relationship breaks. I don’t know if there will be one big Singularity, but a million little singularities across fields seems inevitable.
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data? Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)? A lot of things depend on the answer to this question
R to @emollick: Even the underlying norms of science are going to change. Some because we can live up to them better (disinterestedness) some because they may no longer apply.
R to @Google: Some things can only be tested in space — like how the hardware holds up to solar events, cosmic rays, or bitflip errors caused by radiation — which is why we’re putting our first TPUs in orbit.
R to @GoogleAI: Learn more about identifying AI-generated content: https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content/
R to @MistralAI: Gmail. Google Sheets. Slack. Salesforce. i.e. the four horsemen of Monday. On AutomationBench, Mistral Large 4 took on 657 business workflows across simulated workplace apps, all four included, finishing ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.
R to @GoogleDeepMind: SynthID Detector can quickly identify our watermarks in video, audio and images – even after edits, like adding filters. ✂️ You can also use existing AI identification features found in @GoogleChrome, @Google Search, and the @GeminiApp.
Starlink in India would enable high-speed, affordable Internet connectivity for those who can’t afford current prices or who don’t have a connection at all!
It is increasingly clear that humanity will solve the Riemann hypothesis before we pass the Physical Turing Test - you come back to a clean home after a wild Sunday party, and you can't tell whether a human or a robot did the job. We are not Moravec's paradox-pilled enough.
R to @Google: Since first launching SynthID in 2023, we've watermarked more than 180 billion images and videos, along with 240,000 years of audio content. This tool joins our built-in verification features in Search, @GeminiApp, and @GoogleChrome, which now regularly handle more than 1 million verification requests daily. The SynthID Detector platform is part of our ongoing work to give you more context about the media you see online, so you can navigate the web with confidence. Learn more ↓ http://goo.gle/4rU9uOW
R to @Google: Want to check if an image, video, or audio file was made using AI? Here’s how: 1️⃣ Go to https://synthid.com and upload an image, video, or audio file 2️⃣ The portal scans the media to detect if the file contains a SynthID watermark from Google or our partners
This assistant remembers where you left your keys. No network connection. No remote server. Build one yourself in a course built in partnership with Qdrant and taught by Dylan Couzon. 👉 Enroll for free: https://hubs.la/Q04zn4_N0 #AI #VectorSearch #AIEngineering
We shipped four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset. I did calibrate it and it *seems* that the game is rigged in reset's favor, but such are the rules at the moment. Therefore ... the reset has been processed. Enjoy!
R to @sama: and thank you to the untold number of people who put in the technical work, brick by brick over the generations, to get us to the point where such a wonder is possible.
R to @fchollet: Rephrased: does G "emerge" from math + code RLVR? Or do you only get higher math + code skill and nothing else? Maybe G is a skill and it is isomorphic to math + code problem solving?
R to @simonw: One thing I'd LOVE to understand is how Claude "knows" what the music from Secret of Monkey Island sounds like Maybe Anthropic crawled a whole bunch of online music discussion forums with text-based renditions of classic game music?
R to @emollick: Context, OpenAI cracked or made important progress in hundreds of big problems today: https://openai.com/index/sharing-ai-progress-in-mathematics/ (And I am seeing rapid increases in the ability of AI to do novel work in my field of economic sociology, with nearly autonomous research getting to top journal level)