LangSmith LLM Gateway now supports the OpenAI Decisions API. Your agents get low latency inference, and LLM Gateway gives you a centralized place to control them with model fallbacks, data redaction policies, and spend limits. Read more to get started → https://docs.langchain.com/langsmith/llm-gateway-decision-models
The real world abounds with recursively self-improving systems, but one in particular deserves attention: Science, modeled as a system (perhaps even as an agent, with goals and resources). If you want to really understand AI RSI, science should be your reference point. Science is an intelligent system, and it is obviously recursively self-improving: 1. Scientific discoveries unlock new technology that helps build better experimental tools. This is a top driver of progress in nearly all fields. 2. They unlock new conceptual advances (ideas, theories) that help solve more problems. 3. They increase society's economic output, leading to more resources flowing into science. 4. They unlock better faster tooling (e.g. more compute via better chip & networking technology). As a result, many measures of scientific *input* grow exponentially: 1. Headcount (doubles every ~15 years) 2. Global R&D spending (doubles a bit faster, every ~13 years) 3. Papers and patents (technically this is a measure of headcount) 4. Compute dedicated to science (doubles every ~2 years) But is scientific progress exponential? Historically, the rate of scientific impact over time has remained roughly constant since the start of the industrial revolution (i.e. scientific progress is *linear*). 1850-1900 was about as dramatic as 1900-1950 or 1950-2000. 1850–1900: Evolution, germ theory & antiseptic surgery, thermodynamics, electromagnetic field equations, the periodic table, pharmaceuticals, electricity, telegraph and telephone, internal combustion engine, skyscrapers, mechanized agriculture... 1900–1950: special and general relativity, quantum mechanics, nuclear fission & atomic energy, antibiotics, genetic theory, electronic computers, information theory, synthetic polymers and plastics, the transistor, aviation... 1950–2000: DNA, genetic engineering, integrated circuits, microprocessors & personal computing, the Internet, crewed spaceflight, moon landing, satellite communications & GPS, standard model of particle physics... In real terms, like life expectancy, which has increased in a remarkably linear fashion of roughly 3 months per year since 1840, progress is a straight line. This is especially apparent for fields where impact is easy to measure, like biology, medicine, and agriculture. I first wrote about this phenomenon and its causes in 2012, and a steady stream of research has confirmed it in the years since. Examples include the 2018 paper by Nielsen and Collison, "Science Is Getting Less Bang for Its Buck," and the 2020 economic paper, "Are Ideas Getting Harder to Find?" (In fact, I believe the Nielsen paper stemmed from a conversation I had with him about this exact idea six months earlier) In short, the primary cause is that research solves the highest-impact, easiest problems first, and every subsequent problem is either harder or lower-impact. Exponentially so. The paper that presented information theory wasn't very hard to write (single author!) but you'd have a hard time ever writing a CS paper that beats it in impact. This is why science as a system requires exponential resources (input) to produce linear impact (output). It gets exponentially harder over time. Worth thinking about if you're pondering RSI for AI. I fully believe AI RSI is already happening and will accelerate in the future. But I do not believe this leads to an "intelligence explosion" -- that would fly in the face of everything I know about intelligence and everything I know about recursively self-improving systems.
We're teaming up with Cline and Parallel for an Open Source AI Week hackathon. Join us Oct 19 in San Francisco to build an agent or AI tool, with inference and tool call credits to get you started. Open source your work and share it. Register here: https://luma.com/m8euvq9n
As small open models advance, so does the potential for AI agents to work together. AMD SVP and GM of the Computing and Graphics Group, @JackHuynh explores how teams of agents could multiply individual impact and create new experiences.
🏗️ In this week’s letter from Andrew Ng in The Batch: One of his agents makes 5,000 to 10,000 web searches a day. Agents will need far more data centers. Also inside: 🤖 GPT-6.1 Sol ⚖️ Gemini 4 Argon 🦾 FLUX 3 Action 🧰 HYSET 📬 https://hubs.la/Q04zPZP60 #DeepLearningAI #AIEngineering #AIAgents
What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM? I think MoE might be necessary to get reasonably interactive speeds on the hardware I have access to - I want something faster than 12 tokens/second
It’s Friday 🎉 Here’s a look at this week’s announcements: — http://SynthID.com is now available globally in English, allowing anyone to use our detector portal to easily verify if images, video, or audio files were AI-generated by Google or industry partners — Nano Banana 2.1, featuring leaps in visual design, subject consistency, and precise spot editing that lets you change specific details without altering the whole image — EmbeddingGemma 2, our first natively multimodal open model that unifies text, images, video, audio, and code into a single space for on-device embeddings — @GoogleGemma 4 paired with BOTANIC-1, a specialized AI trained on plant DNA, to rapidly pinpoint the exact mutations that make crops thrive — Gemini agent for your business, a universal cloud agent running 24/7 to get your work done — Guided Vision in Gemini Live, offering real-time visual assistance and dynamic audio descriptions via your camera, built alongside the blind and low-vision community
The challenge for Google post-Gemini 4 will be what they do with a good model. Anthropic & OpenAI show we are moving towards a single interface for many tasks with orchestrator agents. The Gemini 3 era was full of fragmented products for many markets. Won't work for what's next.
New in Managed Deep Agents 0.9: Reactions API for @SlackHQ Channels. When it comes to building effective user agent experiences, loading states are underrated. Now, you can build loading states and receipts using either heuristics, decision models, or custom logic.
Whether agents feel like liberation or loss depends on which joy of engineering you lean on most: making, knowing, or mattering. I've seen engineers say many conflicting things about agents: some are bubbling with excitement and for the first time ever, they feel like they're in an all-star lineup, while others mourn for the death of their craft, leaving them feeling hollow. I'm someone who works with agents every day, so I understand these sentiments. Rather than take sides, my hope is that I help you understand this in a more nuanced way. Too often we see two camps: people who love the joy of coding and people who love the joy of shipping. I think this is a gross oversimplification. Rather than asking you to pick a side, I want to ask you to reflect on the elements of your work that are the most nourishing to you. I think most of us have three sources of joy in our work, which we interweave in multiple ways: - The joy of making (or "flow"): the joy of molding things in our hand and feeling that sweet moment when it all clicks into place. - The joy of knowing: the joy of having such a strong mental model of our systems that we sense where a bug will be just by glancing at files. - The joy of mattering: the joy of knowing that you are the expert that it all relied on. Agents affect each of these elements in subtly different ways, which is why this discussion can be so complicated (we're mourning for different losses, and we're arguing from different places). The joy of making is easy to maintain. it just moves up a level. You're not shaping every line anymore. You're shaping the system or factory, in a loop fast enough that flow still happens. And when you miss the old flow, nothing stops you from writing the code yourself. The joy of knowing is especially at risk of erosion, so it's important to pay attention to it. Sometimes you'll realize you're reaching for a design you should be able to conjure automatically, but your intuition comes up with nothing. Picking an option from an agent's suggestions is a different skill than conjuring an idea in the first place. If you always pick and never conjure, the ability to conjure will erode. You can't direct something you don't even loosely understand. If you don't understand your own codebase, you're not directing an agent to do something, the agent is directing you. You matter. Obsolescence does affect you, but, you, your judgment, and especially your need for judgment, has never been more important. You get to decide what problems to solve, and you get to bring your taste and judgment to the work you do. You get to take ownership of it, all of it, whatever ships. That feeling of sadness isn't a sign of failure to adapt. Grief means that you're attached to something, and that it's natural to feel grief when you see your work changing in ways that feel unfamiliar. Lean into the things that make your work yours, and work on them more.
Right now, using AI agents still feels very single-player. They do your work, but the results are trapped in isolated chat histories and nothing compounds. The Busabase team (@Busabase4agent) is launching an open-source (MIT) database and workspace for agents to fix this. Instead of losing work to ephemeral chats, your agents write directly to a structured workspace to turn their outputs into reusable data, docs, and skills your entire team can access. Details here: https://busabase.com/
Full-stack AI training that scales. Our AMD Data Center team has been working closely with @ZyphraAI to train an advanced reasoning model from scratch. Hear about Zyphra's experience training larger reasoning models more efficiently while supporting longer context windows: https://bit.ly/4ixfdaB
What happens when a table contains tables? Usually, a mess. With LlamaParse, it comes out clean. Doc of the Week: Micron’s latest earnings deck. Even with two business units and same row names, LlamaParse keeps all 18 values under the right header. Try it out on your own earnings deck: https://www.llamaindex.ai/llamaparse
The 3 questions every agent builder should be able to answer: ✅ Where the agent fails ✅ How to reproduce it ✅ How to make it stop Watch @JakeBroekhuizen’s session: https://youtu.be/oVKGKdIAgtg
Randomized trials with old GPT-4o: "AI access raises test scores, and a smaller gain persists a week later. Gains remain for students who use AI as a tutor (“augmentation”) and fade for students who have AI write for them (“automation”)" Across all experiments, positive effects.
Day 5 (dots edition)/ You can now create and text your dot entirely from the ChatGPT mobile app. Impressed so many created them via the desktop/web app previously. Time to scale!
We simplified agent authentication, memory, and channels, and brought in web search as a pre-built tool for Managed Deep Agents. Watch @VictorMoreira16’s full Interrupt NYC session for an inside look: https://youtu.be/XcsspzECTKA
Pro and Max users will soon have Claude Code Projects! I've been enjoying using them. You use a project when you have a stream of related work that outlasts one session. Claude coordinates it, runs each task as a parallel cloud thread and you set the instructions, repos and memory once instead of repeating them. You can also hand off a batch of tasks, walk away and come back to the Overview pane showing which threads finished, which PRs are ready and which need your answer. Check out @delba_oliveira's video for more!
In an urgent care setting, the advice of the obsolete Gemini 2.5 Pro & Gemini 2.5 Flash (without access to patient medical records) were rated of similar quality to doctors by other physicians. There was no safety issues spotted. Models have gotten significantly better since.
✅ 60k queries handled ✅ 500+ customer accounts ✅ 85%+ of sessions resolved without a support ticket ✅ 250+ cases auto-detected + escalated to the right team @SnykSec Assist. Built on LangChain + LangGraph. Observability in LangSmith. Full story: https://www.langchain.com/blog/how-snyk-turned-an-internal-support-agent-into-a-customer-feature
A new agentic research model and harness called AREX checks its answers requirement by requirement, keeps what's verified, and researches only the gaps. 📰✅ 📊 82.5% BrowseComp, 82.0 F1 WideSearch-en 📈 Harness alone: up to +10 pts 🧠 Fine tuned 4B beat untuned 35B on 5 of 6 benchmarks https://hubs.la/Q04ztJKG0 #DeepLearningAI #AIAgents #LLMs
Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality. We’re already testing it across Microsoft for everything from incident response and quality control to scientific discovery.
LangSmith Signals from the last month: ✅ Claude Sonnet 5: #9 to #2 in model adoption. 51% more orgs using it ✅ gpt 5.6 luna: #3 to #1 in call footprints. 65% more calls ✅ Two open-weights models are now in the adoption top 10, but not for call volume ✅ Smaller, faster models dominate call footprints
Day 5/ Composer predictions in the desktop app. Often leads to a double take with how on point they are. Included in the Pro plans without consuming usage.
Just thinking that Paul Simon wrote that we were living in the "days of miracle and wonder" because of the availability of long-distance calls and slow motion cameras.
Its reasonable for mathematicians to point out that proofs are not always the same thing as advancing mathematics. But it suggests a need for new goals for what math is trying to do Same thing will happen everywhere. 100x more PowerPoint or code is not always progress - what is?
Excited to announce that all super intelligence organizations have now jointly agreed to the ultimate in AI/SI safety: Moving all testing to Delta Airlines flights, where accessing the Internet is utterly impossible!
R to @emollick: Google has the advantage, if they use it, of being able to observe where OpenAI and Anthropic are going with Codex/Code and get there now. Otherwise they are going to have to go through that whole learning process themselves, which will make Gemini 4 less impactful along the way.
I built a new feature for my blog entirely by voice with Codex Desktop, while I was cooking dinner https://simonwillison.net/2026/Oct/9/built-using-my-voice/
R to @elonmusk: Also, Starlink has proven to be essential for saving lives during natural disasters throughout the world, when all other communications systems have failed, so you would also be helping save men, women and children throughout India. 🇮🇳
Dear Prime Minister Ambani, Please accept my humble apologies for not realizing that you are the real boss of India. Naturally, you would prefer to maintain your monopolistic exploitation of the great people of India, but would you nonetheless consider allowing Starlink to compete? There are many areas of India, as there are in all countries, with no Internet access, which denies children the opportunity for education and for small businesses the opportunity to sell their goods to a global market. Starlink would change their lives for the better. Thank you 🙏
R to @swyx: most of you are unfortunately not qualified but there is somewhat a path https://x.com/adelwu_/status/2108424235341479984?s=20 wrote down mine in @Coding_Career 6 years ago and you can now get it free or on amazon https://learninpublic.org didnt intend to relaunch CC today but eh why not
Here’s how to get the lowest possible price on your holiday flights, according to five years of aggregated Google Flights data. 👇 ✈️ Connecting flights can also be a big money saver. Nonstop tickets are 20% more expensive on average. ✈️ When you travel is key for saving: Departing on a Monday and returning on a Tuesday or Wednesday saves about 14% over weekend flights (and up to 20% domestically versus Sunday). ✈️ Wednesday beat Tuesday as the absolute cheapest day to book a flight, but the difference is negligible — buying a ticket on a Wednesday only saves you about 1.4% compared to Sunday, the most expensive day.
R to @claudeai: We’ve loved seeing so many founders join the Claude Startups program. Unfortunately, we underestimated demand for the program, and we need to pause the Claude Team and $1,000 API credit offers. With hundreds of thousands of applicants, we're re-reviewing applications so the program can serve as many startups as possible. If you’ve already claimed the Claude offers, they’ll still remain in your account. If you applied or were approved and haven’t claimed the offer, we're re-reviewing your application. Unfortunately, this means that some accounts who had already been approved will not get access to the offer. We know this is disappointing and we’re very sorry. All accepted program members still have access to the Startup Stack and Applied AI office hours. If you're on a Claude Max or Team plan, your monthly API credits are available in Console as usual. We appreciate your patience, and we're working to get you building as quickly as we can. https://x.com/sarahzorah/status/2108401845563806150?s=20
RT by @ylecun: June 1986. It was my first conference in the US and my first time visiting MIT. Had a poster on multilayer nets. I met Marvin Minsky for the first time at the reception hosted by Thinking Machines Inc (the original one that built the Connection Machine). He was surprised that you could train a neural net to compute the product of two binary numbers (I had done the experiment). I was on my way to the first Connectionist Summer School at CMU. The Bell Labs folks, who I met the year before, heard I was in the US and invited me for a talk in Holmdel on my way back to Paris.
R to @simonw: Admittedly I guess this is a pretty incredible troll targeted at people who say "it shouldn't be called artificial intelligent because it's not actually intelligent" Do they now feel obliged to defend the old name against the new one?
R to @simonw: Note that their press release from when he pleaded guilty back on March 19th used "North Carolina Man Pleads Guilty To Music Streaming Fraud Aided By Artificial Intelligence" because of course it did https://www.justice.gov/usao-sdny/pr/north-carolina-man-pleads-guilty-music-streaming-fraud-aided-artificial-intelligence-0
R to @simonw: This tweet inspired by the justice department headline from October 6th: "North Carolina Man Sentenced To 18 Months In Prison For Super Intelligence-Assisted Music Streaming Fraud" https://www.justice.gov/usao-sdny/pr/north-carolina-man-sentenced-18-months-prison-super-intelligence-assisted-music
Excited to support the Genesis Mission @WhiteHouse @ENERGY and Director @mkratsios47 with $150M in investments this week. Looking forward to collaborating on the future of science!
R to @Google: Conducted in partnership with @BIDMC_Medicine, the study demonstrates technology’s potential to strengthen the relationship between doctors and patients. Learn more ↓ https://goo.gle/3VCGhMl
Our medical research system, AMIE, is the first patient-facing conversational diagnostic tool of its kind to be studied prospectively in a real-world clinical setting. In a study, published today in @TheLancet, we found that when patients chatted with AMIE before an in-person appointment, the interaction built their confidence and helped them organize their thoughts. And physicians who reviewed the information from the patient conversation got more time back to focus on collaborative care and shared decision-making instead of digging through data.