Missed a few days of AI news? Here’s what happened in this week’s Data Points 🧵 🔹Google is back at the frontier. Gemini 4 Argon ties GPT-6 Astra at 53 on the Artificial Analysis Intelligence Index, at about 60% of the cost per task. 🔹Europe closes the gap. Mistral Large 4, a 1T-parameter model with open weights coming late October, is now the most intelligent model built outside the US and China. 🔹Reflection AI’s open-weight Beam claims 3–4x less compute to reason than comparable open models. 🔹OpenAI published a batch of new math results from an unreleased model, with proofs formalized in Lean. 🔹Microsoft’s MAI-Transcribe-2-Streaming ranks #1 of 38 for streaming speech-to-text accuracy. 🔹Google’s EmbeddingGemma 2 runs multimodal search on a phone; Cohere’s Embed 5 targets enterprise RAG. 🔹Meta’s Muse agent and your friends’ data, OpenAI’s opt-in text watermarks, arXiv’s new submission caps, and which AI subscriptions give the most value per dollar. Read it all 👇 https://hubs.la/Q04ztBjc0
For most global teams, flexibility matters as agent workloads grow, architectures evolve, and infrastructure requirements become more complex. Join @AWScloud, @NVIDIA, and LangChain to learn how teams can approach the infrastructure layer behind production agent systems. Save your spot: https://aws-experience.com/apj/smb/e/0700a/architecting-flexible-agentic-ai-on-amazon-eks-with-nvidia-and-langchain/
Most measures of AI progress closely fit an exponential over the 2023-2026 period. Meanwhile AI capex over the same period has been slightly super-exponential (i.e. the rate of growth itself is increasing). So if you model AI as a system that takes as input external investment and outputs AI progress, that system has a slightly sub-linear response. But super-exponential injection of *external* resources is not sustainable in the long-term. To achieve sustained exponential progress, the AI industry will need to achieve matching exponential profits -- it will need to close the resource loop.
An astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky. Astronomers have already created complete sky maps from radio to gamma rays. But large regions of the sky have never been observed in UV. In today’s science blog, Brice Ménard explains how he guided Claude to find existing datasets, combine them, and fill in the gaps with statistical inference. The work would have taken humans weeks of work, but, with Claude, just took a few days while Ménard worked on other projects. The resulting map is a useful teaching tool and an example of exactly the sort of slow, low-priority work which scientists never quite get around to, but which AI now makes feasible. Read more: https://www.anthropic.com/research/the-missing-map-of-the-sky
Excited to announce that all super intelligence organizations have now jointly agreed to the ultimate in AI/SI safety: Moving all testing to Delta Airlines flights, where accessing the Internet is utterly impossible!
We're deepening our commitment to scientific research and technological discovery by committing $150 million to the Genesis Mission. We'll also make Claude and our technical support available to more than 15 federal agencies. https://www.anthropic.com/news/genesis-mission-commitment
Humans provide a fun comparison point here. In the 80s, the "cognitive consequences of programming" was a major psychological research focus. The prevailing hypothesis was that learning formal logic and algorithmic structures through computer programming would act as a form of mental gymnastics that could upgrade general reasoning, planning, and novel problem-solving abilities. But foundational research, followed by subsequent meta-analyses, demonstrated that this does not happen. Students who learn programming become highly proficient at algorithmic thinking and coding, but these gains do not transfer to general cognitive tasks or unrelated problem-solving scenarios. Similarly, intensive mathematical training improves mathematical deduction and structural mapping, but does not increase an individual's baseline rate of skill acquisition in unrelated fields. General intelligence seems to be a fundamental property of the brain rather than a skill you can train. Practicing in a domain makes you better at the domain but does not make you generally smarter.
Your agent isn't good enough to take your job, well not yet. Howie Liu (@howietl), Founder of Hyperagent (@hyperagentapp) at Navigate 2026 to discuss how to actually make agents employable.
As a researcher who did early work on the productivity impacts of AI chatbots using RCTs, I’d note a lack of similar studies since the dawn of true agents last fall Partially that is newness & partially research design challenges, but I suspect we are missing some large effects.
You can now build an agent that: 🔎 Finds real products 🛒 Prepares purchases 💳 Pays …with a person approving every order 👨💻 Ask in @slackhq, review the order, approve in @Stripe's Link agent wallet. Built on MPP + Managed Deep Agents. https://www.langchain.com/blog/agents-that-can-pay-with-stripe-link
Day 4/ We have improved steering to be instant, leading to the model now reacting much faster to adjustments you make, allowing you to course-correct direction in realtime and not have the model waste effort. Also releasing GPT-6.1 Sol ultrafast. The two work very well together.
Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.
At #SFTechWeek? Join LangChain, @tavilyai, and @nebius for an afternoon bringing together enterprise leaders, founders, and engineers building with the agentic layer. Info + RSVP: https://luma.com/tavily-gbt7
Less than 4 years later and the Pope & the world’s greatest mathematicians & the President are posting a lot about the direct successor to this model and what it means for humanity.
ClickUp’s Brain² runs every frontier model with full context of your company, writing and running code to build reports, decks, and dashboards across all teams. That code touches sensitive workspace data, has to start instantly mid-conversation, and has to scale to millions of users. That’s why @ClickUp built its execution layer on E2B: every sandbox is its own microVM, starts from a prebuilt template with tools ready, and scales on demand. Enterprise-ready, from SOC 2 to HIPAA to full isolation.
Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️
At Interrupt, The Agent Conference by LangChain, @OVOEnergy will be sharing how they built a flywheel to improve customer support agents. Catch OVO's talk + more: https://interrupt.langchain.com/london
🚀 Pydantic Validation v2.14 is out, with Python 3.15 support. Stabilized MISSING, lazy imports, TypeForm for TypeAdapter, frozendict, more types validating natively in pydantic-core, and create_model namespace. 👉️ Full details on what changed: https://pydantic.io/tEext
We’re launching the Anthropic Cyber Mission, a new effort to secure critical infrastructure and open-source software. https://www.anthropic.com/news/anthropic-cyber-mission
This assistant remembers where you left your keys. No network connection. No remote server. Build one yourself in a course built in partnership with Qdrant and taught by Dylan Couzon. 👉 Enroll for free: https://hubs.la/Q04zGYrF0 #AI #VectorSearch #AIEngineering
At the very start of my grad school in 2005, I wrote paper about the original computer hacking/phreaking/BBS scene: hackers were often driven by curiosity but the tools they made were widely exploited by "Script Kiddies" who caused most damage & chaos Anyhow, about AI hacking...
R to @AnthropicAI: This is part of our broader effort to make the systems we all rely upon more secure and resilient. That work will take time: the Anthropic Cyber Mission will expand and change as we learn what works, alongside our partners in the public and private sectors.
R to @AnthropicAI: To secure open-source software, we’re launching OSS Scanner. We’ll use our frontier models to periodically scan opted-in open-source projects for vulnerabilities, at no cost. Our reports will provide a proof-of-concept, explanation, and suggested fix. https://www.anthropic.com/research/launching-opt-in-vuln-finding-service-for-open-source
R to @AnthropicAI: To help secure systems like power grids and water systems, we’re launching a new Critical Infrastructure Defense Program. We’ll bring frontier Claude models, on-site engineers, and our latest research to these operators’ security, manufacturing and technology providers.
RT by @elonmusk: Congratulations to our CEO @JensenHuang on receiving the National Medal of Science for advancing GPU computing to power scientific breakthroughs. And to fellow honorees Elon Musk, Lisa Su, Michael Dell, Satya Nadella, and Sergey Brin.
R to @Google: Conducted in partnership with @BIDMC_Medicine, the study demonstrates technology’s potential to strengthen the relationship between doctors and patients. Learn more ↓ https://goo.gle/3VCGhMl
Our medical research system, AMIE, is the first patient-facing conversational diagnostic tool of its kind to be studied prospectively in a real-world clinical setting. In a study, published today in @TheLancet, we found that when patients chatted with AMIE before an in-person appointment, the interaction built their confidence and helped them organize their thoughts. And physicians who reviewed the information from the patient conversation got more time back to focus on collaborative care and shared decision-making instead of digging through data.
R to @claudeai: When a question needs deeper analysis, send the dashboard to your analytics tool. When an animation needs finishing touches, open it in your video editor. See the full list of supported tools: https://claude.com/resources/articles/dashboards-and-motion
R to @claudeai: Starting today, Claude Docs, Slides, and Design are out of beta and available on every plan including Free. Your team and Claude can edit the same doc, deck, or design together.
R to @claudeai: Claude Motion turns reports, charts, or product walkthroughs into short animations. Claude writes each one as code, with no video model involved, so you can change any word, number, or timing, then export an MP4. In beta on Team and Enterprise plans.
R to @claudeai: With Claude Dashboards, connect a data platform or CRM tool and ask a question in plain language. Claude writes the query and builds a dashboard that updates as your data changes, and every chart shows that query. In beta on paid plans.
Congratulations to our Chair and CEO Dr. @LisaSu on receiving the National Medal of Science, recognizing her remarkable contributions to semiconductor technology, high-performance computing and American innovation.
Introducing E2B Secrets 🔐 Give agents API access while keeping credentials outside the sandbox. E2B injects keys into configured HTTPS request headers and lets you rotate them without restarting sandboxes. Available now! https://e2b.dev/resources/introducing-e2b-secrets by @GiulioMicheloni
I think I wrote exactly that, and it is still right. Organizations are a form of narrow superintelligence. My university, or Walmart, can do things that a human cannot, often using complex processes no one designed explicitly. I will go further: the failure to deeply consider how to make AI work together with existing organizational superintelligences is a big reason why we have not yet seen any large impact of AI on scientific discovery or economic productivity despite high levels of AI ability.
R to @e2b: “At the end of the day, any infrastructure behind ClickUp Brain² has to pass the same enterprise security review as the rest of ClickUp. E2B clears that bar at every layer, from SOC 2 and HIPAA requirements down to isolated microVMs where every sandbox has its own kernel.” - Alaa Bouayed @alaabouayed, Senior AI Engineer at ClickUp
R to @e2b: “We see code execution as a fundamental primitive of agentic action in the future, and working with E2B has been absolutely fantastic to bring that to market.” - Jay Hack @mathemagic1an, Head of Artificial Intelligence at ClickUp
R to @e2b: "At the end of the day, any infrastructure behind ClickUp Brain² has to pass the same enterprise security review as the rest of ClickUp. E2B clears that bar at every layer, from SOC 2 and HIPAA requirements down to isolated microVMs where every sandbox has its own kernel." - Alaa Bouayed @AlaaBouayed, Senior AI Engineer, ClickUp
R to @e2b: “We see code execution as a fundamental primitive of agentic action, and working with E2B has been absolutely fantastic to bring that to market.” - Jay Hack @mathemagic1an, Head of Artificial Intelligence, ClickUp
Thank you, @POTUS, for this incredible honor today. I'm humbled to stand alongside so many giants of American innovation and looking forward for what we can continue to do together to advance technology and drive American prosperity.
R to @simonw: Related new project from Microsoft is Quicksand, a Python library that bundles QEMU and uses it to run eg an Alpine Linux container - this one is really neat, I've tried it on macOS and Windows and Linux: https://github.com/microsoft/quicksand
New open source cross-platform (Windows, macOS, Linux) sandboxing library from Microsoft - looks very promising, uses processcontainer/bubblewrap/seatbelt under the hood https://github.com/microsoft/mxc
This document is going to be an assigned reading in college classes that cover this moment in time, there's a lot happening in a few paragraphs... https://www.ahmath.org/statements
R to @Google: In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.
RT by @ylecun: The Medal of Science is for scientists. Scientists publish their works in peer-reviewed venues so they can be scrutinized, verified, and reproduced.
R to @Google: From trending flavors to reviewer go-tos that won’t break the bank, here’s a taste of what topped the list this year ↓ https://goo.gle/4AWK3Aa
What are people craving around the world? Our first-ever @GoogleMaps Fan-Favorite Dining List taps into Maps reviews, ratings, and searches for directions to analyze the biggest dining trends and popular restaurants across 10 global food hot spots: Atlanta, Chicago, London, Melbourne, Montréal, New York, San Francisco, Santiago, Seattle, and Tokyo.
R to @swyx: insane humility even at 60% market share personally i think the claude mods discussed in our @trq212 episode have a lot of potential for malleable/domain specific harnesses we stopped using Tags after finding out they cost ~$3k/mo minimum passive ingestion but happy to revisit when it is for the poors
R to @emollick: The few attempts to measure similar things (METR Long Tasks, GDPval) all became saturated earlier this year as AIs began to do days of work at a high quality level.
RT by @ylecun: Toutes les formes d'IA sont de "belles saloperies" ? Vraiment ? Même celles qui dépistent les tumeurs dans les mammographies ? Même celles qui détectent et évitent les obstacles sur la route et sauvent des vies en réduisant les collisions de 40% ? Même celles qui aident à filtrer les spams et les tentatives d'escroquerie par email, messages, ou réseaux sociaux ? Même celles qui détectent et bloquent les tentatives d'influence étrangères sur le processus démocratique ? Même celles qui permettent aux non-voyant d'entendre une description de leur environnement visuel ? Même celles qui connectent les cultures par la traduction automatique des langues ? Même celles qui assistent dans leurs métiers les médecins, les chercheurs, les journalistes ? Il faut toujours éviter de jeter le bébé avec l'eau du bain.
RT by @ylecun: Au contraire. C'est une nouvelle ère qui s'ouvre pour les mathématiques. Une ère où la démonstration formelle est largement automatisée et où l'accent sera reporté sur le développement de nouveaux concepts, nouvelles abstractions, nouvelles définitions, et nouvelles conjectures. L'invention du bateau a réduit l'importance de la nage, mais a permis la découverte de nouvelles terres.
R to @emollick: I suspect there are going to be a lot of bots in this thread who understand number theory better than their operators (or me!). So expect a lot of ostensible AI marketing accounts to have a lot of opinions on Lean proofs all of a sudden.