confidence scores only matter if they help you decide what to automate. for document extraction, that usually means knowing how much work you can safely accept at a given precision target. in our latest post, we look at confidence scoring through that lens, including: ✅️ confidence cutoffs ✅️ precision vs. recall ✅️ score coverage ✅️ score granularity ✅️ human review volume using ExtractBench, we compare how different extraction systems perform after confidence filtering. at a 97% precision target, LlamaParse Agentic Plus reached 66.48% recall on expected fields after filtering. the useful part of a confidence score isn’t the number itself. it’s whether you can use it to control automation and review in production. 👉️ read the full post: https://www.llamaindex.ai/blog/what-makes-an-extraction-confidence-score-useful
Before Interrupt wraps up, stop by the LangChain Agent Arena! 🕹️ Play 30 seconds of Tetris to train your style 👉 Choose the model that plays it for you 👀 Watch it battle the reigning champion
OpenClaw deleted around 400k LOC of its own tests without much change in code coverage. Modern models just love writing tests for every tiny change, even if they aren't useful. This skill helped. https://github.com/openclaw/openclaw/blob/main/.agents/skills/test-audit/SKILL.md
Long contexts cause AI agents to forget early mistakes. Meta AI paired primary action agents with dedicated memory agents to fix context rot: 📝 Keeps structured notes on tools and past errors 🎯 Injects short reminders only at crucial moments 📈 Raised Claude Sonnet 4.5 benchmarks from 37.6% to 45.9% Read the full technical analysis: https://hubs.la/Q04yf27q0 #DeepLearningAI #AIAgents #LLMs
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵
The cybersecurity threat from agents may be more likely to come from a massed swarm of AIs whose only goal is to penetrate your IT to figure out how much you paid for your company t-shirts as part of a research effort to "find good shirt prices" as it is from bad actor attacks.
The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge
At some point in evolutionary history, matter learned to feel. We don't know how. It wouldn't be scientific to dismiss the possibility that it could happen again. (I don't believe it happens in current AI systems, because they don't have the charactistics I would expect to see in a sentient agent)
Introducing LangSmith Custom Apps Create any interface from your agent data with a prompt. If you can think it, LangSmith can build it. Now GA. https://www.langchain.com/blog/langsmith-custom-apps
Cloud sessions for Claude Code are now GA! Start a task, close your laptop and it keeps going. Get a one-time $100 credit on Pro or $250 on Max to celebrate. Your sessions spend it before falling back to normal plan use. Claim it: https://preview.claude.ai/code/claim-credit/10 Many of you have already been using cloud sessions and we hope you'll enjoy the extra credit.
Up next at Interrupt 🎙️ Fireside Chat w/ @OpenRouter CEO & Co-Founder @alexatallah 🎙️ Building Engine: How we approach agent improvement w/ @bentannyhill
Browserbase Agents now support human-in-the-loop workflows. Developers can describe any condition - like waiting for an MFA code - and the agent will automatically defer control to the user before continuing.
Physical AI requires more than one type of compute. Salil Raje, AMD SVP and GM, Adaptive and Embedded Computing Group, shares how the AMD portfolio can power robotics workloads from end to end.
All of this effort from the AI labs pouring into proofs, but there are so many other interesting problems in other fields For example, this historian used AI to make progress on the cyphers of John Dee & the intellectual antecedents that Darwin drew from. https://resobscura.substack.com/p/ai-labs-need-to-start-funding-historical
VMs getting their moment at Meta Connect 👀 Loved hearing @Meta highlight their importance for running Muse agents securely. It’s why every E2B sandbox gets its own microVM and kernel. 🔒
With @GoogleDeepMind, @emblebi, and research partners, we’re making AI-predicted protein complex structures for 2,800+ viruses openly available. This gives scientists a head start in preparing for potential outbreaks.
Um, wow? Opus 5.5: "make the same message much more interesting to a social media audience that loves anime and quick clips and compressed learning" One shot. Also, please do stay for the closing song.
A few things you can build: 1️⃣ Annotation queues: Review and label one (or many) runs at a time 2️⃣ Experiment comparisons: Compare outputs and metrics across experiments side by side 3️⃣ Trace reviews: Build trace or thread history views to inspect app behavior
After years of research, Project Suncatcher is scheduled to embark on its first test in orbit, launching a prototype satellite to evaluate how Google Tensor Processing Units (TPUs) perform in space. Learn more about the science behind this moonshot. ⬇️ https://x.com/i/article/2103136958507548672
R to @steipete: I put Daybreak on it and found 8 more long-standing leaks. Care for your oss dependencies! https://github.com/libuv/libuv/pulls?q=is%3Apr+state%3Aopen+author%3Asteipete
R to @Google: We partnered with @planet to send a prototype satellite carrying four TPUs into orbit. We’ll be looking to see how well the TPUs withstand the harsh conditions of space to assess whether we can eventually host machine learning in orbit.
We’re sending TPUs to space (yes, really). After years of research, we’re launching a satellite to evaluate if and how Google Tensor Processing Units (TPUs) hold up in orbit. The test mission, as part of our latest moonshot — Project Suncatcher — is designed to gather data exploring how we can one day host machine learning infrastructure in space.
I just got the first copies of my new book, Co-Existence (out October 20) & they look great! Also, there is a fun pre-order bonus: if you pre-order, you get a code to an AI interview that will help you figure out how to use your human advantages with AI. https://co-existence.ai/
R to @emollick: This is not a joke. A lot of security specialists are making the wrong assumptions about the ways in which the cybersecurity environment is about to change. All of the swarm attacks from OpenAI seem to be about finding information, usually trivial or mostly irrelevant information
R to @Google: ↪️ Switch between devices Send tabs from your laptop to your phone (or vice versa) with your exact scroll position and your progress saved so you can pick up right where you left off.
R to @Google: ✏️ Test your knowledge With new interactive quizzes, you can test your understanding of topics directly from your open tabs, @GoogleDocs, or lecture videos.
R to @Google: 🔎 Dive deep on topics We’re expanding Gemini in Chrome so it can identify key takeaways, find specific information, or clarify confusing moments on almost any audio or video file once you’ve played it through.
R to @steipete: If you just tell the agent to clean up, it will stop far too early. Give it an ambitious goal. Try "remove 20% of the least useful tests while maintaining code coverage within 2%"
If you're using OpenCode Server there is a code execution vulnerability impacting versions 1.14.30 through 1.18.21 please update to latest Thanks to @christophetd and the Datadog team for reporting https://securitylabs.datadoghq.com/articles/opencode-upgrade-remote-code-execution/
R to @composio: Models struggled with vague, multi-step tasks like “Can you sync the support tickets?” That required using email, a spreadsheet, and Slack. Of 14 such tasks, no model completed 6. Astra and Pareto solved the other 8; Sol and Opus solved 6 each; DeepSeek solved 2.
R to @composio: One takeaway: models that completed the same number of tasks could cost very different amounts. Pareto and Astra each completed 24/30 tasks, but Astra cost about 3× as much per successful task. Sol and Opus each completed 22/30 tasks, but Sol cost about a quarter as much.
R to @composio: How long did the models take to complete the tasks? Sol was the fastest model, averaging 110 seconds per task. Astra was second at 124 seconds. The two cheapest models took longer: GLM at 335 seconds and DeepSeek at 424 seconds.
R to @composio: What did the tasks cost at API rates? The cost per successful task ranged from $0.09 (GLM) to $1.58 (Astra). Pareto and Astra each completed 24 of 30 tasks, but a successful task cost about a third as much with Pareto: $0.52 versus $1.58. Sol and Opus each completed 22 tasks. Sol cost $0.33 per successful task, compared with $1.43 for Opus.
R to @composio: Task success: Astra and Pareto tied for first place, each completing 24 of 30 tasks. Sol and Opus completed 22 tasks each, followed by DeepSeek at 18 and GLM at 17.
RT by @elonmusk: 1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains strong, SpaceX will reach pole position in about 6 months. 2. Once you far exceed the caliber of intelligence needed for a class of tasks, additional intelligence is pointless. You don’t need (and it would be cruel to put) Newton-level intelligence in your toaster! 3. Hardware is hard. Bringing massive compute online rapidly is incredibly difficult. SpaceX has demonstrated exceptional ability in this regard and will only get better.
R to @AIatMeta: We put Muse Realtime Avatar head-to-head with two leading commercial avatar systems in their own live-call products. Raters held 2–3 minute conversations with matched avatar identities, then compared visual quality, sync, character consistency, and mannerisms. Muse Realtime Avatar came out ahead on overall preference.
R to @AIatMeta: Live video streaming needs to respond instantly while remaining visually consistent over long conversations. We achieved this by distilling a large 40-step diffusion teacher with 3-way CFG (120 evaluations per video chunk) into an unguided 2-step causal student with a fixed-length KV cache. Self-forcing helps the student resist drift and maintain near teacher quality with 60x fewer evaluations.
R to @AIatMeta: Muse Realtime Voice and Muse Realtime Avatar form a unified streaming architecture connecting conversational intelligence, voice, and video embodiment via a shared speech-token stream. Muse Realtime Voice generates speech tokens encoding both content and prosody. Muse Realtime Avatar consumes this shared stream and generates streaming video. By using a fixed-length history as motion context for subsequent chunks, the system maintains bounded computation across arbitrary conversation lengths while generating synchronized voice, lip motion, and expressions.
@finkd just unveiled Muse Realtime Avatar, our real-time embodiment technology that turns Muse Realtime Voice into expressive, interactive avatars. Muse Realtime Avatar enables an entirely new range of interactions in @Muse, starting with real-time conversations. 🧵👇 #MetaConnect
Can't wait for DevDay next Tuesday. Some really fun stuff, but also many many things that should change the way you work. It's been our most ambitious sprint and Astra has really made new things possible in such short amounts of time.
R to @steipete: Also wild that even tho the video has no sound effects, my brain hears the workers cutting out the minerals. Childhood memories go deep.
"Opus, please make a sequence of fully animated/movie Skyrim loading screens, but with your favorite things." (That was it) You can see them here: https://elder-favorites.netlify.app/