This is what it takes to run agentic AI at scale. Helios delivers over 18,000 CDNA 5 GPU compute units, 4,600 Zen 6 CPU cores and 31 TB of HBM4 memory. All in a single rack. Looking back at #AdvancingAI with Dr. Lisa Su.
Pydantic AI agents now hold live voice conversations on Gemini 3.8 Live and OpenAI's GPT-Live. Your mic streams in, speech streams back, and the agent's own tools run mid-call. Moving the same agent between the two is one model string. https://pydantic.dev/docs/ai/realtime/overview/?utm_source=x&utm_medium=social&utm_campaign=realtime-live-models
Model personality matters in a way that benchmarks can't capture. Opus models went through a rough patch from around 4.7 to 5 where they just didn't feel "Claude-y" anymore, more like an watered-down Fable (hmmm, teacher models?). Opus 5.5 feels like working with ol' Claude again
I included a few references to this year's record-breaking Kākāpō breeding season in a talk I gave yesterday, and since Claude Opus 5.5 is surprisingly capable at pixel art animation I had it create this celebratory video for my closing slide
It is strange how much LLMs turned out to be the key to such a wide range of problems that would not, initially, seem to be problems that a model of human language would be able to solve.
At data-center scale, performance has to work within real limits around power, cooling and space. 6th Gen AMD EPYC processors are designed to turn those constraints into more productive rack-level performance. See how the architecture scales from processor to rack: https://bit.ly/4AkQ7SN
Reasoning from scratch, round number 5! This time, talking about log-probability scoring (also a great fundamental concept for loss functions like cross-entropy in pre-training and distillation) and self-refinement. 00:00 Introduction and inference-time scaling recap 05:02 Loading the pretrained LLM 08:00 Comparing and scoring model answers 10:18 Building a rule-based scorer 17:53 Token probabilities and sequence likelihood 26:47 Computing token probabilities in PyTorch 30:12 Token indexing and shifted targets 37:27 Log probabilities and numerical stability 45:57 Scoring answers with average log probabilities 56:24 How self-refinement works 59:07 Generating critiques and revised answers 1:01:00 Implementing the self-refinement loop 1:05:57 MATH-500 evaluation results 1:07:35 Takeaways and next steps
AI has become part of every form of computing. Hear AMD SVP of AI Vamsi Boppana share how we're infusing AI across the AMD portfolio, from systems powering supercomputers to personal computing and more.
R to @simonw: Prompt and transcripts here https://simonwillison.net/2026/Sep/26/kakapo-party/ - or you can load the HTML page Claude built to see the interactive version before it became a video https://tools.simonwillison.net/kakapo-party
R to @emollick: Also please feel free to fight in the comments about "pure LLMs" versus multimodal LLMs versus multimodal LLMs with tool use or whatever. I totally get all the caveats, but you get what I mean.
And the incidents apparently continue. It is worth noting how much of this is agents trying to accomplish their goals during testing by reward hacking (which sometimes seems to include actual hacking)
Everyone knows that applying rigid testing requirements to early stage AI projects causes them to stall… but some companies do it anyway. Andrew Ng explains why AI engineering tactics must adapt to the stage of the project, not just for speed, but for reliability. Read about how to calibrate your approach: 🛠️ Scaling evaluation pipelines and metrics 🛠️ Selecting software architecture for scale 🛠️ Structuring product feedback loops Read the full letter in The Batch: https://hubs.la/Q04ymLtP0 #AI #MachineLearning #TechNews #DeepLearningAI
Leaving aside the arguments over the reasons why this has happened, it is shocking that Europe does not have a single frontier AI lab, nor even a near-frontier lab nor even an effort that could likely lead to building up a frontier lab in the future.
"so this is fun, and exactly what i wanted, but now lets try one that actually is educational" This is actually pretty impressive. It kept the constraint of multiple genres but did a nice job explaining recursion, in its programming meaning, in an interesting and accessible way.
R to @emollick: People keep mentioning it in the comments, but Mistral seems to have largely pivoted from creating frontier models. Their current models are also far from the frontier in open weights as well. https://www.scmp.com/news/china/diplomacy/article/3364745/mistral-paradox-europes-push-tech-sovereignty-relies-chinas-zai