Qualitatively, there is now a larger gap between open & closed models than there has been in awhile. Fable/Astra class models are agentic in a way pre-Fable models are not, for better or worse. None of the open models have crossed that line, yet. When they do, it will be a jump.
From per-core strength to enterprise and cloud-native performance. Across enterprise and cloud-native testing, AMD EPYC 9996 delivers 2.4x to 3.7x performance gains. See the full performance story: https://bit.ly/4AkQ7SN
A little Ember-1 tl;dr. Seems like a great model! (Was recently asked on a podcast, given $ xx million, what's the best way to develop a frontier LLM today? My recommendation was: start with an existing one and spend that budget on post-training. Great example here.)
Engineer-to-engineer collabs are the best 🙌 Congrats to @AIatMeta on Muse Realtime Avatar! We worked with their team to make the model more efficient so the avatar can keep up when you talk to it.
Glad Claude limits feel different! The lower Opus 5.5 price goes straight into your limits. They go about 25% further than Opus 5. Cache reads cost 60% less and output is generated 30% faster: https://claude.dev/blog/what-a-task-costs-on-opus-5-5/
There is always talk on X about how networks like BlueSky are full of people who believe AI doesn't work (this is less true than it used to be), but LinkedIn is full of insane semi-technical slop advice like this. Why are they not reporting BLEU, ROUGE & BERTScore for Astra? 🤔
AI is not just changing what CPUs need to do. It’s changing how they are designed. AMD Senior VP, Corporate Fellow and Chief Architect of AMD CPUs Mike Clark shares how engineers are using AI to explore more design possibilities, accelerate verification and narrow down options faster. Learn how engineers can spend less time on repetitive tasks and more time applying their expertise: https://bit.ly/4h7Jig1
Same. @useblacksmith has been an amazing sponsor but we need to distribute the load. My plan is to let codex decide which tests actually need to run and drastically nix CI and run tests hourly.
Model personality matters in a way that benchmarks can't capture. Opus models went through a rough patch from around 4.7 to 5 where they just didn't feel "Claude-y" anymore, more like an watered-down Fable (hmmm, teacher models?). Opus 5.5 feels like working with ol' Claude again
I included a few references to this year's record-breaking Kākāpō breeding season in a talk I gave yesterday, and since Claude Opus 5.5 is surprisingly capable at pixel art animation I had it create this celebratory video for my closing slide
It is strange how much LLMs turned out to be the key to such a wide range of problems that would not, initially, seem to be problems that a model of human language would be able to solve.
Code freeze isn’t really a thing anymore before releases and in the future the code might even be generated online per request according to some constraints.
R to @emollick: I can't believe people think SimHat is made up. Would I also make up the gritty reboot in the early 2000s? And the online multiplayer version ("its a brimmunity!")? And the recent exploitative mobile game?
R to @emollick: Anyhow, it is almost like some sort of general intelligence, but made with artificial means. Very hard to know what we might call that.
R to @simonw: Prompt and transcripts here https://simonwillison.net/2026/Sep/26/kakapo-party/ - or you can load the HTML page Claude built to see the interactive version before it became a video https://tools.simonwillison.net/kakapo-party
R to @emollick: Also please feel free to fight in the comments about "pure LLMs" versus multimodal LLMs versus multimodal LLMs with tool use or whatever. I totally get all the caveats, but you get what I mean.