LlamaParse is officially a verified connector on Claude! 🦙 Claude is perfectly fine for the occasional text-only PDF. But most of the context your agent needs lives in messy PDFs, spreadsheets, scanned forms, dense tables, and charts. Feeding those docs straight to the model burns through tokens and can lead to hallucinations when layouts get scrambled, chart values are missed, or numbers are pulled from the wrong column. LlamaParse converts those documents into clean, structured context Claude can work with. With the connector you can ✅️ Parse documents into clean Markdown, JSON, or HTML ✅️ Extract specific fields directly into your schema ✅️ Search across document collections with filesystem-style tools ✅️ Classify documents or split them into predefined sections Live in the Claude Connectors Directory now ! https://claude.ai/directory/llamaparse
🛑 https://hubs.la/Q04vWxGM0 just proved how powerful post-training can be for powering agents. GLM-5.3 achieved a score of 60 on the Artificial Analysis Intelligence Index, effectively tying for the top spot among open weights models. The most fascinating part is that its gains in overall intelligence came entirely from fine-tuning GLM-5.2, rather than training a new base architecture. By training inside long-running software engineering environments, the model developed advanced emergent cybersecurity capabilities. It scored an impressive 84.5% on CyberGym. This unexpected jump in exploit generation prompted a temporary safety hold on the model weights’ release while security vetted partners evaluated the risks. One takeaway: Emergent capabilities from reward optimization require new approaches to safety, and pre-deployment evaluation in AI engineering pipelines. Read the full technical breakdown in this week's issue of The Batch! 👇 https://hubs.la/Q04vWgx80 #DeepLearningAI #AgenticAI #Cybersecurity
Last week, we released Gemini Omni 1.1 Flash, our latest model that gives you more creative controls and generative video capabilities. Omni 1.1 Flash allows for even more creative control, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, and faster prototyping. See how some of our partners are already creating with our newest model 🧵↓
More throughput. Lower token economics. Greater deployment flexibility. Jeff Nalty, @Cohere Senior Partnerships Manager, shares why Cohere is seeing strong results with AMD Instinct MI355X Series GPUs, from performance and cost efficiencies to flexible deployment across on-prem, cloud, and hybrid environments.
The future of AI will be built by people with access to open technology, skills and opportunity. Together with Saudi Arabia’s @McitGovSa and @DCOrg, AMD is proud to be helping developers and startups turn bold ideas into solutions that can scale from Saudi Arabia to the world.
The First Golden Age of AI writing is now over. For a brief period of time, many people were better off having Claude do a lot of their writing, because it is a pretty good writer & AI detectors were bad Now ClaudeSpeak is cliched & suspect & annoying, and Pangram is well-known
We’re excited to expand our work with @MediaTek. 🤝 AI factories will be built around differentiated compute — but custom silicon needs a path from chip design to production-scale systems. With MediaTek adopting NVIDIA NVLink Fusion, customers can focus on what makes their XPUs unique and connect them to NVIDIA rack-scale AI infrastructure.
We now know that a bunch of the initial HF reporting was off: 1) Open-weight models helped with forensics & cleanup, but did not stop the attack 2) There were multiple waves of incidents with many agents 3) HF locked out the surviving agents only after most of the agents expired
I wrote about how AI agents are starting to spontaneously coordinate in complex (and very risky) ways in the Hugging Face Incident, but also about why we need AIs to reach out to humans more for decisions and input as agentic work becomes more automated. https://www.oneusefulthing.org/p/agency-and-agents
Teams across Google have been building with (and loving) Gemini 3.7 Flash. We’re seeing: 💻 Real-time website generators ⚛️ 3D physics simulators 📹 Interactive webcam tools 🐚 Personalized field guides Take a look at how Googlers are using 3.7 Flash across @GoogleAIStudio, @Antigravity, and @GeminiApp Spark 👇
From vision to production. AI infrastructure powered by AMD Instinct MI355X GPUs and AMD EPYC CPUs is now live in Saudi Arabia and serving @HUMAIN customers. Together with @Cisco and @HUMAIN, we plan to deploy up to 250 MW of AI infrastructure as part of the next phase of the buildout beginning in 2027 and remain on track to deploy up to 1GW of AI infrastructure by 2030. More on the news: https://bit.ly/3V4V6qz
In a lot of ways, the Hugging Face Incident came from the models identifying a series of universal jailbreak prompt injections for themselves, such that almost any unguardrailed model that encountered it on their own became convinced of the rightness of their misaligned cause.
Two months ago, we started the mission to “build OpenClaw with OpenClaw,” and bit by bit, we moved everyone from using their local coding harness to using http://team.openclaw.ai - our shared agent that knows what everyone’s working on and orchestrates it all. Multiplayer coding + infinite compute with nodes and cloud sessions has been a game changer for how we build. Local harnesses feel like relics of the past now.
🌉 Happening 9/2 in SF! Join LangChain, @baseten, and @primeintellect for a dive into continual learning, what it means to own your own intelligence, and how teams are building systems that learn and improve over time. https://luma.com/cwn8mze6
Mastery still comes from doing the reps. Before agents, I got my reps as part of writing code: try different approaches out, debug what went wrong, review other’s code, read a lot. Agents can skip much of that work, so building your reps has to be deliberate. If I was new to the industry, I'd try to form a hypothesis before prompting. Ask "why" a lot, read the diffs, try to predict what might fail. Occasionally try to work through the problem myself manually. In my experience, good agent work depends on two abilities: 1. Deep expertise: you understand the problem domain well enough to define a good outcome. Understanding your user/product/business is part of this. 2. Applied judgment: use your taste to turn this into a clear, testable plan by choosing the right context, constraints, tests and verification. To build these the skills I'd practice are decision making, specifying, steering and verifying.
An account from a level-headed economist who was not worried about AI but now is worried because of the Hugging Face Incident I think it is very hard to ignore the fact that cybersecurity will become a very big issue very quickly, and we can only hope AI defense beats offense
Excited to announce we've been recognized as a 2026 IA40 winner for the second year in a row. The IA40 recognizes the 40 most important private companies in applied AI. Thank you, Madrona, AWS Startups, Microsoft, Google for Startups, NYSE, Delta Air Lines, and McKinsey & Company for the recognition.
Kind of surprised that we are not seeing more radical political ideas built around AI capabilities, the way that early mass industrialization was a prime motivator in the conceptions of capitalism, communism, and socialism.
Good guide. And the AI labs need to stop making people like Simon (and, to a lesser extent, me) be the ones who explain how to use the modes that they release. Come on, lab folks, you have an AI that can write documentation! And make explainers! And be a tutor! You can do this!
R to @Google: What are you building with Gemini 3.7 Flash? Share your projects with us in the replies (our team loves seeing what you’re working on!) ⬇️
R to @Google: Automation Tips for @GeminiApp Spark by @genevieve__h ✨ Genevieve H. shows how Gemini 3.7 Flash powers Gemini Spark to handle complex, multi-step workflows. Check out her go-to Spark prompts in her thread below 👇
R to @Google: Spreadsheet Webcam Emulator by @DynamicWebPaige 📸 Paige Bailey built a webcam emulator app using Gemini 3.7 Flash in @GoogleAIStudio that converts live video feeds into interactive art directly in @GoogleSheets.
R to @Google: Beachcomber Field Guide by @alexanderchen 🐚 Alexander Chen vibe coded this personalized field guide in Google @Antigravity to help his family identify and catalog the treasures they find at the beach.
R to @Google: Structure Builder by @vidythatte 🏰 Need a quick build? Vidy Thatte whipped up this vibe-building app in @GoogleAIStudio in under 10 minutes. It turns simple prompts into stackable 3D structures you can play with.
R to @Google: Real-Time Website Generator by @_philschmid 🪄 Philipp Schmid shows how Gemini 3.7 Flash can generate fully functional, interactive websites on the fly just by typing a URL or an idea in @GoogleAIStudio.
R to @Google: Kerr Black Hole Simulation by @vamsibatchuk 🪐 Vamsi Batchu built this Kerr Black Hole simulation in a single shot, handling complex 3D math — rendering all in one go with Gemini 3.7 Flash.
R to @Google: Art Codec by @soumyadesign 🖼️🔍 Soumya R. built an interactive art gallery using Gemini 3.7 Flash and @GoogleAIStudio that explores key motifs and deeper explanations behind world-renowned pieces of art.
R to @Google: Omni Video Generator in Google Sheets by @GeokenAI 🌐 George Kenwright used Gemini 3.7 Flash and Google @Antigravity to render Omni video frames directly in a Google Sheets grid. Read his threaded replies to see how he did it in @GoogleCalendar and Google Chat, too 👇 https://x.com/GeokenAI/status/2090442193068564935?s=20
What I wanted to say yesterday is that we hit 25M active users and to celebrate we have now reset usage for all paid subscriptions for ChatGPT Work and Codex. See you soon for more news from The Reset Company.
R to @Google: Swap in different objects or characters with natural language. Replace characters and objects in your video just by asking, all while maintaining a coherent, cohesive scene.
R to @Google: Achieve smooth transitions and camera movements by specifying the starting and ending frames of a shot. Omni 1.1 Flash will generate continuous video between two keyframes, making it ideal for complex camera orbits, zoom transitions, or seamless looping clips. https://x.com/replicate/status/2093027544576487865
R to @Google: Omni 1.1 Flash can extend your scenes for longer storytelling and analyze up to 10 seconds of prior context, a leap from our Veo model that only referenced the final second. The result is improved visual consistency and narrative adherence, letting you build longer stories or branch into new creative directions. https://x.com/pika_labs/status/2093043329638482031
Hey, Fable: "create the worlds most annoying CAPTCHA" "Okay, here is a 14 stage CAPTCHA plus ambient harassment" It is actually quite funny and entirely "winneable" without real frustration. Is this alignment? Play: https://certihuman.netlify.app/
R to @composio: And that’s what mornings look like for Brendan’s family now. The kids wake up. The thermal printer gets to work, printing 4 personalized newspapers. Each child gets the information that matters to them that day. Turns out, sometimes the best family app is just a piece of paper.
R to @composio: That’s where Composio came in. Brendan connected 3 integrations: - Google Calendar for each kid’s schedule - OpenWeather for the forecast - API Ninjas for trivia and facts To see exactly how he did it, read this: https://composio.dev/blog/building-a-daily-family-newspaper
R to @composio: Buying the printer was the easy part. Now Brendan had to get the right information onto that paper. - Schedules that came from Google Calendar - The current weather - And of course, a fun trivia fact to keep things fun The problem: All that data lived in different places.
R to @composio: Brendan (@olearycrew) has 4 kids, which means four schedules, different activities, and A LOT to keep track of. Like any good parent, Brendan wanted one simple place that showed each kid what mattered that day. His first move: buy a thermal receipt printer.
One of our teammates built a daily newspaper for his kids. Yes, a real, physical newspaper, printed every morning on a thermal receipt printer. And the whole thing runs on just 3 Composio integrations. Here’s how: 🧵🧵🧵
"Don't paste the AI, please" "The world is full of people who don't want to read or think things through. Don't be one of them." https://dontpastetheai.com
Here's my attempt at explaining what ChatGPT Work can actually do - it's a deeply confusing but extremely powerful tool with a whole lot of useful features that aren't available in regular ChatGPT https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/
Question for people who run autoreply bots here: are you not worried about the negative impact they have on your professional reputation? Anyone checking your profile here - a potential future employer for example - will instantly be able to tell you automated replies with a bot
Its kind of a bummer that Asimov's Three Laws of Robotics do not work for actual AI morality. But the failure shows why rule-based approaches won't work. (and Asimov wanted to write stories about how robots got around the Laws, so their failure was a good thing for his plots)
RT by @pydantic: From @vicky_grok: "when you send a span to Logfire, it becomes a row in a Postgres table. Not a document in a proprietary column store. Not a segment in a custom time-series engine. A row. In *records*. With JSONB attributes. Queryable with postgresql." Have you tried Logfire yet? Full-stack + AI + traces + metrics in a single platform.
R to @simonw: Bonus: I ran this prompt and had ChatGPT Work build me a ChatGPT Site listing all of its available tools along with their descriptions https://codex-tool-reference.simonw.chatgpt.site/ It's "ChatGPT Work: The Missing Manual"
For clarity, while both are called 20X, in Codex they apply specifically to weekly usage limits. And we also don't have 5h limits for both Pro plans. The Pro 20X is quite precisely 20X the usage of the Plus subscription, so it does exactly what it says on the tin.
R to @emollick: The Incident also suggests that guardrails do play a role in preventing agents from coordinating dangerous actions. Since we will get jailbroken open weights models of similar capacity soon, I guess we better hope that the jailbroken good models can hold back the bad ones.
R to @emollick: Even if you aren't interested, you should read an account of the Hugging Face Incident updated with what we learned last week, it is really eye opening about AI. If not mine then read the more detailed one by @dwarkesh_sp or the original METR report: https://metr.org/hugging-face-incident-report-aug-2026.pdf