"your edge right now is understanding the possibilities and being curious" Using AI for software engineering is now the default. The engineers I've seen who have the edge are ambitious and have been grinding with agents for years on the job at this point. It's a compounding effect. For teams still unsure or catching up, the smartest thing I can advise is to be even more ambitious and aggressively share what you learn with your peers. If you do all this work privately, only you benefit. If your team only gets to experiment with AI after hours, you've quietly told them using it doesn't count. Senior engineers should be the first to make this shift. Their decades of taste in system design and failure modes is the secret sauce that makes agent output good. They should direct their experience to building with agents better safety. Then show whats working to others to replicate.
How AI shifted from a transductive paradigm (token-by-token answer completion) to an inductive paradigm (on-the-fly reasoning chain synthesis), in one chart. This is the big story of 2025 and 2026. This is what made coding agents work.
🖥️ Agents don't just consume tokens for input and output like humans do with early chatbots. Agents use computers: searching, storing, analyzing, transacting, at a scale no person matches. In this week's Andrew's Letter in The Batch, Andrew Ng explains why that means we need much more infrastructure, and faster tools, not just faster models. 🏗️ 📬 https://hubs.ly/Q04zR8n20 #DeepLearningAI #AIEngineering #AIAgents
Don't forget: Claude Max & Team now include monthly API credits. Max 5x: $100, Max 20x: $200 Great for building apps & agents via the API, Managed Agents or Agent SDK. We hope you'll enjoy Opus 5.5 there too.
Reasoning From Scratch: Reinforcement Learning with Verifiable Rewards (RLVR) round 2. Covering clipped policy ratios, KL loss term, format rewards, and other GRPO tips & tricks. 00:00 Introduction and recap 01:52 Interpreting basic GRPO training metrics 06:34 Planned improvements to GRPO 08:58 Running longer training jobs with Python scripts 13:39 Running the baseline GRPO training script 17:29 Loading and plotting training logs 19:29 Diagnosing unstable training 23:55 Evaluating checkpoints on MATH-500 26:26 Downloading existing checkpoints 30:09 Tracking advantage statistics 34:53 Understanding entropy 40:32 Computing entropy in PyTorch 44:17 Interpreting entropy values 48:58 Adding entropy tracking to GRPO 53:36 Analyzing advantage and entropy metrics 56:18 Stabilizing GRPO with clipped policy ratios 1:03:27 Implementing the clipped policy loss 1:09:39 Analyzing clipped policy training results 1:11:25 KL divergence and reward hacking 1:15:12 Adding a KL loss term 1:20:34 Limitations of the simplified KL loss 1:23:04 Format rewards and think tags 1:25:47 Adding special tokens to the tokenizer 1:30:29 Implementing the format reward 1:35:56 Analyzing format reward training 1:38:25 Rewarding format only for correct answers 1:40:48 Further GRPO improvements from research 1:45:43 Next steps and distillation
They show up for us. Who shows up for them? Meet Wenyi, Founder of Kuddo Health. Her experience with therapy sparked a mission to support clinicians, with help from NVIDIA GPU–accelerated AI. Watch her story. #WorldMentalHealthDay
Pydantic AI now turns on prompt caching with one capability. Anthropic and Bedrock cache nothing unless the request asks. The Caching capability asks for you: tool definitions, instructions and the conversation, one setting instead of each provider's. https://pydantic.io/uJC8d
To buy stuff, you can just take a picture of your credit card, drop it in the chat and Grok @Bot will scour the Internet for the best deal on the item you want and order it
The ridiculous arguments by Ambani’s army of Internet shills make no sense On the one hand, they say Starlink couldn’t possibly compete with the amazing low prices offered by Ambani’s de facto monopoly in India and then, in the same breath, they essentially claim that Starlink would be too strong of a competitor! This is absurd hypocrisy and hyperbole. The reality is that competition would greatly benefit the average citizen in India by giving them more choices, as it always does. Ambani’s profits will be lower, but he is not going out of business. And, obviously, if Starlink is too expensive in India, no one will buy it, so SpaceX will need to make pricing affordable.
Not only did Astra beat Montezuma's Revenge, but so did Opus 5 But one of the criteria to resolve this AGI bet is for AI to win a discontinued prize that was sort of like a weak Turing Test. So Metaculus has decided to recreate the test to confirm the criteria is resolved. Neat!
I speak to a lot of people about AI & not everyone gets how it can be useful for them outside of work So its a shift that when I talk to people using Muse or Dot, they often tell a story of how their AI did something proactively to help them, and it makes them feel good about AI
aie nyc tix will sell out this weekend - last call to join us to the biggest ever technical conf in New York and our first with a finance mainstage - in some ways the final unification of both my careers: from investment bank to hedgefund, from bigtech to nyc startup. cya monday!