We are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!
Everyone has done it: You build a new application with coding agents and love the toy version so much that you wind up deploying it without first evaluating trade-offs like latency, uptime, and compute costs. And we all know vibe coders who do this every day and only discover the production consequences later. When a developer lacks the proper grounding in software engineering fundamentals, they can’t steer an agent to make the structural decisions needed for building great applications. In this week's letter, Andrew Ng outlines why understanding the full software development stack, from data management to user interfaces, remains essential for AI engineering. Here’s Andrew’s list of the key technical capabilities requiring a knowledgeable human’s guidance: 🏗️ Full-Stack Development: Planning out API design, session management, caching strategies, and asynchronous processing. 🗄️ Data Lifecycle Management: Picking data models and storage infrastructure to maintain consistency and clean feeds for downstream AI systems. 📐 System Architecture Design: Choosing everything from the monolith versus microservice dilemma to load balancing. 🔒 Security and Reliability: Building in smart failure-handling policies to ensure graceful degradation, unit and integration testing, and security and testing — early and often. 🚀 Production Operations: Configuring CI/CD pipelines, managing databases, and creating observability tools for understanding workloads. Read Andrew’s full letter to explore how software fundamentals shape effective AI engineering, and how those fundamentals fit into our overall map of AI Engineering Skills: https://hubs.la/Q04vJ_V50
It might soon be disrespectful to use a weaker model for human-facing content: “you saved 6 cents to make me read through error-filled & badly written AI slop? At least send me high quality and low-error slop that doesn’t waste my time.” (This is how I feel about AI X comments)
I've enjoyed working with Cursor pre-acquisition and have respect for the team and what they have built. The 5% here should have come with strong caveat and I would love for Michael to share the math. Tokens are not a proxy for revenue nor value created and the OpenAI models are on the very frontier of token efficiency. Smaller or less strong models require many more tokens to achieve a task and therefore will inflate traffic share significantly.
Are you measuring AI infrastructure performance on how agents actually run? @SemiAnalysis_ built AgentX to do just that. Vera Rubin NVL72 performance measured by NVIDIA on the SemiAnalysis AgentX workload shows up to 30x better throughput per megawatt than GB300 NVL72. Learn more ➡️ https://nvda.ws/4x2WyXX
We unfortunately have decided that we cannot continue providing access to our models through Cursor and are ending our partnership. It boils down to trust and we’ve asked that this takes effect on November 12 to give you some time to plan. Many have used the GPT models through Cursor and here are options we know should work in the future: - We will continue to allow using your own OpenAI API key and similarly will continue to provide access through our IDE extensions for Cursor. - We will keep working with the broadest range of tools and harnesses, some of which are OSS, but also many many closed-source ones. We are as committed as ever to continue supporting developers and the flourishing ecosystem of tools, harnesses and products. We will also continue to invest in our own open-source initiatives and believe in broad optionality for developers. You can read more about our decision in the blog: https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/
We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/
Consensus estimate is that ~15GW of AI compute produced in 2027 cannot be turned on in 2027. This is harder than just finding power, as you also need to build out all the transformers, wiring, liquid-cooling, (massive) chillers & complex networking. https://grok.com/share/bGVnYWN5_4f06d401-90e9-4ed4-bad8-89e3b4fadde2
This paper does a good job of showing the promise and gaps of autonomous AI scientists. Big question is how much more advanced models close those gaps.
Thanks everyone for all the nice feedback on Build a Reasoning Model (From Scratch) so far! I am also flattered that 2 book clubs are discussing it. I look forward to join for live Q&As on Thu, Sep 3, at 10 am (and 2 pm CT). Please join us & bring your questions!
Some early evidence that Google AI Overviews may be doing to Wikipedia the same thing that coding agents did to StackExchange... https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6164926
Was just skimming through the webpages of some of my favorite hard science fiction authors and wow the majority of them hate LLMs: a surprising number of them because they think it is a useless stochastic parrot, some because of IP, and a minority because of existential risk.
Software fundamentals still matter with agents because tradeoffs still exist. Agents will pick a tradeoff. Expertise is knowing the menu - latency, consistency, cost etc - and steering the agent to the one that's right for your system.
RT by @elonmusk: SpaceX and Tesla are each building 100GW/year of solar production capacity as fast as possible, but natural gas will still be needed to supplement and bootstrap solar for several years. The limiting factor for nat gas turbine production is casting the blades & vanes. By doing in-house casting at SpaceX, we can accelerate nat gas turbines coming online by up to 18 months, which is a profound game-changer.
What used to take years now takes a week. @TAMU's VISION, an NVIDIA DGX SuperPOD and the most powerful academic supercomputer on the June 2026 TOP500, is giving researchers across seven institutions the compute to tackle challenges at unprecedented scale and speed.