Building reliable AI out of unpredictable components requires a new playbook: continuous iteration and disciplined eval loops. To help developers bridge the gap from quick demo to production, @AndrewYNg mapped out Pillar 1: Building and deploying AI Applications, of the AI Engineering Skills Map: 👇🧵👇 🧠 LLM Foundations: Understand model mechanics to predict failures and select the right architecture. 📊 Grounding Models with Data: Architect reliable context through clean data pipelines and retrieval structures. 🤖 Building Agentic Systems: Design the agent harness, including tool integrations, context memory, and production guardrails. 🧪 Evaluation-Driven Development: Build tailored evaluation loops to drive systematic, measurable progress. ⚙️ Operating in Production: Maintain reliability using real-time observability, security defenses, and statistical evaluation. 📈 Machine Learning Foundations: Use core deep learning principles to evaluate model trade-offs and engineer better data. Read the full technical breakdown of Pillar 1 here: https://hubs.la/Q04vhX3N0 #AIEngineering #MachineLearning #LLM
OpenWorker -- an open source agent that doesn't just chat but completes tasks on your laptop -- just released a new version with many features for security workflows. After our initial release, many users found it especially useful for cybersecurity. Attackers are already using AI; OpenWorker is committed to giving defenders the same leverage. Running an agent requires both (i) A model and (ii) A harness (the software around the model). Because the OpenWorker harness is fully open source, security teams can audit it to make sure we haven't built any backdoors that exfiltrate your code and data to some company or even a foreign adversary. OpenWorker now comes with built-in cybersecurity agents for (i) Scanning your code for vulnerabilities. (ii) Scanning dependencies for supply chain injections. (iii) Checking your cloud security configuration for attack surfaces. This enables developers to do much more security work before deployment (part of what's called the "shift left" movement). You choose the model: you can run open weight models fully locally so sensitive code never leaves your machine. This helps with legitimate security work (like reproducing a known exploit to defend against it) that can trigger refusals in leading closed models. Or use your ChatGPT subscription, or stealth preview models like Ox Alpha, or any model via API key. Thanks also to all the open source contributors! Join work with @rohitcprasad so please follow him too to get more frequent updates. Try it out: https://openworker.com/ Code: https://github.com/andrewyng/openworker
Since announcing Jalapeño, our first custom inference chip, we’ve been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without sacrificing efficiency.
Introducing ChatGPT Business Premium Seats The new $100 Premium seat is a game changer for small businesses and startups—giving lean teams better tools, faster…
Introducing ChatGPT Business Premium Seats The new $100 Premium seat is a game changer for small businesses and startups—giving lean teams better tools, faster workflows, and capabilities once reserved for big companies. A flexible plan that scales with your team’s ambition. https://chatgpt.com/pricing/?type=team
Meet Portable Computer, Perplexity's new local-first agent stack on NVIDIA DGX Spark. When running locally, Portable Computer offers one-click local inference setup and an optimized agentic experience for DGX Spark. Learn more and get started today: https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/#perplexity-spark
Every extraction API demos well on a clean invoice. But what about the scanned form, the nested table, the 40-page financial report with merged headers? We tested 14 frontier systems to find out. ExtractBench evaluates schema-guided extraction across 370 enterprise documents, 67 document types, and 4,800+ pages. Join Simon Suo, CTO and co-founder of LlamaIndex, for a technical walkthrough of the results: what makes extraction hard, where existing benchmarks miss it, and how cost trades off against accuracy across VLMs, coding agents, and specialized APIs. Wednesday, Aug 26 · 9:00 AM PT / 12:00 PM ET Save your seat 👉 https://lnkd.in/gMbFNHPg
Excited about our Jalapeno results today. An incredible achievement from the team, taking a new chip from concept all the way to very impressive performance on real workloads in the lab. Tomorrow’s fast will feel like today’s ultrafast. As I’ve mentioned before, we’re pushing to bring this to as many people as possible. We’ve seen a TON of demand for /ultrafast, made possible by our deep partnership with Cerebras and their unique hardware, which will push the frontiers of tomorrow's ultrafast even further. I’m very much looking forward for the collaboration to continue pushing the absolute limits of how fast we can run our most capable models on Cerebras and bring this to the most demanding customers.
So much demand for this one. Works similar to the Pro $100 plan but designed for teams and small companies. ✅ All ChatGPT, ChatGPT Work, and Codex features ✅ Connect to Google Workspace, Slack, GitHub, Microsoft 365, and more ✅ Secure workspace with SAML, SSO, and MFA ✅ Centralized billing and administration ✅ Usage analytics and spend controls ✅ No 5h limits
When an LLM engine crashes, a cold restart can mean minutes of lost capacity. Shadow engine recovery, a new preview feature in NVIDIA Dynamo, keeps a standby engine warmed up and ready to take over. In our GLM-5.2 test, it restored capacity in 7.3 seconds, nearly 39x faster than a cold restart. Read all about it: https://nvda.ws/45OhcA0
Agent 從 demo 走向 production 後,問題通常不只在模型輸出,而是提示、工具、資料與程式碼交互造成;這類工具若成熟,可把除錯從人工追 log 推向半自動化。
原文節錄
LangChain · @LangChain
Accelerate every step of the agent development lifecycle: 🔎 Detect production issues 💡 Identify root causes 🔧 Propose fixes to prompts and code 🔄 Monitor…
Accelerate every step of the agent development lifecycle: 🔎 Detect production issues 💡 Identify root causes 🔧 Propose fixes to prompts and code 🔄 Monitor for recurrences
LangSmith Engine now offers >2x performance on key internal benchmarks. Engine has already helped identify tens of thousands of issues in our customers’ agents. Now it offers: ✅ More accurate issue detection, clustering, and remediation ✅ Support for SaaS and self-hosted deployments ✅ Reduced Analysis mode for cost-sensitive customers ✅ Integrations with @slackhq and @Linear ✅ Automatic closing of stale issues Learn more → https://www.langchain.com/blog/new-in-langsmith-engine-2x-better-issue-detection
Search gives your agents read access to the internet, extending their context with the most up‑to‑date information. Now Search is built into our dashboard, so you can sample queries before taking them to production.
.@Vtrivedy10 + @nickhollon10 on why you should separate deciding what a task should look like from building it, and how to let a coding agent turn what it learns into a reusable "world spec" for every future task. A closer look ⤵️
Claude now has one memory across chat and Claude Cowork, and you decide what's in it. Hand Cowork a task and it starts from what Claude already knows from your chats: the project you talked through, your manager's preferences, or the client from last quarter.
What does it take to move customer experience agents from an initial use case into a production system that improves over time? Our latest guide brings together lessons from @Lyft, Fastweb + Vodafone, and @LATAMAirlines. Download our guide to explore the full architectures, results, and lessons from each team: http://www.langchain.com/resources/customer-experience-cx-agents-in-production
NVIDIA Vera Rubin NVL72 production racks are here. The compute tray is engineered for fast compute, assembly, and serviceability. Manufacturing is 100% automated, and every tray goes together in one minute. Congratulations to @Microsoft on the first operational Vera Rubin NVL72 racks, now rolling off @HonHai_Foxconn Ingrasys lines.
Join us next week for The Learning Loop in SF on September 2nd: https://luma.com/cwn8mze6 Hear from: ✅ @jakebroekhuizen- LangChain ✅ @willcb- @PrimeIntellect ✅…
Join us next week for The Learning Loop in SF on September 2nd: https://luma.com/cwn8mze6 Hear from: ✅ @jakebroekhuizen- LangChain ✅ @willcb- @PrimeIntellect ✅ @oneill_c- @baseten We’ll dive into continual learning, what it means to own your own intelligence, and how teams are building systems that learn and improve over time. After the talks, we’ll head to the patio where you can meet the LangSmith Engine team, enjoy food and drinks, slot car racing, an AI photo booth, and more. Spots are limited, so be sure to register!
R to @OpenAI: We plan to begin deploying Jalapeño in OpenAI’s compute infrastructure by year-end. It’s the first step in a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will push efficiency and speed further. https://openai.com/index/jalapeno-first-results/
R to @claudeai: Topics some consider sensitive, like health or religious beliefs, stay out of memory unless you turn them on in Settings. Memory is on by default on Free, Pro, and Max plans. Review yours anytime in Settings > Memory. Read more: https://claude.com/blog/claudes-memory-works-everywhere-and-you-decide-whats-in-it
R to @claudeai: Everything Claude remembers is saved as a list of topics in Settings, where you can read, edit, or delete each one. Memory also updates on its own as you chat, saving new details as they come up. You can also say "remember this" to save something specific.
Tomorrow we will bring back the 5h limit for Plus accounts across ChatGPT Work and Codex. I had mentioned this a while ago, but then postponed it. This is necessary as (a) the 5h limit allows us to smoothen the load on our compute, allowing to keep the plan generous in terms of weekly usage and (b) users on the Plus plan are relatively casual and new users, but then also just accidentally eat through their whole weeks usage and then are confused, making it not a great experience. We are for the upcoming months keeping the 5h limit not enabled for Pro $100 and Pro $200 subscriptions.
The impact of computing is shaped by more than technology. It also depends on the people who can learn, research and innovate with the latest solutions.
Tibo 發文稱「OpenAI DevDay 2026 will be our best DevDay in the history of the company. It will not be close.」語氣強烈,但貼文沒有說明議程、產品發布、API 更新或講者資訊。若 Tibo 與 OpenAI 內部或開發者生態有關,這仍只是個人式預告,證據不足以確認 DevDay 2026 的實際內容。
一龍馬判讀
開發者大會通常會影響 API 使用者、工具鏈廠商與新創規劃時程;但目前只有宣傳性判斷,團隊不應據此調整技術路線或商業承諾。
原文節錄
Tibo · @thsottiaux
OpenAI DevDay 2026 will be our best DevDay in the history of the company.…
RT by @elonmusk: Megatons to orbit & beyond is a critical piece of the K2 puzzle. The other parts are: massive amounts of solar cell production, AI chip production, millions of humanoid (general purpose) robots and a mass driver on the Moon. My guess is that ~1M Optimi plus ~1GW of solar will constitute the first Von Neumann Probe, a system capable of replicating itself entirely from local materials on any sufficiently large planet or moon. This would enable expansion of civilization throughout the galaxy and reaching K3.
Starbase Louisiana will ultimately have over a dozen launch towers, enabling more than 30 Starship flights per day and making it the biggest launch site…
Starbase Louisiana will ultimately have over a dozen launch towers, enabling more than 30 Starship flights per day and making it the biggest launch site on Earth! SpaceX makes sci-fi real.