Lightning fast to customize. Lightning fast to run. NVIDIA Nemotron 3.5 Lightning is a compact, customizable open model built to help always-on agents complete specialized tasks faster. Kari Briski joins @MTSlive to explain how Lightning helps always-on agents work faster.
New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures
Why @Box chose Deep Agents: 1️⃣ Complete model agnosticism Customers can choose LLM providers, and Deep Agents allows this flexibility at the platform level.…
Why @Box chose Deep Agents: 1️⃣ Complete model agnosticism Customers can choose LLM providers, and Deep Agents allows this flexibility at the platform level. 2️⃣ Speed of iteration Open agent harness = less time spent on building core agent infrastructure, more time solving enterprise-specific problems. Inside the Box Agent ⤵️ https://www.langchain.com/blog/building-box-ai-how-an-enterprise-content-platform-went-ai-native-with-deep-agents
10 million downloads for NVIDIA Warp 🎉 Warp started with a simple idea: you shouldn’t have to leave Python to get real GPU performance for physics and simulation. Since then, developers have used it to accelerate work across physics simulation, computational engineering, geometry processing and robotics. Thank you to everyone who downloaded it, broke it, filed an issue or sent a PR. On to the next 10M!
該期內容還列出 Andrew Ng 談 AI engineering 的全端技能、GLM-5.3 的開放權重資安能力、OpenAI/Google/Nvidia 改善即時互動吞吐、DeepSeek-V4-Pro 的開源 harness,以及 Self-GC 用 LLM 修剪長上下文。貼文是週報導讀,未提供每項技術的原始數據或獨立驗證。
Without strong software engineering fundamentals, coding agents often default to bad trade-offs that hurt system latency, reliability, and cost. This week in The Batch: ▪️ Andrew Ng on full-stack skills for AI engineering ▪️ GLM-5.3 brings advanced cybersecurity capabilities to open weights ▪️ OpenAI, Google, & Nvidia speed up throughput for real-time interaction ▪️ DeepSeek-V4-Pro ships with an open source harness ▪️ Self-GC uses an LLM to better prune long contexts Read the full details here: https://hubs.la/Q04vJC_w0 📱
.@PodiumHQ's agents seemed broken. LangSmith showed the real story: the agent was behaving rationally based on the context it had. Principal Software Engineer Walker Ward on tracing agent reasoning end to end.
Figuring out the origins of cosmic rays (e.g. supernova, black hole...) from ground-based detector arrays is a notoriously difficult inverse problem. This post walks through how astroparticle physicists are using Keras to replace the usual hand-crafted features and directly model the raw spatio-temporal waveforms https://developers.googleblog.com/decoding-cosmic-signals-with-deep-learning-and-keras/
GLM-5.3 is a good model, and as the open weights models get better and better it becomes increasingly important that they actually publish model cards,…
GLM-5.3 is a good model, and as the open weights models get better and better it becomes increasingly important that they actually publish model cards, do red teaming, etc. Since you can break the guardrails with any open model, we need a sense of what the risks are as well.
How have software engineering fundamentals changed with agentic coding? Here is our AI Engineering Skills map for software engineering fundamentals. https://x.com/i/article/2093384274372419585
Already running an inference engine? So where does NVIDIA Dynamo fit in? In five minutes, we break down how Dynamo sits around engines like @sgl_project, @vllm_project and TensorRT-LLM to scale inference across GPUs and nodes. Full video in the comments 🔽
Google AI 彙整本週推出的多項 Gemini 相關更新:Gemini 3.5 Transcribe 被稱為其最精準的語音轉文字模型,Gemini Omni 1.1 Flash 擴充影片生成與編輯控制,Gemini App 的 Live 體驗加入 Daily Brief、Gemini Spark、Personal Intelligence 與 Gmail 收件匣管理等任務功能。Google 也提到新的跨 Google 計畫 Expert Intelligence,讓使用者能與可信來源互動並結合洞察,起點是 Gemini Notebook 中符合資格的 Google Play 電子書。這些都是官方發布說法,證據中未附基準測試、支援語言範圍或實際可用地區。
一龍馬判讀
Google 正把 Gemini 從聊天與單點生成推向語音、影片、個人工作流與內容來源整合;使用者與企業要留意功能可用性、資料授權與 Gmail 等個人資料接入後的隱私邊界。
原文節錄
Google AI · @GoogleAI
Here’s what launched this week: — Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed to deliver intelligent transcriptions — Gemini Omni…
Here’s what launched this week: — Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed to deliver intelligent transcriptions — Gemini Omni 1.1 Flash, bringing expanded creative capabilities and controls for video generation and editing — @GeminiApp’s Live experience, moving beyond conversation to complex tasks with new features like Daily Brief, Gemini Spark, Personal Intelligence, and @Gmail inbox management — Expert Intelligence, a new cross-@Google initiative that lets you engage with and combine insights from trusted sources, starting with eligible @GooglePlay ebooks in @Gemini_Notebook
In 13 minutes, @jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs a month. Topics covered: ✅ Their four quadrant eval framework ✅ Why closing the production-to-eval loop is the hardest part ✅ How a data lake and long context changed what agents can do with data
Two weeks before launch, @unifygtm's agent was burning so much compute that one message could instantly eat through a customer budget. Co-Founder & CTO @HeggieConnor on how they cut costs by 90-95%
Agents need their own identity to do real work on the web. We partnered with @tryramp to automate our event logistics workflow. Our agent logs into sites with a Browserbase Context, downloads receipts, and submits them via the Ramp CLI.
Aside from waiting for economic impact, we are now at the place where AI video, image & music models are good enough to be real tools for letting more people who never could produce new kinds of art How long until we see a creativity boom among the slop flood? Is it happening?
On September 17th, the LangChain Academy team is hosting a live workshop on Deep Agents. RSVP today and get ready to learn how to build your own Deep Agent and see how planning, memory, and subagents let an agent run complex, multi-step tasks. https://events.langchain.com/webinar/intro-to-deep-agents/
Open development helps AI move faster. When we can learn from, build on and improve each other’s work, we can make better technology together. Hear @ctnzr on the MAD Podcast with @mattturck
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use.…
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use. Here’s how you can try it in @FlowbyGoogle and more → https://goo.gle/4zHEzJ2
We sat down with the Composio team to ask them one question: "How do YOU use Composio?" Watch as they break down the workflows they run every day, helping them do their work more efficiently.
R to @LangChain: Watch or listen to the latest Max Agency on your favorite podcasting platform. 🎧 Apple: https://podcasts.apple.com/us/podcast/how-unify-cut-its-ai-agent-costs-95-in-two-weeks/id1891551672?i=1000783151404 🎧 Spotify: https://open.spotify.com/episode/6kWQouc2QmiHGk0vdiZEtd?si=ba6e241ca4a24faf ⏯️ YouTube: https://youtu.be/6898VdRtKDE
R to @AnthropicAI: Claude can reliably fix measurable misalignment. But subtle or rare failures may have no benchmark at all—so everything hinges on measuring the right things. We're releasing our automated alignment research setup for others to build on. Full report: https://alignment.anthropic.com/2026/automated-alignment-researchers/
R to @AnthropicAI: Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training.
R to @AnthropicAI: Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities. Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger.
R to @AnthropicAI: Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve…
R to @AnthropicAI: Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities. We then tested its best methods on held-out benchmarks to see if they'd generalize.
In verifiable domains, model capability scaling should remain unbounded. Models will simply keep improving by "absorbing more and more of the computational universe", which is infinite by construction.
.@Airbnb is joining us at Interrupt NYC. Pedro Rodriguez will share how Airbnb's Trust org went from prototype to a standardized production stack on LangChain and LangGraph, in a domain where being wrong is expensive. Catch his talk + more @ Interrupt NYC, Sept 24. https://interrupt.langchain.com/nyc
You can use Composio to connect GrokBot to thousands of apps. We just found this recent @nateherk’s video on 9 GrokBot hacks, and were pleasantly surprised to see Composio make the list: https://youtu.be/TMPUUyQC5aM?si=2enGlVXvefGMpawh&t=323 Thanks for the love, Nate! <3
R to @fchollet: To note, this isn't intelligence. This is skill. Superhuman skill. Of course it will *feel* like intelligence to anyone who equates intelligence with skill, which is probably almost everyone. Intelligence in my definition is (and has always been) the efficiency with which you extract and operationalize the patterns you need to achieve a given level of skill. It's basically the ratio between your resources and what you can do with them, an information conversion ratio. It is not tied to any capability threshold. You can always achieve arbitrarily high skill with arbitrarily low intelligence, given arbitrarily high resources. Humans will be left behind capability-wise in all verifiable domains, but remain many orders of magnitude more intelligent than current AI -- you didn't need 1,000,000x the code volume of all of GitHub in order to learn to code.
R to @fchollet: One way to think of current AI is as a big sponge for patterns. It absorbs and operationalizes any pattern it is exposed to (albeit with very low data efficiency at training time). Once you can programmatically enumerate the complete space of patterns in a domain, saturating the domain becomes purely a matter of computational resources.