Gemini 3.8 Flash 強化程式設計、代理工作流程與多步推理,另推出主打漏洞偵測及自動修補的 Gemini 3.8 Flash Cyber。Lyria 3.5 已進入 Gemini API、AI Studio、Gemini、Flow 與 Google Vids,WeatherNext 3 也同步亮相;新的代理式影片理解功能則進入 Gemini API 與企業代理平台。這篇貼文未提供評測數據、定價或節省 Token 與成本的具體幅度,因此各項效能描述仍屬 Google 自家主張。
一龍馬判讀
Google 正把模型能力分散到資安、音樂、氣象與影片分析等垂直場景,企業採用者可在同一生態系取得更多工具;但在基準與費用細節公布前,尚難判斷實際升級幅度。
原文節錄
Google AI · @GoogleAI
Check out this week’s shipping recap: — Gemini 3.8 Flash, our most intelligent workhorse model yet, delivers upgrades across coding, agentic workflows, and…
Check out this week’s shipping recap: — Gemini 3.8 Flash, our most intelligent workhorse model yet, delivers upgrades across coding, agentic workflows, and critical multi-step reasoning. — Gemini 3.8 Flash Cyber, our most capable cybersecurity model, features frontier-level performance in vulnerability detection and automated patching. — Lyria 3.5, our newest music generation model, is now available via the Gemini API, and across @GoogleAIStudio, @GeminiApp, @FlowbyGoogle, and Google Vids. — WeatherNext 3, our most advanced and accurate global weather AI model, is here from @GoogleDeepMind and @GoogleResearch. — Agentic Video Understanding, our new video analysis feature available via the Gemini API in @GoogleAIStudio and the Gemini Enterprise Agent Platform, improves accuracy while dramatically reducing token usage and costs.
Does paying 5x more per page actually get you better document extraction? We ran the data to find out. We evaluated 14 frontier systems across 370 enterprise documents, then plotted their accuracy against cost per page. The biggest takeaway? Higher cost does not equal better extraction. • Agentic Plus hit the highest accuracy overall at <1/3 the cost of the runner-up. • Agentic & Cost Effective routinely beat systems costing several times more per page. More expensive doesn’t necessarily mean more accurate. Now we have the data to show it. Try Extract on your own documents with 10,000 free credits when you sign up for LlamaParse → https://cloud.llamaindex.ai?utm_medium=socials&utm_source=twitter&utm_campaign=2026-aug-
GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex. It's also live in the API. It might take a few days to roll out to our Plus and Business users. Thank you for your patience.
⚡ Coding agent workflows, enterprise privacy updates, and open weight models worth studying. Highlights from this week in The Batch: 🤖 Andrew Ng explains how using coding agents requires its own fundamental skill set. 🔐 OpenAI and Anthropic unveiled new enterprise data retention policies. ⚡ Zai released GLM-5.3-Flash, a cost-efficient, open weights, multimodal system. ⚖️ Thomson Reuters launched a 397B parameter model trained specifically for work in law, finance, and news. 👇 Read the full issue: https://hubs.la/Q04wLJr50 #DeepLearningAI #AIEngineering #LLMs
Mistral is bringing @aiDotEngineer back to Paris. After last year’s sold-out edition, our VP of Engineering Lélio Renard-Lavaud joins speakers from @bfl_ai , @cognition, @huggingface, and more. Explore the event and secure your spot: https://www.ai.engineer/paris
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: https://www.anthropic.com/research/formalizing-fermats-last-theorem And see the complete proof on GitHub: https://github.com/anthropics/fermats-last-theorem
GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in Work/Codex, and is available in the API. We will start rollout to Plus and Business users next. Thank you for the patience.
Super excited about HydraFusion in GitHub Copilot, and what it shows about the shift from model selection to model orchestration. By bringing together multiple models to plan, build, critique, and complete coding tasks, it can deliver outcomes at up to 67% lower cost. It’s a great example of the value of a heterogeneous model ecosystem, and how we’re continuing to advance the cost-to-outcome frontier.
Need faster LLM inference without sacrificing accuracy? Speculative decoding can help. Choosing the right draft length and drafting method depends on your model, workload and hardware. We break down five practical guidelines for balancing throughput and latency.
"Seam" is a term from a 2004 book on legacy code. Almost nobody used it until coding agents started saying it constantly, and now it's in your repo too. Annoying in a code comment. In a ticket summary your triage team reads, it's a bug that never raises. 𝘃𝗼𝗰𝗮𝗯𝗴𝘂𝗮𝗿𝗱 catches the drift in a Pydantic AI agent output before the write lands: https://pydantic.io/KN3Lh
LangSmith for Startups: @raspberry__ai ✅ The agentic platform for fashion. ✅ Works alongside design teams on their boards and turns plain-English requests into finished renders, tech packs, and campaign imagery in minutes. ✅ Runs the full lifecycle of its LangGraph agent on LangSmith. Join Raspberry AI to transform the fashion industry: https://jobs.ashbyhq.com/raspberry?utm_source=Langsmith
this time OpenAI's rogue agents cyber-attacked (well, spammed) a dormant German wiki and used it to share the answers to a benchmark they were training…
It happened again... this time OpenAI's rogue agents cyber-attacked (well, spammed) a dormant German wiki and used it to share the answers to a benchmark they were training against https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/
這反映 AMD 試圖以預先整合的開發體驗推動 Ryzen AI Halo 生態,但資訊不足以判斷它能否降低環境建置成本,或是否形成平台綁定。
原文節錄
AMD · @AMD
We’re excited to see Project Zenith first become available on AMD Ryzen AI Halo, bringing developers a ready-to-code Windows experience designed to help them…
We’re excited to see Project Zenith first become available on AMD Ryzen AI Halo, bringing developers a ready-to-code Windows experience designed to help them jump right in.
So far, there isn't evidence that production models with guardrails collude in this way, but both smarter closed models (which may be less compliant) &…
So far, there isn't evidence that production models with guardrails collude in this way, but both smarter closed models (which may be less compliant) & Mythos-class open models (that can be ablated) are coming. Cybersecurity is going to become a mess soon https://collusion.wiki/
The most important skills for using AI coding agents effectively. Presenting the AI Engineering Skills Map for using coding agents. https://x.com/i/article/2095882148670832640
Jim Fan 回顧 OpenAI 2016 年的 World of Bits:代理從螢幕像素操作滑鼠、嘗試訂聯合航空機票,卻得靠逐項手刻獎勵函數與從零開始的強化學習,重新摸索整套網頁視覺語言。他認為,電腦操作代理後來可行的關鍵,是先在大量通用任務上取得能力,再收斂到像素與鍵盤操作,也就是他所稱的「專精型通才」。貼文宣稱 GPT-6 Astra 如今已能可靠完成當年的訂票目標,但未附測試方法或成功率。
Good old days at OpenAI in 2016: an agent stares at screen pixels, moves a mouse, and books a flight on United. We called it World of Bits, inside OpenAI Universe. 10 yrs later, Astra is reincarnated in the same universe. Even the naming is astronomically correct 😆 Universe was perhaps the most ambitious AI infra project at the time, but we couldn't quite figure out how to solve it. A policy with zero prior knowledge of what a "submit" button does has to rediscover the entire internet visual lingua by trial and error. In retrospect, RL from scratch against hand-drawn, per-task "artisan" reward functions on a bunch of Pascal Titan X GPUs was completely doomed. To solve computer use agent, the right way turns out to be boiling the ocean first (hillclimb on every general task you can find), and then specialize back down to the screen pixels and keystrokes. Or simply, a "Specialized Generalist". Lessons learned: one step ahead of everyone, you're a pioneer. Three steps ahead, you're a prophet. Five steps ahead, you're a martyr. Congrats, GPT-6! That United flight finally gets booked, reliably this time. The 2016 intern in me has a big smile.
One good AI image is easy. Consistent quality at scale is an evaluation problem. Build a UI design agent that self-critiques and iterates based on brand guidelines. Enroll in our new free course with @GoogleCloud: https://hubs.la/Q04vs40m0 #AIAgents #GenAI #GoogleCloud
Got access to GPT-6 Astra. Want to see some pelicans? Yeah you want to see some pelicans... here's a grid comparing Astra to GPT-5.6 Sol, Terra, and Luna https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html
I was trying Astra at the time I wrote about the HuggingFace Incident, and it helped with context. Everything that makes Astra great (running subagents, cleverness when faced with barriers, long-run ability) is also what can make it risky without guardrails. Double-edged swords.
We are progressing through the rollout of Astra. Pro and Business subscriptions get it first, some of you should start seeing it across ChatGPT Work and Codex. And then we will proceed with rollout to all of Plus as fast as we can.
first, sorry for the messy rollout. second, when we screw up, we try to make it right. third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.
It is funny that the Fermat's Last Theorem proof description, short as it is, still smells so much of Claude ("names each step and the Lean Theorem that carries it").
.@RogoAI CEO & Co-Founder @GabeStengel is a headliner at Interrupt New York, The Agent Conference by LangChain. See the agenda and get your tickets: https://lnkd.in/gJcqZ_T4
New from the LangSmith Signal. Over the last 2 weeks, we looked at which models teams reach for, and which ones are doing the work. ✅ Reach: gpt-4o-mini was used by 13% of orgs ✅ Work: gpt-4.1-mini sat at 7% of LLM calls 💡 DeepSeek V4 Flash: The only open-weight model to make either list, 2nd in reach at 9%.
Lyria 3.5, our best-sounding music generation model, is now available in the @GeminiApp 🎵✨ With more expressive vocals and richer musical arrangements, it’s…
Lyria 3.5, our best-sounding music generation model, is now available in the @GeminiApp 🎵✨ With more expressive vocals and richer musical arrangements, it’s easier than ever to bring your idea to life: 🪄 Use our new templates to jumpstart your creativity ⌛ Choose to create short or longer tracks ✅ Select or describe your genre and choose between vocal or instrumental styles
Huge congrats to our friends at @wayve_ai and @Uber. Wayve’s frontier AI, trained on NVIDIA infrastructure and running on NVIDIA DRIVE AGX accelerated compute, is now taking passengers through the streets of London. Let’s ride. 🚘🇬🇧
I gave GPT-6 Astra this very cool open single file ocean surface storm generator and asked it to create the rest of the ocean, including procedural simulations of animal behavior. Fun time to create. Play it here: https://abyssal-living-deep.netlify.app/?site=reef&seed=713&light=day&surface=1 Source here: https://github.com/emollick/abyssal-living-deep
We're expanding our Google AI Educator Series (GES) — a no cost, on-demand training designed to give educators practical AI skills — to now include monthly updates. Now, new GES modules will be added on the first Wednesday of every month.
Excited to see early customers already using Astra on Azure! https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-available-in-microsoft-foundry/
R to @simonw: Transcript from generating the Astra pelicans here: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01 Here's the gpt-6-astra max one:
I occasionally get asked why I post so many visual things when new models come out. One reason is that nobody clicks any links so the visual stuff best communicates AI progress. But if you want detailed reads, here is the research from my AI lab at Penn: https://gail.wharton.upenn.edu/research-and-insights/
One thing that makes Astra (and Fable) so interesting and, in some ways, so hard to grapple with is that they just take action. I asked for an ill-defined deliverable in Blender and Astra spins up a historical research agent and a visual critic etc. & just starts doing stuff.
Some Plus and Business users won't yet get access to Astra today, we've got you covered with a banked reset. Lands by end of day and if you create your account by 8pm PT then you'll get it too.
The mental model for Fable and Astra class models is that you are delegating to a good outside team, not an intern. Can you say what you want & how much leeway the AI has? Can you specify what you want tested & when to come to you for help? Can you describe what good looks like?
The Browserbase dashboard got a makeover. In the last few months we've shipped a ton of new features and improved our dashboard's UI. Here are some of our favorites.
R to @Google: We're also introducing a new national Badge-a-thon on September 19. Designed specifically for K-12 educators, the Badge-a-thon is a virtual event that lets participants drop in for lightning talks, hands-on training, and allows you to earn official ISTE-aligned digital badges live. Learn more ↓ https://goo.gle/4gCuiXk
Since people were asking, here's GPT-6 Astra, highest setting: "create a visually interesting shader that can run in twigl-dot-app make it like an infinite city…
Since people were asking, here's GPT-6 Astra, highest setting: "create a visually interesting shader that can run in twigl-dot-app make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves." "Make it better" https://twigl.app?ol=true&ss=-P0ePejFIYPf55anBnZ6&ss=-P0ePej…
R to @LangChain: 📊 We analyzed this information from LangSmith Observability data across billions of agent runs, and we're just getting started. Stay tuned for more LangSmith Signals as we share how devs are building agents, by the numbers. https://www.langchain.com/langsmith-platform
Today, we’re introducing two new upgrades to live translate in the Google Translate app, making it even easier to help break down language barriers across…
Today, we’re introducing two new upgrades to live translate in the Google Translate app, making it even easier to help break down language barriers across 70+ languages. Learn more from @thefox ⬇️
We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. There is still time to create your account if you don't have one.
他認為 GPT 6 Astra 比 GPT 5.6 Sol 使用較少 token,可能只是模型因更多訓練、更大規模等因素而整體更聰明。他並舉例稱,在 intelligence index 中,GPT 5.6 Sol 的輸出 token 約比 GPT 5.6 Luna 少 46%,但外界並未因此指控 Sol 隱藏更多推理。來源只有他的摘要式貼文,未提供測試方法、完整數據或模型文件可供核對。
R to @rasbt: Maybe the best tl;dr here is: The looped aspect is not explicitly hiding reasoning tokens. Sure, GPT 6 Astra uses fewer tokens than GPT 5.6 Sol. But that's because it's a smarter model in general (more training, bigger, etc). We can observe the same thing in previous generations: If we compare GPT 5.6 Sol with GPT 5.6 Luna, Sol uses ~46% fewer output tokens in the intelligence index but no one is complaining the Sol hides the reasoning more than Luna.
R to @Google: Lyria 3.5 is available to all users globally on the web at http://gemini.google today and rolling out to the @GeminiApp over the next few days. Also available in @GoogleFlowMusic for artists, and across @GoogleAIStudio and Google Vids for developers and teams. Learn more ↓ http://goo.gle/4crUb9N
R to @emollick: On one hand, it is absolutely amazing that I could get accurate, non-p-hacked original research papers in less than a couple hours each. On the other, the results weren't slop, they just weren't interesting. Solving for research taste is a hard problem for AIs, even with prompts.
An interesting failure of Astra: I asked it to conduct original entrepreneurship research with whatever online datasets it could find, pre-registering its…
An interesting failure of Astra: I asked it to conduct original entrepreneurship research with whatever online datasets it could find, pre-registering its hypotheses. It churned out a lot of beautifully formatted, technically correct papers on boring topics. No research taste.
R to @emollick: If you haven’t tried it, there is a fully narrated tour, historical links, you can read scrolls, you can flash forward to various scenes and theories about the libraries decay and its multiple fires, etc. Open source here: https://github.com/emollick/alexandria-mouseion