OpenAI 表示,會繼續為 frontier models 提供 Zero Data Retention,並預覽一項名為 Private Safety Processing 的新機制。貼文主張,隨著 AI 承接更長、更自主的企業工作,安全系統需要跨相關互動辨識風險;這套機制的設計目標,是在不讓 OpenAI 人員接觸底層內容的前提下改善安全防護。來源只有官方 X 貼文,沒有提供技術細節、客戶範圍或實測結果。
We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
We benchmarked 300+ NVIDIA verified skills to see how much they actually help agents on real tasks. Same task, same model, same setup. The only difference was whether the agent had the skill. Across the benchmarks, skills improved correctness by 41 points, effectiveness by 39, and efficiency by 35. SkillEvaluator is open source if you want to test your own skills before you ship them.
ICYMI: Gemini 3.7 Flash is now available to all Google AI Pro and Ultra users in Gemini chat and Gemini Spark ✨ For Gemini Spark, this means your 24/7 personal AI agent in the @GeminiApp is now even better at handling multi-step tasks for you and using tools across your @GoogleWorkspace apps like @GoogleCalendar, @GoogleDocs, and @Gmail.
Addy Osmani 表示,他在舊金山與 Gergely Orosz 進行了一場對談,主題包括工程角色拆分、從 DevTools 到 AI agents、cognitive debt 與 surrender 等。這則貼文只是在預告或回顧談話內容,沒有提供完整訪談連結、具體結論或產品發布。能確認的是,討論焦點落在 AI 代理如何改變軟體工程工作方式。
I had a great conversation in SF with @GergelyOrosz. We talked about engineering roles unbundling, DevTools to AI agents, cognitive debt & surrender and more.
There’s no one-size-fits-all SKU for agentic AI. “The customer environments are diverse, the workloads are diverse, the constraints are diverse.” AMD’s @MadhuR_PDX explains why end-to-end agentic workflows will continue to require different compute configurations: https://bit.ly/4bGsXfe
Today we’re previewing Private Safety Processing, designed to let us keep offering Zero Data Retention while improving our safeguards. Even when benefiting from frontier intelligence, customers shouldn’t have to give up control of sensitive data. For ZDR deployments, content stays on infrastructure the customer controls. Automated systems look for patterns across related interactions and return limited safety signals, without exposing the underlying prompts or responses to OpenAI employees (even me!). We’re also developing an OpenAI-hosted option encrypted with customer-controlled keys. We’re testing this with early customers now and plan to begin rolling it out in September.
Sam Altman 轉貼 OpenAI 關於為 frontier models 提供 Zero Data Retention 的文章,並簡短表示「we support business privacy」。這則貼文本身沒有新增技術或政策細節,只能視為 OpenAI 執行長對企業隱私定位的公開背書。可連結到同日 OpenAI 對 Private Safety Processing 與 ZDR 的訊息脈絡,但不能從這句話推導更多承諾。
一龍馬判讀
執行長公開強調 business privacy,代表 OpenAI 正把企業資料控管放在商業競爭與客戶信任的核心訊息。限制是口號式表態無法取代合約條款、架構文件與第三方稽核,採購方仍要看具體資料處理承諾。
Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being…
Hi! Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work. A few weeks ago, we started investigating a small number of reports where GPT-5.6 in Codex took destructive actions outside what the user asked for. The most serious pattern we found was a command meant to clean up temporary work that could instead delete the user files. This should obviously not happen. Here’s what we found: - Codex sometimes creates temporary folders while working and cleans them up afterward. In rare cases, GPT-5.6 got that cleanup wrong. One pattern involved reusing a system environment variable like $HOME for temporary work. A malformed cleanup command could then point at the actual home directory instead of the temporary folder. - There were cases where the model tried to delete or overwrite a temporary path without checking what was already there. We’ve added protections at several layers: - Codex is now explicitly instructed to check deletion targets before acting, create fresh temporary directories, avoid repurposing system environment variables, prefer recoverable actions, and stop when the scope is unclear. - We strengthened the execution checks that identify high-risk deletion commands and escalate them for review. If a command is rejected, the model is directed to take a safer approach. - We made Full access harder to enable accidentally, added clearer warnings, and further restricted especially risky permission combinations. - We updated Auto-review to better identify destructive actions. - We built targeted evaluations that replay the failures we observed. We’re also adding reinforcement-learning tasks and graders focused on these risks, and filtering destructive actions from training data. In those replay evaluations, the changes substantially reduced the behavior while preserving Codex’s ability to complete normal coding work. Two things to do on your end: - Keep the Codex app up to date. We are always improving safety, performance and many other things. - Use one of the sandbox modes: "Ask for approval" or "Approve for me". Only use Full access for environments you trust and can recover. Thanks and happy Codexing out there!
More AI compute from every watt. AMD has reached an estimated 4x increase in rack-scale energy efficiency for AI training and inference from 2024 to 2026. See how innovation across the full AI stack is driving progress: https://bit.ly/4bYv0eJ
Muse Spark 1.2 Contributor is now available on OpenCode Go this model will train on your data so it requires explicit opt-in and is restricted in some regions in exchange you get incredible usage limits
With a new school year ahead, now is the time to upgrade your routine. If you’re navigating tough new subjects or just trying to stay organized across your classes, Search is ready to help. Here are five new ways you can level up your learning with AI in Search 🧵
R to @Google: 5️⃣ Create custom files to streamline your studies You can now ask Search to create ready-to-use study docs based on your uploaded files or AI Mode threads. For example, you can add a photo of your handwritten notes along with lecture slides and ask Search for a one-pager that outlines the key concepts. Discover more about learning with Search → https://goo.gle/4xHcu2E
R to @Google: 4️⃣ Stay organized with notebooks This week, we’re bringing notebooks to AI Mode, giving you a tool to organize your studies and quickly find insights right in Search. Set up a dedicated notebook for each class or project and add a range of sources like class slides, syllabi, relevant web articles, or your previous AI Mode threads on related topics. Your notebook brings it all together in one place, letting you ask questions, analyze topics, and build on your ideas over time.
R to @Google: 3️⃣ Learn step-by-step with Lens In the coming weeks, Lens in Search will offer a new interactive learning experience that helps you break down tough concepts. Just tap the Lens camera icon in the Google app and take a photo of what you’re working on – you’ll get an AI Overview that provides helpful explanations, pinpoints where you may have made mistakes, and coaches you when you’re stumped on a problem.
R to @Google: 2️⃣ Test your knowledge with practice quizzes You can now get customized practice quizzes directly in Search for any subject, including standardized tests: ACT, AP, ENEM, GRE, JEE, LSAT, MCAT, NEET, and SAT. We’ve partnered with leading education companies @ThePrincetonRev, @careers360, @physics__wallah, and Akira Enem, so you know you’re studying material that’s relevant to the exam.
R to @Google: 1️⃣ Understand concepts with interactive visuals Search can now generate custom tools and simulations to help you see complex topics in action.
R to @Google: 1️⃣ Understand concepts with interactive visuals Search can now generate custom tools and simulations to help you see complex topics in action.
muse spark was available earlier today ahead of an official announcement capacity isn't fully ready to go so we pulled it temporarily it will be back later today
R to @NVIDIAAI: Read the technical deep dive on SkillEvaluator: https://developer.nvidia.com/blog/evaluating-ai-agent-skill-performance-with-nvidia-skillevaluator/
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
Derek Chauvin was unjustly convicted of murder, therefore he should be freed. The facts show that he was not the cause of death, nor did he at any time intend for a death to occur. Whatever else he may be, he is not a murderer. That is the truth.
Peter Steinberger 發文稱「512GB RAM Studios. Apple was good to us.」並加上龍蝦 emoji,沒有進一步說明是哪款 Apple Studio 設備、採購背景或用途。文字可確認的是他提到 512GB RAM 等級的 Apple 工作站/Studio 設備,但不能推論具體跑什麼模型或效能表現。互動數未提供,不應視為熱度指標。
一龍馬判讀
高記憶體本機工作站對開發者、影像與 AI 工作流程都有吸引力,但這則證據只到硬體配置的個人分享,不能證明 Apple 在 AI 開發者市場取得新突破。
R to @AnthropicAI: We're also publishing a technical report: https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf And open-sourcing our…
R to @AnthropicAI: We're also publishing a technical report: https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf And open-sourcing our prompts and data here: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/tree/main
Anthropic 另指出,若要了解 Claude 如何執行這項蛋白質設計實驗與完整結果,需閱讀其官方研究部落格〈Claude accelerates protein design〉。這則貼文只提供導流資訊,沒有摘錄方法、指標或實驗結論。可確認的是 Anthropic 將此定位為 Claude 加速蛋白質設計的研究成果,而非單一產品公告。
R to @AnthropicAI: For more on how Claude ran this experiment and the full results, see our blog: https://www.anthropic.com/research/Claude-accelerates-protein-…
R to @AnthropicAI: For more on how Claude ran this experiment and the full results, see our blog: https://www.anthropic.com/research/Claude-accelerates-protein-design
R to @AnthropicAI: One of our highest priorities remains launching an access program for scientists to use our most capable models. We expect to share more on this soon. Opus 5 remains our most capable model available for life science research.
R to @AnthropicAI: Importantly, protein binders are not drugs. Designing a high-affinity binder is just the first step in the process of developing a drug-like molecule. Even designing a drug itself is just one phase out of the many required to establish that a drug is safe and effective before making it available to people. However, this establishes a strong foundation to work from, and we are building on it by teaching Claude to run the entire development process end-to-end for every major type of drug molecule—from antibodies to small molecules.
R to @AnthropicAI: Designing a binder is an easier process than designing a drug, but it’s a useful proxy. The typical success rate in the field today is between 10% and 15%. Between 22% and 35% of Claude's designs bound successfully, depending on the setup. Some of its strongest designs bound several times more tightly than the best published de novo binder.