As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://openai.com/index/pacing-model-development-cyber-capabilities/
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. https://openai.com/index/pacing-model-development-cyber-capabilities/
We just released TensorRT Model Connect in Public Preview. You can take a supported @huggingface model to end-to-end TensorRT inference in just two commands. No intermediate ONNX export, and the resulting bundle can run through native C++ APIs. We also built the entire project with @OpenAIDevs Codex agents, with humans directing and reviewing the work. That includes model implementations, performance tuning, tests, integrations, and docs. It’s open source, so go try it out, dig into the implementations, or contribute support for a new model: https://github.com/NVIDIA/TensorRT-Model-Connect
What are you building with open models? Show us what you’re working on and you could be headed to #NVIDIAGTC Berlin. 🎫 Your Golden Ticket includes: → Free conference pass → VIP seating for Jensen’s keynote → Exclusive NVIDIA merch → Access to special events Submissions are open Aug 18 – Sep 10, 2026. Details and how to enter: https://nvda.ws/4wFdunb
If you're building a software factory, code good enough to ship still needs human taste and ownership. You'll likely need humans in the loop upfront for deciding on product intent, system design (if you care) and your quality bar. Do review code (lights-on factory) but be intentional with where it's needed the most. I've found you want to watch out for where automated back-pressure breaks. Or where maintainability trade-offs need to be made. Aim for quality checks to happen as early and continuously as possible. Not all of them have to, but this includes type systems, automated tests, mutation testing, security scanners and linting for architecture rules. Number of checks != quality. You'll likely need to experiment with what checks give you the best signal to noise ratio. Be ready to tighten or relax your constraints deliberately. You want to build your factory so some aspects of human taste get encoded in the environment, the agent gives you evidence of its work being right and where a human still "owns" what ships to production.
Your team's using Claude Code, Codex, Cursor? @coderhq runs the agent on your own infra - isolated, any model, fully audited. Real diffs, real control. https://fandf.co/4foRRlL
we've been doing a lot of a/b testing of @aiDotEngineer youtube thumbnails. i always hated that it is such an opaque process. open sourcing/crowdsourcing our learnings today! https://ai-engineer-thumbnail-lab.swyx.chatgpt.site/#insights doing this in hopes that people can share their experience or learn from ours. at the end of the day we just want to get good educational content to rise above the noise online. please lmk what you think
Claude 官方宣布,Claude 現在可連接 Gmail 與 Google Drive:使用者可要求它回覆郵件串,並由 Claude 起草與寄出回覆;Google Drive 則可進行檔案管理。官方強調使用者可控制哪些動作需要批准,入口在 connectors menu。此功能開放給所有付費方案。
一龍馬判讀
這把 Claude 從聊天助理推進到可操作個人工作資料的代理工具,受影響的是重度使用 Google Workspace 的付費用戶。風險也更直接落在郵件誤寄、檔案權限與核准流程設計上。
原文節錄
Claude · @claudeai
Claude can now send emails in Gmail and manage files in Google Drive.…
Claude can now send emails in Gmail and manage files in Google Drive. Ask Claude to reply to a thread, and it drafts and sends the response. You control when it needs your approval. Connect Gmail or Google Drive from the connectors menu to try. Available on all paid plans.
We've just surpassed 3 million models on the Hub 🤗 the community is accelerating towards an open, distributed future where open AI is everywhere, for everyone 🚀
What is an obvious thing that we should do with Codex, API or our models that we should just do but haven't yet? What is 100% within reach, but we just seem to be missing?
R to @OpenAI: We’re sharing the concrete changes we’re making to strengthen monitoring, security, and alignment as capabilities advance. We’ve introduced stronger workload and network isolation, continuous security testing, and expanded multistage monitoring for higher-risk training, evaluations, and tool-using inference. These safeguards are designed to detect concerning behavior quickly and limit what systems can access or affect.
NVIDIA AI 宣傳一場「Ask the Experts: What's New in the Nemotron Open Family」直播,主題是 Nemotron 開放模型家族的新內容。證據只有一則 X 貼文與直播連結,沒有列出新模型規格、授權、效能數據或發布內容。現階段只能確認 NVIDIA 正在以 Nemotron Labs 名義做公開說明,不能推斷具體產品更新。
How we innovate matters. So does the impact we make. Explore the progress we’ve made in our 2025-26 Corporate Responsibility Report: https://bit.ly/4mRH2cW
Gimme, gimme, gimme Codex after midnight Won’t somebody make these failing tests all go away? Gimme, gimme, gimme Codex after midnight Ship it through the darkness by the start of the day