What does it take to move from isolated agent experiments to production AI at enterprise scale? Our latest guide brings together lessons from @SchneiderElec, V@odafone, and @mondaydotcom across: - Shared agent platforms and LLMOps - Production observability and evaluation - Multi-agent architectures - Clearer boundaries between agents, tools, and environments - Security, governance, and data residency - Scaling adoption across teams and business units https://info.langchain.com/guide/scaling-agents-in-europe-and-the-middle-east
A thing to know about the AI business is labs that have frontier models can release half-built products and they work surprisingly well because the AI can just figure stuff out and improvise. Its like including a forward deployed engineer & customer service agent in the product.
Meet AMD Ross™, the agentic AI assistant for embedded developers. AMD Ross automates complex, repetitive workflows, scales expertise, and helps teams move faster, from design intent to execution through natural language interactions. Learn more: https://bit.ly/4daG3Cb
Igor Nikolaienko, AI Architect at @DeutschePostDHL is taking the stage at Interrupt, the Agent Conference by LangChain. Catch Igor’s talk + more: https://interrupt.langchain.com/london
Pydantic AI tools now run commands and edit files through one API, ctx.workspace, without caring where. Your machine, an SSH host, a Bubblewrap sandbox, or a Modal, E2B or http://Fly.io Sprites sandbox: swap one capability and your tools stay the same. https://pydantic.io/teW92
Pydantic AI agents now run on Jev from TypeSafe, and Pydantic Logfire shows each answer's probability and the options it almost picked. For yes/no or pick-from-a-list agents, Jev skips generating text, so it's faster and cheaper than a frontier model. https://pydantic.io/RctKd
Leaving aside Anthropic's incentives for publishing this research, there is no doubt that open weights models will soon create the same security threats that closed source models have been demonstrating, except without guardrails. We are close. Probably good to plan accordingly.
A video world model can make a convincing clip and still get the physics wrong. Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions. The captions explain why and how a scene unfolds. We use them to fine-tune world models and add that reasoning to prompts when generating video. Adding physics reasoning to the prompt alone improved NVIDIA Cosmos 3’s PhyGenBench score by 5.62 points without retraining. Paper and project: https://physis-intelligence.github.io/physis-lang-web/
Codex Security Cloud is getting a major upgrade, with access to cyber-capable models through Daybreak Blue included by default. It scans entire GitHub repos, continuously reviews new commits, investigates and deduplicates findings, and prepares fixes for review – even when your laptop is closed. Available as a plugin in Codex desktop and web.
It does feel like every year OpenAI releases a variation on the same third-party ecosystem only to semi-abandon it: Plugins in 2023, GPTs and the GPT Store in 2023-2024, Apps in 2025, and now Plugins (same name, different thing than before) in 2026.
After a brief period where OpenAI seemed to be unifying work around the ChatGPT app, between Dot and Spaces and Pages and local/cloud ChatGPT Work and Scheduled Tasks in the Cloud and Scheduled Tasks on your computer, everything is getting quite confusing and overlapping again.
More users shouldn't have to mean slower AI. In @Signal_65 RAG testing hosted via @TensorWave, AMD Instinct MI355X delivered ~42% lower p99 latency and roughly 2x the concurrent-user headroom before breaching the evaluated SLA: https://bit.ly/45hAvl3
R to @LangChain: Engine learns how your agent works by reading traces and repos, then tests for agent-specific weaknesses and flags any it confirms. You get a list of verified issues – from hallucinations to violations of system prompts – so you can fix them before they impact your users.
What do you want from AI? We’re launching a new study with Anthropic Interviewer to learn more about your experiences using AI, what role you want it to play in your life and the world, and what you want from the companies building it. Last December, 81,000 people told us about their hopes and fears about AI in the largest qualitative study ever done. This time, we’re giving participants the option to make their responses public so that anyone, not just Anthropic, can learn from them. What you tell us will shape The Anthropic Institute’s research and inform the decisions we make. If many people say companies like Anthropic should be doing something differently, that will be on the record, where anyone can point to it. The study runs Sept 29 to Oct 6 and is open to Free, Pro, and Max users on Claude and Claude Code. Take part here: http://claude.ai/anthropic-interviewer/your-thoughts-on-ai?from=social
R to @emollick: A broken early version of whatever product they are shipping is still better then what a well-built traditional product a lot of the time because LLMs are just unreasonably effective tools at doing a vast array of stuff and people & companies are learning to just roll with it.
Every single agent will need to verify themselves to prove their intentions on the internet. That's why we're pioneering Web Bot Auth, a protocol that allows for good Agents to use the web. We wrote a blog about the internals and how it works: https://browserbase.run/wba
What is a way to test human-level generality in artificial intelligence? "The meta-benchmark of being able to pass ARC-AGI-(n+1) immediately upon release"
Launching a new Codex Cloud, much improved from last year. With configurable cloud environments, it’s impossible to go back to building on your laptop once you’ve taken the time to configure it. Agents API, the same tech powering all our cloud agents, including dots, is also now in preview and supports computer use. It allows you to build the same incredible products we are making available today.
Trajectories in LangSmith are a chronological, human-readable view of an agent session. ✅ Every run rolls into a trace, and traces from a session link into a thread: the full execution tree, nested runs and all. ✅ A trajectory strips that out, keeping each message once, in order, so you read the session as the path the agent took.
You can now use your ChatGPT subscription directly in over 16 partners products. No little rules, you can just use all your included usage right there. This includes things such as Devin, OpenCode, Notion, and many more. Sign in with ChatGPT and go.
We are opening up our platform and you can now build full native apps with plugin extensions, and ship them right in ChatGPT. We have over 1.2B weekly users and will surface relevant plugins right in the conversations.
I only used Dots briefly before launch, so can’t offer a detailed review, but I found it to be good & very much in the rapidly-expanding Clawlike category with Muse & Grokbot. Using a capable model with access to your data as an assistant & second opinion is remarkably useful.
A lot of messaging around AI (by its proponents) enthusiastically frames it as a form of destruction, obliteration even: "AI just killed XYZ" - "It's over for XYZ" (in reality, this is practically never accurate) The way regular folks perceive AI is very much shaped by such takes. The backlash is inevitable.
R to @emollick: I don't even know which tool has permission to do which things on which devices. Like Dot can create threads and tasks in Codex/ChatGPT app & Voice mode also creates threads and sometimes those threads continue and sometimes they are abandoned and none of this is very clear now.
We will have a few million dots online within days, working on all sorts of things across such a diverse and large community. Excited to learn from all of you on what you love and what doesn’t yet feel magical. Personally I felt a jump after 2-3 days of use after teaching it more about my preferences and things on my mind. It learns very quickly to be most useful and it can take on surprisingly ambitious tasks on its own. We’re learning from how you all use your primary dot before releasing the ability to create an entire team of them.
It was so nice to connect with so many of you in person at DevDay. Immaculate vibes from the crowd and fun to see so many build things throughout the day.
I got community noted, but the note is wrong 😅 Your primary dot is included in your plan and will be available 24/7. If you ask it to create a codex task for you, that one will be drawing usage as usual. But when your dot does work directly, that uses nothing and is all on top of your plans usage. In the future you will be able to increase the speed and allow your dot to have more bandwidth, this will be a paid feature, but the baseline functionality will always just be included in your plan.
I'm at OpenAI's DevDay event in San Francisco today - as I have for the past three DevDay events, I'm running a live blog where I'll be posting updates during the keynote, which starts in five minutes https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/
RT by @steipete: Congrats to the OpenClaw community on OpenClaw Enterprise! 🦞 Organizations can use NVIDIA OpenShell as an open source option for governing agents with OpenClaw Enterprise. Glad to keep working with the community to make agents safer.
Workera is being acquired by Pearson! It's been a privilege to support CEO @kiankatan and the whole Workera team through this journey. Back in 2019, Kian had the insight that rigorous skills measurements would be important. With recent advances in AI, this is now more true than ever. Workera's technologies for rigorously measuring people's skills in different tasks and job roles are helping many businesses understand where their employees are strong and where there are areas for development. With Workera and Pearson combining forces, I'm delighted that under @omarabbosh and Kian's leadership. Workera now has the potential to support an even larger number of people and businesses. Thank you, Kian and the Workera team for having me serve as the company's chairman, for launching Workera out of http://DeepLearning.AI and AI Fund, for your friendship over the years, and for your important work advancing assessments.
Dots are here! A new way to use AI that works 24/7 for you; get more of your time and attention back to work at a higher level. https://openai.com/index/introducing-dots/
R to @fchollet: Also if you're going to brag about obliterating something, maybe first check whether that thing is regarded by most of the population as an important source of meaning and joy. I don't know, I'm not a PR person
R to @NVIDIAAI: Physis-Lang with NVIDIA Cosmos 3 Super and Nano rank #1 and #2 on the Physics-IQ Verified image-to-video leaderboard. The benchmark tests how well generated videos follow real-world physics. Full rankings: https://physics-iq-verified.anates.ai/?track=i2v
Live demos suffered from rolling out all the updates at the same time, but were committed to keep doing these live in the future. Almost all the things we announced will be available today. Can’t see what you will do with it.
R to @OpenAI: We’re also reopening Pro 200 subscriptions, with continued access to frontier models like Astra, including our new GPT-6.1 Sol model which brings near-Astra capabilities to a model you can use every day. Our commitment is that your subscription will continue to help you get more work done with an increasing level of quality. We are also committing to not reintroducing the 5hr limit so that you can fully use your weekly usage when you want.
R to @OpenAI: Ultrafast is available today for GPT-6 Astra in Codex, ChatGPT Work, and the API, with GPT-6.1 Sol coming soon. To access it in Codex and ChatGPT Work, we’re introducing Pro 500—a new plan with our highest usage limits (25x Plus) and access to Ultrafast. https://chatgpt.com/pricing
Token traffic on responses API has increased 100X compared to today last year. Meanwhile the team improved reliability with 99.9% uptime and improving performance significantly.
R to @OpenAI: GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. https://openai.com/index/introducing-gpt-6-1-sol
R to @OpenAI: GPT-6.1 Sol shows major improvements over GPT-6 Sol in our alignment evaluations, bringing it more in line with GPT-6 Astra. It’s more transparent about its limitations and more reliable at respecting user intent and safety constraints.
R to @OpenAI: GPT-6.1 Sol is a significant upgrade over GPT-6 Sol across coding, computer use, and complex professional work—approaching GPT-6 Astra on several benchmarks at substantially lower cost.
R to @OpenAI: GPT-6.1 Sol makes frontier intelligence more affordable, so you can use it for more of the work that matters and developers can build and run applications at scale. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT-6 Sol’s cached input pricing.
R to @LangChain: ✅ Debug long-running agents faster, skip the reconstruction ✅ Give SMEs a readable view to review agent behavior ✅ Score agent behavior with online evals ✅ Turn production sessions into fine-tuning datasets Try Trajectories in LangSmith today. https://www.langchain.com/blog/langsmith-trajectories-tracing
R to @OpenAI: Dots will be available in ChatGPT on web, mobile, and desktop across Pro, Business Premium, and Enterprise users in eligible markets. To get started, create your first dot in the ChatGPT desktop app or your desktop browser, connect your apps, and let it introduce itself.
R to @OpenAI: You always stay in control. Set boundaries and specify what your dot can do on its own, when it should ask first, and what it should never do. Safety is built in. You choose which apps to connect, and your dot runs on its own cloud computer—so connecting yours is entirely optional. https://openai.com/index/how-we-build-safety-security-and-privacy-into-dots/
R to @OpenAI: Your dot can do simple things like book a table—or take on your most ambitious work with the initiative of a high-agency engineer or chief of staff. It has its own computer and works across 4,000+ apps through ChatGPT, using the plugins you connect. It learns what you need and gets to work before you ask, quietly in the background, 24/7. https://openai.com/index/introducing-dots/
R to @AnthropicAI: Making your interview public is completely optional. Our blog post covers the benefits and possible risks of doing so. Read it here: https://www.anthropic.com/research/your-thoughts-on-ai
R to @emollick: It also has a natural analog on teams as a persistent entity with a job that you can delegate to using existing team communication tools like Slack. I expect more Clawlikes soon, though I don’t think they are the final form of this type of AI.
for friends of @latentspacepod we are gonna be asking all your burning Dots, Sol 6.1, CUA, and Decisions API questions for @AriX and @nikunjhanda 🔜 at the gateway pavilion on a few mins, come by!