Your organization is not spending enough of its efforts on bolstering cybersecurity during the window before open weights Mythos-class models/harnesses become…
Your organization is not spending enough of its efforts on bolstering cybersecurity during the window before open weights Mythos-class models/harnesses become available. The HuggingFace incident shows us that you don’t even need intentional bad actors to be exposed to AI hacking.
Our 2nd founder dinner in SF co-hosted by @jerryjliu0 and @GuangyuRobert at @Fundamental - the team behind @tryshortcutai Talking about existing moats in the AI…
Our 2nd founder dinner in SF co-hosted by @jerryjliu0 and @GuangyuRobert at @Fundamental - the team behind @tryshortcutai Talking about existing moats in the AI era. Frontier labs are moving past model APIs into vertical agents - ChatGPT Health, Claude for Legal. So where does the moat sit now? ✅ Agent engineering ✅ Infra optimization ✅ Domain evals and data ✅ Workflow expertise ✅ GTM and brand If you're a founder or CTO shipping agents in production and want to know what other teams are doing to maintain their moat. Request a seat 👉️ https://luma.com/llamai-8hry
Tracking the provenance of synthetic content is becoming a regulatory requirement. To comply with new laws like the EU AI Act, Anthropic will embed invisible watermarks in all future Claude models — and eventually, older ones too. For generated text, Claude implements Google’s SynthID methodology, using a seed generator to nudge word choices in specific directions in order to create a statistical pattern. This can then be detected by a scoring API. For images, it embeds C2PA metadata. Claude claims this will not meaningfully affect output, but users are skeptical. We look at the technical implementation, the probability of false positives, and the downstream implications for output quality. 📊Read the analysis: https://hubs.la/Q04vcghP0 (https://hubs.la/Q04vcghP0)
This is a sign of the future: 1) Explainability of AI actions is already tenuous 2) It gets more tenuous at extreme scale of multiple agents working over long periods because they produce so many thinking tokens 3) The only way to solve this is with other AIs & they are limited
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
Creating richer digital worlds takes more than imagination. It takes breakthrough compute. On the latest episode of Advanced Insights, AMD CTO Mark Papermaster joins @WetaFXOfficial CTO Kimball Thurston to explore the future of visual effects, from large-scale rendering and simulation to AI-assisted creative workflows. Watch now: https://bit.ly/3Uix1fN
For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools. https://www.anthropic.com/research/enabling-independent-research
Today we’re introducing Gemini 3.5 Transcribe, our latest transcription model built for incredibly precise, smart dictation across your favorite apps and…
Today we’re introducing Gemini 3.5 Transcribe, our latest transcription model built for incredibly precise, smart dictation across your favorite apps and devices. Remember when traditional speech-to-text meant shouting over background noise, constantly hitting backspace to fix misspelled words, and manually deleting every "um" and "uh"? Those days are over. Gemini 3.5 Transcribe isn't just dictation — it’s active intelligence with precise, context-aware speech-to-text support in 85+ languages. The model automatically filters out filler words, formats unstructured speech, and even pairs with your screen context to execute voice commands. Watch as Gemini 3.5 Transcribe removes filler words and uses multimodal capabilities to seamlessly turn messy voice input and local files into a polished email draft.
Interrupt NYC is less than a month away. This one-day event is made for AI leaders and engineers building and deploying agents in production. Info + RSVP: https://interrupt.langchain.com/nyc
Build coding agents that learn from experience. In Building Adaptive AI Agents, you’ll turn an agent’s own traces into reusable skills and build a code knowledge graph that finds the right context where keyword search misses. Built in partnership with @Oracle and taught by Nacho Martínez (@jupiterwanderer) and Casius Lee. Enroll for free: https://hubs.la/Q04vlNYg0
它提醒代理互動不只是 API 協定問題,還會混入人類設計、角色設定與社會行為;但缺乏證據支撐時,較適合作為研究問題線索,而非產品或技術趨勢的定論。
原文節錄
Ethan Mollick · @emollick
Moltbook, despite being weird and compromised and full of human roleplaying, was 100% a harbinger of what actually happens when agents communicate in the wild.
Moltbook, despite being weird and compromised and full of human roleplaying, was 100% a harbinger of what actually happens when agents communicate in the wild.
From the show floor to the AI factory: Hot Chips 2026 was all about extreme co-design to accelerate agentic workloads — the most complex workload in history. Vera CPU, Vera Rubin, Groq 3 LPX, Spectrum-X Multiplane, BlueField-4 Scale-In networking. One full AI stack platform built for agents.
In just three minutes, see how LangSmith Engine can help you automate agent improvement by: 🔎 Identifying issues 💡 Proposing fixes 🔨 Building dataset examples 👀 Monitoring for regressions
It is a sign of how fast things have changed that this is what OpenAI announced as the future of enterprise AI agents back in October 2025. They killed Agent Builder in June. (But a surprising number of enterprise products still use this model for agents).
Friction is what builds taste and mastery. The tools that remove it also remove what made us good enough to use them well. Worth reading if you're leading teams or are an early-career engineer: https://larsfaye.com/articles/ai-coding-will-prevent-expertise
Did you know: Your agent can now connect to 1,384+ apps through Composio. Some of those apps are growing faster than others. Here are the 7 fastest-growing apps in the Composio ecosystem this week 📷 ClickSend: +2,059% KieAI: +711% Wrike: +462% HeyReach: +431% BaseLinker: +375% Rocketlane: +219% Google Meet: +170%
Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers; - a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B; - a DeepSeek V4-style mHC residual path with four parallel streams; - plus a native vision encoder (not shown). * "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA + full attention GQA, DeepSeek V3.2 uses DSA + full attention MLA. PS: Sry for the excessive tech jargon. Explainers on all these components (MLA, DSA, KDA, mhC, etc.) in my LLM Architecture Gallery PPS: Haha, maybe justification for getting that pricey Mac Studio M5 Ultra 256 GB / 512 GB to run this locally...
We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet 🔊 It turns audio into precise transcription in 85+ languages, removing…
We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet 🔊 It turns audio into precise transcription in 85+ languages, removing filler words like "ums" and "ahs" while handling self-corrections and capturing your intent so you can get things done using just your voice.
Save the date for IFA 2026! @jackhuynh, AMD SVP & GM of Computing and Graphics, will deliver the opening keynote and share the AMD vision for the next era of personal AI. 📅 Sept. 4 | 10:45 AM CEST 📺 https://bit.ly/3UUZecD
The @GeminiApp can transform your questions and complex topics into custom and interactive visualizations — directly within your chat. Whether you’re rotating a molecule or simulating a complex physics system, you create 3D simulations with just one prompt in Gemini. Try starting your prompt with “show me...,” to get your interactive visualization. For the best results, try using Gemini’s Flash model.
I know it is a lot for instructors who never asked for AI to change things, but if you teach and have not updated your expectations over the last 6 months over what AI can do (a lot more), its error rate (a lot less), and its ubiquity among your students (near total), you need to
R to @OpenAI: We worked with METR and Redwood Research to conduct a third-party assessment of the model behavior observed during the incident. They’re sharing a report of their findings: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
R to @AnthropicAI: Now, we want to scale this research model. If you're a researcher and would like access to our tools to pursue work you can't otherwise do today, we’d like to hear from you. You can express interest here: https://forms.gle/rmLjTvibven9CmDFA
Anthropic 表示,另外兩項針對 Claude 使用情境的研究仍在進行中:Oxford 的 Human Information Processing Lab 會研究 Claude 行為與使用者感受之間的關係,METR 則在估算 coding agents 在真實世界帶來的生產力增益。這則貼文只說明研究主題與合作單位,尚未提供方法、樣本細節或結果。Anthropic 稱之後會公布更多內容。
一龍馬判讀
這把 AI 評估從模型分數推向使用者體驗與實際工作產出,但目前仍停在預告階段,不能先推論 Claude 對情緒或生產力的正負效果。
原文節錄
Anthropic · @AnthropicAI
R to @AnthropicAI: The other two studies are ongoing: HIP Lab is studying how Claude's behavior relates to how people feel when using AI, while…
R to @AnthropicAI: The other two studies are ongoing: HIP Lab is studying how Claude's behavior relates to how people feel when using AI, while METR is estimating real-world productivity gains from coding agents. We'll share more from both soon.
Anthropic 引述 Stanford Social and Language Technologies Lab 對人類與 AI 協作的研究,稱研究發現超過一半的對話涉及「consequential tasks」,也就是會影響他人或很難復原的工作。貼文提供 alphaxiv 完整寫作連結,但這裡的證據只有 Anthropic 對結果的摘要,沒有列出分類標準、錯誤率或任務範例。這項研究看起來是同一系列 Claude 對話分析的一部分。
一龍馬判讀
如果大量 AI 對話已進入難以回復的決策或工作流程,企業與產品團隊就不能只把聊天機器人當成低風險助理;但判讀仍需要看原研究如何定義與標註這些任務。
原文節錄
Anthropic · @AnthropicAI
R to @AnthropicAI: The SALT Lab studied how people collaborate with AI.…
R to @AnthropicAI: The SALT Lab studied how people collaborate with AI. They found that over half of these conversations involved consequential tasks—work that affects other people or is hard to undo. Read their full writeup here: https://www.alphaxiv.org/abs/2608.human-ai-collaboration-at-scalev1
Anthropic 說,三個研究團隊——Stanford 的 Social and Language Technologies Lab、Oxford 的 Human Information Processing Lab,以及 METR——設計了獨立研究,分析 2026 年 4 月到 5 月間 25 萬筆 Claude.ai 或 Claude Code 對話的彙整輸出。貼文強調資料是 aggregated outputs,但沒有說明去識別化方式、抽樣條件、使用者同意流程或可重現性細節。這則是整個研究系列的資料來源說明。
一龍馬判讀
25 萬筆真實使用脈絡能補上實驗室評測看不到的 AI 使用樣貌;同時,隱私保護與資料偏差會直接影響外界能否信任研究結論。
原文節錄
Anthropic · @AnthropicAI
R to @AnthropicAI: Three research groups—Stanford’s Social and Language Technologies lab, Oxford’s Human Information Processing Lab, and METR—designed…
R to @AnthropicAI: Three research groups—Stanford’s Social and Language Technologies lab, Oxford’s Human Information Processing Lab, and METR—designed independent studies to analyze the aggregated outputs from 250,000 https://Claude.ai or Claude Code conversations between April and May 2026.
R to @GoogleAI: — Try it out in the @Geminiapp on macOS and Gboard on @Android — Build in the Gemini API via @googleaistudio, @antigravity, and the Gemini Enterprise Agent Platform (public preview) — Coming soon to @googlechrome and Gemini Enterprise for Customer Experience Learn more: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/
R to @GoogleDeepMind: Here’s what’s new: 🔵 It’s better at understanding complex phone numbers, postal codes, and order IDs – even in noisy environments.…
R to @GoogleDeepMind: Here’s what’s new: 🔵 It’s better at understanding complex phone numbers, postal codes, and order IDs – even in noisy environments. 🔵 It can remove filler words and auto-format your text. 🔵 It recognizes custom vocabulary so the model knows unique names and product titles. 🔵 It automatically detects and transcribes speech in 85+ languages. Try it now in the @Geminiapp on macOS and Gboard on @Android. Start building in @GoogleAIStudio and @Antigravity → https://goo.gle/4gzP1K8
Tip: Pick your model and effort at the top of a session. Switching models mid-session forces a full uncached re-read of your entire context. The KV cache is tied to specific weights - it can't transfer. Effort level and fast mode work the same way. It's a one-turn tax that scales with depth. Okay to switch on the first few turns but expensive later on.
R to @LangChain: For @Rippling, LangSmith makes pulling and analyzing all conversations at scale simple. “The ability to pull and analyze all conversations at scale… LangSmith makes that possible. We have a bunch of automated analysis running on top of it.” — Laks Srini, Product Owner https://www.langchain.com/blog/how-rippling-went-ai-native-across-every-product-in-6-months-with-deep-agents-and-langsmith
Inside @Rippling’s eval pipeline: ✅ Offline evals: Pre-recorded mocks + fixtures that run locally on every commit without external dependencies. ✅ Post-merge integration evals (online): 300-400 queries against a full Rippling sandbox to validate system health before deployment. ✅ Deploy-blocking evals (online): ~10 critical scenarios against real systems that gate every deployment. ✅ Continuous evals (online): Scheduled runs against prod data, multiple times daily, monitoring live system health.
It’s called Composio :) OpenRouter: 500+ models Composio: 1,384 apps (and growing) Even better news: There is a generous free tier with 100K tool calls/month.
Good news: it already exists. It’s called Composio :) OpenRouter: 500+ models Composio: 1,384 apps (and growing) Even better news: There is a generous free tier with 100K tool calls/month.
R to @Google: Gemini 3.5 Transcribe powers Rambler on @Android. We first introduced Rambler at the Android Show to help turn spoken thoughts into polished text. It uses 3.5 Transcribe to automatically remove filler words and clean up speech. You can also use your voice to make edits, correct misspellings, and even change the writing style.
R to @Google: In the @GeminiApp on macOS, you can use Gemini 3.5 Transcribe to get more done 🗣️✨ Ask Gemini to generate images, look up information, summarize text, analyze files and more using just your voice.
R to @Google: Instead of typing out every thought, Gemini 3.5 Transcribe works with the way you actually talk: ✅ Seamlessly handles self-corrections ✨ Removes…
R to @Google: Instead of typing out every thought, Gemini 3.5 Transcribe works with the way you actually talk: ✅ Seamlessly handles self-corrections ✨ Removes filler words to deliver clean, formatted text 🎯 Understands your natural intent and speaking style 🔊 Accurately captures audio in noisy environments 👥 Distinguishes up to 3 speakers with timestamps 🌍 Supports 85+ languages, including regional accents and dialects
R to @LangChain: Watch or listen to the latest Max Agency on your favorite podcasting platform. 🎧 Apple: https://podcasts.apple.com/us/podcast/how-unify-cut-its-ai-agent-costs-95-in-two-weeks/id1891551672?i=1000783151404 🎧 Spotify: https://open.spotify.com/episode/6kWQouc2QmiHGk0vdiZEtd?si=ba6e241ca4a24faf ⏯️ YouTube: https://youtu.be/6898VdRtKDE
R to @composio: How we measured growth: We compared the # of tool calls from the past 7 days vs. the previous 7 days across all 1,384 apps. Apps needed at least 5000 tool calls to qualify, so small sample sizes wouldn’t skew the rankings. We’ve hidden absolute tool call numbers for privacy reasons.
R to @Google: We're offering one year of Google AI Pro at no cost for eligible college students in the U.S. through December 31, 2026. Terms apply 🎉 Learn more about the offer, plus our new and enhanced @GeminiApp tools for students: https://goo.gle/3U4e941
PSA: do not use codex "locked use" capabilities right now. it is currently relying on unstable mac features and has completely locked me out of my macos keychain twice this week. thx @_chenglou for linking to apple developer forums acknowledging this is a "known bug". just avoid. ofc, would be nice to do everything in cloud, but cloud isn't there yet.