一龍馬/AI 情報站讀懂消息背後的脈絡
星期四
搜尋

他還提到 DeepSeek V4 風格的 mHC 四路殘差路徑與原生視覺編碼器;但貼文是個人技術解讀,沒有附官方模型卡、訓練資料或 benchmark

中文摘要

Sebastian Raschka 稱,先前受關注的 Ox Alpha LLM 其實是 GLM-5.3-Flash,並整理其相對 GLM-5.2 的架構變化:採用類 Kimi Linear 的 3:1 混合注意力,包含 34 層 Kimi Delta Attention 與 11 層 MLA/DSA,MoE 主幹也從 744B-A40B 縮到 320B-A18B。他還提到 DeepSeek V4 風格的 mHC 四路殘差路徑與原生視覺編碼器;但貼文是個人技術解讀,沒有附官方模型卡、訓練資料或 benchmark。

一龍馬判讀

若屬實,這代表高階模型架構正在把多種高效率注意力與稀疏 MoE 混搭,以壓低推理成本並保留多模態能力;但在缺少官方評測前,還不能把架構描述直接等同於效能優勢。

原文節錄

Sebastian Raschka · @rasbt

Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash...…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers; - a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B; - a DeepSeek V4-style mHC residual path with four parallel streams; - plus a native vision encoder (not shown). * "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA + full attention GQA, DeepSeek V3.2 uses DSA + full attention MLA. PS: Sry for the excessive tech jargon. Explainers on all these components (MLA, DSA, KDA, mhC, etc.) in my LLM Architecture Gallery PPS: Haha, maybe justification for getting that pricey Mac Studio M5 Ultra 256 GB / 512 GB to run this locally...

收錄日期
2026-08-27
來源
Nitter RSS(公開貼文)
抓取時間
2026/08/27 06:12(台北)