一龍馬/AI 情報站讀懂消息背後的脈絡
星期四
搜尋

jundot/omlx 是 Apple Silicon 專用的本機 LLM inference server

Python · ⭐ 19,805 · 🍴 1,689 · 🔥今日 +467

中文摘要

主打 continuous batching、分層 KV cache,以及可從 macOS 選單列管理模型與伺服器。README 說它支援 macOS App、Homebrew 與 source 安裝,OpenAI-compatible client 可連到 `http://localhost:8000/v1`,並有 `/admin` Web UI、模型管理、聊天、benchmark 與多語介面。限制相當明確:需要 macOS 15.0+、Python 3.11–3.13、Apple Silicon;GLM-5.2/MiniMax M3/Qwen3.5 等模型若未建置 native custom kernels 會退回較慢路徑,README 以 M3 Ultra 上 GLM-5.2 prefill 845 vs 約 29 tok/s 作為例子。

一龍馬判讀

這對想把 coding agent、OpenAI-compatible 工具接到本機 Mac 模型的開發者很實用,特別是重視離線、低延遲或資料不出機器的情境。門檻是硬體與系統版本綁定明顯,部分效能還依賴 full Xcode 或官方 DMG 預編譯 kernel。

原文節錄

jundot/omlx

oMLX LLM inference, optimized for your Mac Continuous batching and tiered KV caching, managed directly from your menu bar.

取得部分原文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

收錄日期
2026-08-20
來源
github.com/trending
抓取時間
2026/08/20 05:50(台北)
來源資料
Python · ⭐ 19,805 · 🍴 1,689 · 🔥今日 +467