一行命令在本地运行 Llama、Qwen、DeepSeek 等大模型,自带模型仓库与 OpenAI 兼容 API。
Rapid-MLX
Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Up to 4× faster than Apple's MLX (mlx-lm
- 仓库
- raullenchai/Rapid-MLX
- Stars
- ★ 3,959
- Forks
- 430
- 语言
- Python
- 最近更新
- 2026-10-11
- 访问次数
- 0
本地部署
同类项目
纯 C/C++ 的大模型推理引擎,支持 CPU/GPU/Apple Silicon 与 GGUF 量化,自带 OpenAI 兼容服务。
高吞吐、低显存占用的大模型推理与服务引擎,PagedAttention 是其核心技术。
Never stop coding. Free MIT AI gateway: one endpoint, 359 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-awa
OpenAI 兼容的本地自托管 AI 引擎,一个服务同时支持文本、图像、语音与向量。
用家里闲置的手机、电脑和平板组成集群,一起运行大模型。