Open Machine Learning Compiler Framework
🖥️ 本地推理与部署
Ollama、llama.cpp、vLLM:本地与生产环境的模型推理
以 OpenAI 兼容接口运行任意开源大模型,并可一键部署到云端。
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
单文件即可运行的 GGUF 模型工具,自带 Web 界面,适合角色扮演与写作。
NVIDIA 的通用模型推理服务器,支持多框架、多模型并发部署。
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Hugging Face 的生产级大模型推理服务,支持张量并行与连续批处理。
llama.cpp 的 Python 绑定,提供 OpenAI 兼容的 Web 服务。
High-speed Large Language Model Serving for Local Deployment
统一部署与服务各类开源大模型、嵌入与多模态模型,一行代码替换 OpenAI。
🤖 AI Gateway | AI Native API Gateway
在 Intel CPU / GPU / NPU 上加速运行与微调大模型。
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
Accessible large language models via k-bit quantization for PyTorch.
A Datacenter Scale Distributed Inference Serving Framework
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Fast, flexible LLM inference
在 Mac 上用 MLX 运行与微调大模型的工具包。
A framework for efficient model inference with omni-modality models
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
FlashInfer: Kernel Library for LLM Serving
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.