Build, Manage and Deploy AI/ML Systems
🏋️ 训练微调与 MLOps
微调、分布式训练、强化学习与实验管理
易用、高性能的开源 RLHF 框架。
轻松在任意设备上做分布式训练与混合精度。
Containers for machine learning
Easily fine-tune (SFT/RL), evaluate, and deploy Qwen, Gemma, or any open agentic LLM/VLM!
魔搭社区的模型即服务平台开源库。
构建、部署与扩展 AI 应用和模型服务的统一框架。
slime is an LLM post-training framework for RL Scaling.
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
Type-safe, distributed orchestration of agents, ML pipelines, and real-time inference on your k8s — in pure Python with async/await, also other languages (rust, go and ts)
Tools for merging pretrained large language models.
The Open Source Feature Store for AI/ML
A repository of models, textual inversions, and more
自动化机器学习实验追踪、编排与部署平台。
Efficient Triton Kernels for LLM Training
Aim 💫 — An easy-to-use & supercharged open-source experiment tracker.
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
A Cloud Native Batch System (Project under CNCF)
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
Democratizing Reinforcement Learning for LLMs
PyTorch 原生的大模型后训练库。
A PyTorch native platform for training generative AI models