开源 LLM 工程平台:追踪、评测、提示词管理与指标。
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
- 仓库
- open-compass/VLMEvalKit
- Stars
- ★ 4,437
- Forks
- 776
- 语言
- Python
- 许可证
- Apache-2.0
- 最近更新
- 2026-10-10
- 访问次数
- 0
LLM
同类项目
测试提示词、模型与 RAG 的评测与红队工具,CLI 与 CI 友好。
大模型少样本评测框架,Open LLM Leaderboard 背后的工具。
Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
用众多可复用的提示词模式增强人类能力的开源框架。
An AI prompt optimizer for writing better prompts and getting better AI results.