一行命令在本地运行 Llama、Qwen、DeepSeek 等大模型,自带模型仓库与 OpenAI 兼容 API。
🖥️ 本地推理与部署
Ollama、llama.cpp、vLLM:本地与生产环境的模型推理
纯 C/C++ 的大模型推理引擎,支持 CPU/GPU/Apple Silicon 与 GGUF 量化,自带 OpenAI 兼容服务。
高吞吐、低显存占用的大模型推理与服务引擎,PagedAttention 是其核心技术。
Never stop coding. Free MIT AI gateway: one endpoint, 359 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-awa
OpenAI 兼容的本地自托管 AI 引擎,一个服务同时支持文本、图像、语音与向量。
用家里闲置的手机、电脑和平板组成集群,一起运行大模型。
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
微软 1-bit 大模型的官方推理框架,让 CPU 也能高效跑大模型。
开放的聊天机器人训练、服务与评测平台,Chatbot Arena 背后的项目。
Hundreds of models & providers. One command to find what runs on your hardware.
面向大模型和多模态模型的高性能服务框架,前后端协同设计,推理延迟低。
AirLLM 70B inference with single 4GB GPU
7.4 billion tokens per month. 34 free LLM providers. 635 free model endpoints. All behind one /v1 endpoint, plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experime
苹果官方为 Apple Silicon 打造的数组与机器学习框架。
把大模型打包成单个可执行文件,下载即可跨平台运行。
ncnn is a high-performance neural network inference framework optimized for the mobile platform
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
基于 TVM 编译的通用大模型部署引擎,可在手机、浏览器与各类 GPU 上本地运行。
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
High-performance In-browser LLM Inference Engine
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
NVIDIA 官方大模型推理优化库,在 GPU 上提供极致吞吐与低延迟。
Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, ASR, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.