Inference Engines
High-performance model inference and serving runtimes
llama.cpp
8.3
★ 126.7k◇ 22.6kC++
vLLM
8.6
★ 90.7k◇ 21.6kPython
nomic-ai/gpt4all
7.1
★ 77.4k◇ 8.3kC++
ray-project/ray
8.6
★ 43.7k◇ 8.0kPython
gitleaks/gitleaks
8.1
★ 29.1k◇ 2.2kGo
liguodongiot/llm-action
6.7
★ 25.0k◇ 2.8kHTML
Lightning-AI/litgpt
7.9
★ 13.6k◇ 1.5kPython
halfrost/Halfrost-Field
7.1
★ 13.2k◇ 1.9kGo
bentoml/OpenLLM
7.4
★ 12.5k◇ 838Python
mistralai/mistral-inference
7.0
★ 10.8k◇ 1.1kJupyter Notebook
openvinotoolkit/openvino
8.2
★ 10.8k◇ 3.3kC++
Tiiny-AI/PowerInfer
6.7
★ 9.8k◇ 598C++
bentoml/BentoML
7.9
★ 8.8k◇ 1.0kPython
InternLM/lmdeploy
7.7
★ 8.0k◇ 734Python
ai-dynamo/dynamo
7.5
★ 7.9k◇ 1.5kRust
algorithmicsuperintelligence/openevolve
6.9
★ 7.3k◇ 1.1kPython
katanemo/plano
7.5
★ 7.0k◇ 487Rust
FareedKhan-dev/kimi-k3-in-c
7.0
★ 6.9k◇ 1.1kC
drumih/turbo-fieldfare
6.1
★ 6.5k◇ 409Swift
flashinfer-ai/flashinfer
7.8
★ 6.3k◇ 1.4kPython
kserve/kserve
8.0
★ 5.8k◇ 1.6kGo
Michael-A-Kuykendall/shimmy
6.6
★ 5.8k◇ 561Rust
gpustack/gpustack
7.1
★ 5.6k◇ 632Python
lemonade-sdk/lemonade
7.3
★ 5.6k◇ 480C++
1 / 2next →