Inference Engines
High-performance model inference and serving runtimes
xlite-dev/Awesome-LLM-Inference
6.6
★ 5.5k◇ 429Python
FellouAI/eko
6.7
★ 5.0k◇ 442TypeScript
ruvnet/ruvector
7.2
★ 4.5k◇ 590Rust
ruvnet/RuVector
7.2
★ 4.5k◇ 590Rust
algorithmicsuperintelligence/optillm
6.6
★ 4.3k◇ 385Python
predibase/lorax
6.6
★ 3.8k◇ 327Python
hemansnation/AI-Engineer-Headquarters
5.3
★ 3.7k◇ 698Jupyter Notebook
neuralmagic/deepsparse
5.9
★ 3.2k◇ 193Python
spiceai/spiceai
7.3
★ 3.1k◇ 225Rust
b4rtaz/distributed-llama
6.1
★ 3.0k◇ 248C++
FasterDecoding/Medusa
5.5
★ 2.8k◇ 205Jupyter Notebook
ovg-project/kvcached
5.8
★ 1.2k◇ 138Python
nobodywho-ooo/nobodywho
6.5
★ 1.1k◇ 77Rust
jjang-ai/mlxstudio
5.6
★ 960◇ 65
zhihu/ZhiLight
5.1
★ 908◇ 104C++
openinfer-project/openinfer
6.5
★ 669◇ 103Rust
pegainfer-project/pegainfer
6.5
★ 669◇ 103Rust
andrewkchan/yalm
3.6
★ 596◇ 64C++
zjhellofss/KuiperLLama
3.9
★ 571◇ 143C++
zengxiao-he/tessera
4.3
★ 562◇ 9Python
zhongkaifu/TensorSharp
6.0
★ 396◇ 39C#
interestingLSY/swiftLLM
3.8
★ 333◇ 37Python
avifenesh/memra
3.6
★ —◇ —OpenEdge ABL
← prev2 / 2