STACKQUADRANT

Inference Engines

High-performance model inference and serving runtimes

47 repos

llama.cpp

8.3

llama.cpp — a leading open-source project in the AI/LLM ecosystem.

126.7k22.6kC++

vLLM

8.6

vLLM — a leading open-source project in the AI/LLM ecosystem.

90.7k21.6kPython

nomic-ai/gpt4all

7.1

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

77.4k8.3kC++

ray-project/ray

8.6

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

43.7k8.0kPython

gitleaks/gitleaks

8.1

Find secrets with Gitleaks 🔑

29.1k2.2kGo

liguodongiot/llm-action

6.7

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

25.0k2.8kHTML

Lightning-AI/litgpt

7.9

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

13.6k1.5kPython

halfrost/Halfrost-Field

7.1

✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记

13.2k1.9kGo

bentoml/OpenLLM

7.4

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

12.5k838Python

mistralai/mistral-inference

7.0

Official inference library for Mistral models

10.8k1.1kJupyter Notebook

openvinotoolkit/openvino

8.2

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

10.8k3.3kC++

Tiiny-AI/PowerInfer

6.7

High-speed Large Language Model Serving for Local Deployment

9.8k598C++

bentoml/BentoML

7.9

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

8.8k1.0kPython

InternLM/lmdeploy

7.7

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

8.0k734Python

ai-dynamo/dynamo

7.5

A Datacenter Scale Distributed Inference Serving Framework

7.9k1.5kRust

algorithmicsuperintelligence/openevolve

6.9

Open-source implementation of AlphaEvolve

7.3k1.1kPython

katanemo/plano

7.5

Plano is an AI-native proxy server and data plane for agentic apps - centralizing orchestration, safety, observability, and smart LLM routing so you can deliver agents faster.

7.0k487Rust

FareedKhan-dev/kimi-k3-in-c

7.0

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

6.9k1.1kC

drumih/turbo-fieldfare

6.1

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

6.5k409Swift

flashinfer-ai/flashinfer

7.8

FlashInfer: Kernel Library for LLM Serving

6.3k1.4kPython

kserve/kserve

8.0

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

5.8k1.6kGo

Michael-A-Kuykendall/shimmy

6.6

⚡ Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary. FREE now, FREE forever.

5.8k561Rust

gpustack/gpustack

7.1

Performance-optimized AI inference on your GPUs. Unlock superior throughput by selecting and tuning engines like vLLM or SGLang.

5.6k632Python

lemonade-sdk/lemonade

7.3

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

5.6k480C++