STACKQUADRANT

Inference Engines

High-performance model inference and serving runtimes

47 repos

xlite-dev/Awesome-LLM-Inference

6.6

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

5.5k429Python

FellouAI/eko

6.7

Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai

5.0k442TypeScript

ruvnet/ruvector

7.2

RuVector is a High Performance, Real-Time, Self-Learning, Vector Graph Neural Network, and Database built in Rust.

4.5k590Rust

ruvnet/RuVector

7.2

RuVector is a High Performance, Real-Time, Self-Learning, Vector Graph Neural Network, and Database built in Rust.

4.5k590Rust

algorithmicsuperintelligence/optillm

6.6

Optimizing inference proxy for LLMs

4.3k385Python

predibase/lorax

6.6

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

3.8k327Python

hemansnation/AI-Engineer-Headquarters

5.3

A collection of scientific methods, processes, algorithms, and systems to build stories & models.

3.7k698Jupyter Notebook

neuralmagic/deepsparse

5.9

Sparsity-aware deep learning inference runtime for CPUs

3.2k193Python

spiceai/spiceai

7.3

A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.

3.1k225Rust

b4rtaz/distributed-llama

6.1

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

3.0k248C++

FasterDecoding/Medusa

5.5

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

2.8k205Jupyter Notebook

ovg-project/kvcached

5.8

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

1.2k138Python

nobodywho-ooo/nobodywho

6.5

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

1.1k77Rust

jjang-ai/mlxstudio

5.6

MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)

96065

zhihu/ZhiLight

5.1

A highly optimized LLM inference acceleration engine for Llama and its variants.

908104C++

openinfer-project/openinfer

6.5

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

669103Rust

pegainfer-project/pegainfer

6.5

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

669103Rust

andrewkchan/yalm

3.6

Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O

59664C++

zjhellofss/KuiperLLama

3.9

校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。

571143C++

zengxiao-he/tessera

4.3

From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

5629Python

zhongkaifu/TensorSharp

6.0

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

39639C#

interestingLSY/swiftLLM

3.8

A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).

33337Python

avifenesh/memra

3.6

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

OpenEdge ABL