STACKQUADRANT

Model Serving

Platforms for deploying and serving ML/AI models at scale

42 repos

mosecorg/mosec

6.3

A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

90273Python

pipeless-ai/pipeless

5.0

An open-source computer vision framework to build and deploy apps in minutes

85252Rust

bentoml/Yatai

5.9

Model Deployment at Scale on Kubernetes 🦄️

84076TypeScript

ServerlessLLM/ServerlessLLM

5.5

Serverless LLM Serving for Everyone.

71176Python

kossisoroyce/timber

5.2

Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.

68723Python

eightBEC/fastapi-ml-skeleton

4.4

FastAPI Skeleton App to serve machine learning models production-ready.

60393Python

underneathall/pinferencia

4.7

Python + Inference - Model Deployment library in Python. Simplest model inference server ever.

54383Python

ome-projects/ome

6.5

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

50294Go

AI-Hypercomputer/JetStream

4.8

JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).

45667Python

intel/xFasterTransformer

4.2

xFasterTransformer — open-source AI/LLM project.

43576C++

NVIDIA/gpu-rest-engine

3.7

A REST API for Caffe using Docker and Go

42295C++

Lightning-Universe/stable-diffusion-deploy

4.6

Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.

39138Python

Epistates/pmetal

4.8

PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.

31026Rust

containers/podman-desktop-extension-ai-lab

5.7

Work with LLMs on a local environment using containers

29883TypeScript

BMW-InnovationLab/BMW-YOLOv4-Inference-API-GPU

4.1

This is a repository for an nocode object detection inference API using the Yolov3 and Yolov4 Darknet framework.

27568Python

raketenkater/ggrun

5.1

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

26816Go

raketenkater/llm-server

4.9

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

26816Go

BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU

3.9

This is a repository for an nocode object detection inference API using the Yolov4 and Yolov3 Opencv.

21759Python
Model Serving — AI/LLM Repositories — StackQuadrant