STACKQUADRANT

Model Serving

Platforms for deploying and serving ML/AI models at scale

42 repos

jundot/omlx

8.2

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

21.2k1.8kPython

TensorRT-LLM

7.2

TensorRT-LLM — a leading open-source project in the AI/LLM ecosystem.

14.5k2.7kPython

vllm-project/vllm-omni

7.5

A framework for efficient model inference with omni-modality models

6.5k1.6kPython

beclab/Olares

7.3

Olares: An Open-Source Personal Cloud to Reclaim Your Data

5.2k321Go

ahkarami/Deep-Learning-in-Production

4.5

In this repository, I will share some useful notes and references about deploying deep learning-based models in production.

4.4k685

HuaizhengZhang/AI-Infra-from-Zero-to-Hero

6.1

🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for Machine Learning, LLM (Large Language Model), GenAI (Generative AI). 🍻 OSDI, NSDI, SIGCOMM, SoCC, MLSys, etc. 🗃️ Llama3, Mistral, etc. 🧑‍💻 Video Tutorials.

4.3k409

ModelTC/LightLLM

7.0

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

4.3k359Python

containers/ramalama

7.5

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

3.0k361Python

thu-pacman/chitu

7.0

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

3.0k258Python

superlinked/sie

6.9

Superlinked Inference Engine is an Open-source inference server and production cluster for embeddings, reranking, and extraction.

2.9k294Python

vllm-project/vllm-ascend

7.6

Community maintained hardware plugin for vLLM on Ascend

2.7k2.2kC++

roboflow/inference

7.2

Turn any computer or edge device into a command center for your computer vision projects.

2.4k312Python

tensorchord/envd

6.7

🏕️ Reproducible development environment for humans and agents

2.2k168Go

microsoft/aici

4.8

AICI: Prompts as (Wasm) Programs

2.1k86Rust

mlrun/mlrun

7.1

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.

1.7k318Python

waybarrios/vllm-mlx

6.8

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

1.6k219Python

kitops-ml/kitops

7.0

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

1.4k185Go

alibaba/rtp-llm

6.1

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

1.3k270Cuda

logicalclocks/hopsworks

5.8

Hopsworks - Data-Intensive AI platform with a Feature Store

1.3k160Java

basetenlabs/truss

6.9

The simplest way to serve AI/ML models in production

1.2k122Python

aiptimizer/TurboOCR

6.0

Fast GPU OCR server. 270 img/s on FUNSD. TensorRT FP16, PP-OCRv5, HTTP + gRPC.

1.0k100C++

sgl-project/sglang-omni

6.6

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

1.0k410Python

efeslab/Nanoflow

4.6

A throughput-oriented high-performance serving framework for LLMs

97452Jupyter Notebook

openvinotoolkit/model_server

6.5

A scalable inference server for models optimized with OpenVINO™

923274C++