STACKQUADRANT

raketenkater/llm-server

Model Serving

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

5.1
GitHub Metrics
Stars
267
Forks
15
Open Issues
Watchers
4
Contributors
3
Weekly Commits
8
Language
Go
License
MIT
Last Commit
Aug 27, 2026
Created
Mar 11, 2026
Latest Release
v3.2.8
Release Date
Aug 21, 2026
Synced: Aug 28, 2026
Quality Scores
Documentation Qualityw: 20%
3.9

No dedicated docs site. Description: 275 chars. Stars signal: 267. Contributors: 3. Score: 3.9/10

Community Healthw: 20%
2.7

Stars: 267. Contributors: 3. Watchers: 4. Forks: 15. Issue ratio: 0.0%. Score: 2.7/10

Maintenance Velocityw: 15%
9.1

Last commit: 0d ago. Weekly commits: 8. Latest release: v3.2.8. Score: 9.1/10

API Design & DXw: 20%
7.1

Stars/issues ratio: 267. Typed language: Go. No dedicated API docs. Permissive license: MIT. Popularity signal: 267 stars. Score: 7.1/10

Production Readinessw: 15%
3.5

Battle-tested: 267 stars. Peer review: 3 contributors. Versioned: v3.2.8. Licensed: MIT. Age: 0.5 years. Maintenance: last commit 0d ago. Score: 3.5/10

Ecosystem Integrationw: 10%
4.7

Fork interest: 15. Major ecosystem: Go. Integration-friendly: MIT. Adoption: 267 stars. Score: 4.7/10

Tags
cudaggufgolanginference-serverllama-cppllamacppllmlocal-llmlocalllamametal
Radar
Documentation Quality
Community Health
Maintenance Velocity
API Design & DX
Production Readiness
Ecosystem Integration