STACKQUADRANT

avifenesh/memra

Inference Engines

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

6.2
GitHub Metrics
Stars
327
Forks
38
Open Issues
2
Watchers
25
Contributors
7
Weekly Commits
250
Language
Rust
License
MIT
Last Commit
Aug 28, 2026
Created
Jul 5, 2026
Latest Release
v0.117.0
Release Date
Aug 28, 2026
Synced: Aug 28, 2026
Quality Scores
Documentation Qualityw: 20%
6.1

Has docs site (https://inference.tiyuvta.ai). Description: 265 chars. Stars signal: 327. Contributors: 7. Score: 6.1/10

Community Healthw: 20%
4.6

Stars: 327. Contributors: 7. Watchers: 25. Forks: 38. Issue ratio: 0.6%. Score: 4.6/10

Maintenance Velocityw: 15%
9.7

Last commit: 0d ago. Weekly commits: 250. Latest release: v0.117.0. Score: 9.7/10

API Design & DXw: 20%
7.9

Stars/issues ratio: 164. Typed language: Rust. Has documentation site. Permissive license: MIT. Popularity signal: 327 stars. Score: 7.9/10

Production Readinessw: 15%
3.5

Battle-tested: 327 stars. Peer review: 7 contributors. Versioned: v0.117.0. Licensed: MIT. Age: 0.1 years. Maintenance: last commit 0d ago. Score: 3.5/10

Ecosystem Integrationw: 10%
5.3

Fork interest: 38. Major ecosystem: Rust. Integration-friendly: MIT. Adoption: 327 stars. Has web presence. Score: 5.3/10

Tags
blackwellcudagemmaggufgpu-kernelsinference-enginellmllm-inferencellm-servingmoe
Radar
Documentation Quality
Community Health
Maintenance Velocity
API Design & DX
Production Readiness
Ecosystem Integration