Evaluation & Testing
Frameworks for evaluating, benchmarking, and testing AI systems
confident-ai/deepeval
8.1
★ 18.0k◇ 1.9kPython
Ragas
7.2
★ 15.6k◇ 1.7kPython
ifixai-ai/iFixAi
7.5
★ 12.4k◇ 1.3kPython
NVIDIA/garak
7.6
★ 9.1k◇ 1.2kPython
jeinlee1991/chinese-llm-benchmark
6.2
★ 6.4k◇ 262
Tencent/AI-Infra-Guard
7.6
★ 6.1k◇ 568Python
Q00/ouroboros
7.8
★ 5.7k◇ 578Python
PacktPublishing/LLM-Engineers-Handbook
6.4
★ 5.3k◇ 1.3kPython
Agenta-AI/agenta
7.7
★ 4.7k◇ 660TypeScript
EvolvingLMMs-Lab/lmms-eval
7.7
★ 4.4k◇ 649Python
truera/trulens
7.5
★ 3.5k◇ 335Python
lmnr-ai/lmnr
7.2
★ 3.2k◇ 229TypeScript
BlazeUp-AI/Observal
6.2
★ 2.4k◇ 471Python
future-agi/future-agi
6.7
★ 1.9k◇ 569Python
huggingface/aisheets
6.0
★ 1.6k◇ 139TypeScript
cyberark/FuzzyAI
5.3
★ 1.6k◇ 216Jupyter Notebook
bug0inc/passmark
5.8
★ 1.3k◇ 184TypeScript
microsoft/prompty
6.8
★ 1.3k◇ 126Rust
cvs-health/uqlm
6.7
★ 1.2k◇ 131Python
Jwuthri/Tracely-ai
5.8
★ 1.2k◇ 91Python
JudgmentLabs/judgeval
6.6
★ 1.1k◇ 97Python
juanjuandog/FinSight-AI
5.6
★ 1.0k◇ 62Java
MGdaasLab/WHartTest
6.4
★ 1.0k◇ 161Python
langwatch/scenario
6.2
★ 960◇ 80Python
1 / 2next →