Evaluation & Testing
Frameworks for evaluating, benchmarking, and testing AI systems
CopilotKit/aimock
6.7
★ 905◇ 65TypeScript
vostride/agent-qa
5.6
★ 897◇ 15TypeScript
benchflow-ai/awesome-evals
5.1
★ 854◇ 89
darkrishabh/agent-skills-eval
5.2
★ 719◇ 35TypeScript
onejune2018/Awesome-LLM-Eval
4.5
★ 656◇ 84
ValueByte-AI/Awesome-LLM-in-Social-Science
5.3
★ 647◇ 52
Pacific-AI-Corp/langtest
6.3
★ 559◇ 52Python
PacificAI/langtest
6.3
★ 559◇ 52Python
faiscadev/fakecloud
6.1
★ 537◇ 40Rust
relari-ai/continuous-eval
5.6
★ 517◇ 38Python
rhesis-ai/rhesis
5.6
★ 391◇ 34Python
ai-dashboad/flutter-skill
5.7
★ 362◇ 54Dart
JonathanChavezTamales/llm-leaderboard
4.6
★ 356◇ 40JavaScript
palico-ai/palico-ai
4.5
★ 343◇ 31TypeScript
PetroIvaniuk/llms-tools
4.8
★ 327◇ 51
athina-ai/athina-evals
4.0
★ 301◇ 23Python
testdriverai/testdriverai
4.7
★ 240◇ 35JavaScript
PramodDutta/qaskills
4.5
★ 214◇ 23TypeScript
naodeng/awesome-qa-skills
5.3
★ 191◇ 25Python
buer2233/ai-api-test-skill
4.3
★ 142◇ 17Python
blackhaiyu-sudo/spec2case
4.1
★ 118◇ 6Python
petrkindlmann/qa-skills
4.4
★ 103◇ 20Python
Circleoipillar/xmind-vault
2.1
★ 34◇ —
← prev2 / 2