STACKQUADRANT

Evaluation & Testing

Frameworks for evaluating, benchmarking, and testing AI systems

47 repos

CopilotKit/aimock

6.7

Mock everything your AI app talks to — LLM APIs, MCP, A2A, vector DBs, search. One package, one port, zero dependencies.

90565TypeScript

vostride/agent-qa

5.6

The self-improving Agentic QA harness with Memory. Write tests in natural language.
 Catch regressions before releases ship.

89715TypeScript

benchflow-ai/awesome-evals

5.1

A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.

85489

darkrishabh/agent-skills-eval

5.2

A test runner for agentskills.io-style AI agent skills

71935TypeScript

onejune2018/Awesome-LLM-Eval

4.5

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

65684

ValueByte-AI/Awesome-LLM-in-Social-Science

5.3

Awesome papers involving LLMs in Social Science.

64752

Pacific-AI-Corp/langtest

6.3

Deliver safe & effective language models

55952Python

PacificAI/langtest

6.3

Deliver safe & effective language models

55952Python

faiscadev/fakecloud

6.1

Free, open-source AWS emulator. LocalStack alternative: 26 services, 1,924 operations, 100% conformance. No account, no auth token, no paid tier.

53740Rust

relari-ai/continuous-eval

5.6

Data-Driven Evaluation for LLM-Powered Applications

51738Python

rhesis-ai/rhesis

5.6

The testing platform for AI teams. Bring engineers, PMs, and domain experts together to generate tests, simulate (adversarial) conversations, and trace every failure to its root cause.

39134Python

ai-dashboad/flutter-skill

5.7

AI-powered E2E testing for 10 platforms. 253 MCP tools. Zero config. Works with Claude, Cursor, Windsurf, Copilot. Test Flutter, React Native, iOS, Android, Web, Electron, Tauri, KMP, .NET MAUI — all from natural language.

36254Dart

JonathanChavezTamales/llm-leaderboard

4.6

A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)

35640JavaScript

palico-ai/palico-ai

4.5

Build, Improve Performance, and Productionize your LLM Application with an Integrated Framework

34331TypeScript

PetroIvaniuk/llms-tools

4.8

A list of LLMs Tools & Projects

32751

athina-ai/athina-evals

4.0

Python SDK for running evaluations on LLM generated responses

30123Python

testdriverai/testdriverai

4.7

Computer-Use SDK for E2E QA Testing

24035JavaScript

PramodDutta/qaskills

4.5

QA Skills Directory QA Skills is a curated directory of testing-specific skills for AI coding agents (Claude Code, Cursor, Copilot, etc.).

21423TypeScript

naodeng/awesome-qa-skills

5.3

Awesome QA Skills — a bilingual (zh/en) AI testing Agent Skills library for Codex, Cursor, Claude Code, Kiro, OpenCode, and Trae. Ships 4 testing workflows and 25 testing-type skills (58 skill folders with language parity): independently installable, composable, and eval-ready with skill-up. Covers requirements, strategy, cases, API/performance/sec

19125Python

buer2233/ai-api-test-skill

4.3

AI接口自动化测试 Skill:面向 Python + pytest + requests,驱动 Codex / Claude Code 生成、维护和调试接口用例(AI API test automation skill for Python + pytest + requests)

14217Python

blackhaiyu-sudo/spec2case

4.1

Spec2Case 是生产级 AI 测试用例生成智能体,支持图片/文本需求理解、人工确认、LangGraph 流程编排和 Excel 用例导出。

1186Python

petrkindlmann/qa-skills

4.4

50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime.

10320Python

Circleoipillar/xmind-vault

2.1

XMind Vault

34