STACKQUADRANT

Evaluation & Testing

Frameworks for evaluating, benchmarking, and testing AI systems

47 repos

confident-ai/deepeval

8.1

The LLM Evaluation Framework

18.0k1.9kPython

Ragas

7.2

Ragas — a leading open-source project in the AI/LLM ecosystem.

15.6k1.7kPython

ifixai-ai/iFixAi

7.5

The open-source diagnostic for AI misalignment. 32 tests across fabrication, manipulation, deception, unpredictability, and opacity. Provider-agnostic. Runs against OpenAI, Anthropic, Bedrock, Azure, Gemini, and more. Letter grade in under 5 minutes, content-addressed manifest for bit-identical replay. Built by iMe.

12.4k1.3kPython

NVIDIA/garak

7.6

the LLM vulnerability scanner

9.1k1.2kPython

jeinlee1991/chinese-llm-benchmark

6.2

ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括335个大模型,覆盖chatgpt、gpt-5.2、o4-mini、谷歌gemini-3-pro、Claude-4.5、文心ERNIE-X1.1、ERNIE-5.0-Thinking、qwen3-max、百川、讯飞星火、商汤senseChat等商用模型, 以及kimi-k2、ernie4.5、minimax-M2、deepseek-v3.2、qwen3-2507、llama4、智谱GLM-4.6、gemma3、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

6.4k262

Tencent/AI-Infra-Guard

7.6

A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

6.1k568Python

Q00/ouroboros

7.8

Agent OS: Stop prompting. Start specifying. A Socratic interview gates the spec on an ambiguity score, then one command drives execution, a 3-stage evaluation gate, and a budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

5.7k579Python

PacktPublishing/LLM-Engineers-Handbook

6.4

The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices

5.3k1.3kPython

Agenta-AI/agenta

7.7

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

4.7k661TypeScript

EvolvingLMMs-Lab/lmms-eval

7.7

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

4.4k649Python

truera/trulens

7.7

Evaluation and Tracking for LLM Experiments and AI Agents

3.5k335Python

lmnr-ai/lmnr

7.2

Laminar - open-source observability platform purpose-built for AI agents. YC S24.

3.2k229TypeScript

BlazeUp-AI/Observal

6.2

Observal is an AI agent registry with first in class observabilty and eval framework

2.4k471Python

future-agi/future-agi

7.0

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

1.9k573Python

huggingface/aisheets

5.9

Build, enrich, and transform datasets using AI models with no code

1.6k139TypeScript

cyberark/FuzzyAI

5.3

A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM APIs.

1.6k216Jupyter Notebook

bug0inc/passmark

5.8

The open-source Playwright library for AI browser regression testing with intelligent caching, auto-healing, and multi-model verification.

1.3k184TypeScript

microsoft/prompty

6.8

Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.

1.3k126Rust

cvs-health/uqlm

6.7

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

1.2k131Python

Jwuthri/Tracely-ai

5.7

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

1.2k91Python

JudgmentLabs/judgeval

6.8

The open source post-building layer for agents. Our environment data and evals power agent post-training (RL, SFT) and monitoring.

1.1k97Python

juanjuandog/FinSight-AI

5.6

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

1.0k62Java

MGdaasLab/WHartTest

6.4

WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.)

1.0k161Python

langwatch/scenario

6.0

Agentic testing for agentic codebases

96080Python