evaluation-framework
Tracked open-source repos tagged evaluation-framework, sorted by stars.
Related topics
Topics that frequently appear alongside evaluation-framework on the same repo.
Recent risers
Repos created in the last 90 days, tagged evaluation-framework.
No new repos tagged with this topic in the last 90 days.
- #1
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
★ 24,664+215Star change over the last 7 days - #2★ 17,953+184Star change over the last 7 days
- #3
A framework for few-shot evaluation of language models.
★ 13,827+89Star change over the last 7 days - #4
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
★ 5,040+7Star change over the last 7 days - #5
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
★ 2,533+8Star change over the last 7 days - #6
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
★ 2,037-2Star change over the last 7 days - #7
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
★ 1,890+89Star change over the last 7 days - #8
Repo for AI Agents The Definitive Guide
★ 1,528+422Star change over the last 7 days - #9
AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
★ 627+2Star change over the last 7 days - #10
Data-Driven Evaluation for LLM-Powered Applications
★ 517+1Star change over the last 7 days