Skip to main content
buildradar
Sign in
Topic · evaluation-framework

evaluation-framework

Tracked open-source repos tagged evaluation-framework, sorted by stars.

Repos
10
Total stars
70,407
Avg. stars
7,041
Share
0.00%

Topics that frequently appear alongside evaluation-framework on the same repo.

Recent risers

Repos created in the last 90 days, tagged evaluation-framework.

No new repos tagged with this topic in the last 90 days.

  • promptfoo@promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

    24,664+215Star change over the last 7 days
  • deepeval@confident-ai

    The LLM Evaluation Framework

    17,953+184Star change over the last 7 days
  • lm-evaluation-harness@EleutherAI

    A framework for few-shot evaluation of language models.

    13,827+89Star change over the last 7 days
  • Kiln@Kiln-AI

    Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

    5,040+7Star change over the last 7 days
  • lighteval@huggingface

    Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

    2,533+8Star change over the last 7 days
  • EvalAI@Cloud-CV

    :cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI

    2,037-2Star change over the last 7 days
  • future-agi@future-agi

    Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

    1,890+89Star change over the last 7 days
  • ai-agents-the-definitive-guide@Nicolepcx

    Repo for AI Agents The Definitive Guide

    1,528+422Star change over the last 7 days
  • AgentLab@ServiceNow

    AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.

    627+2Star change over the last 7 days
  • continuous-eval@relari-ai

    Data-Driven Evaluation for LLM-Powered Applications

    517+1Star change over the last 7 days
← Back to topics