Skip to main content
buildradar
Sign in
Topic · evaluation

evaluation

Tracked open-source repos tagged evaluation, sorted by stars.

Repos
80
Total stars
291,588
Avg. stars
3,645
Share
0.02%

Topics that frequently appear alongside evaluation on the same repo.

Recent risers

Repos created in the last 90 days, tagged evaluation.

  • fable-method@Sahir619

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    2,267
  • Tracely-ai@Jwuthri

    Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

    1,159
  • HarnessEval-W@MirroS-Lab

    HarnessEval-W: Agentifying the Evaluation of Visual Worlds

    256
  • EvoTrace@jinzijian

    Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.

    160
  • langfuse@langfuse

    🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

    33,909+372Star change over the last 7 days
  • mlflow@mlflow

    The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

    27,728+114Star change over the last 7 days
  • promptfoo@promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

    24,664+215Star change over the last 7 days
  • opik@comet-ml

    Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

    21,668+143Star change over the last 7 days
  • WeKnora@Tencent

    Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

    20,899+581Star change over the last 7 days
  • ragas@vibrantlabsai

    Supercharge Your LLM Application Evaluations 🚀

    15,541+123Star change over the last 7 days
  • oumi@oumi-ai

    Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

    9,380+3Star change over the last 7 days
  • opencompass@open-compass

    OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

    7,374+48Star change over the last 7 days
  • helicone@Helicone

    🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

    6,116+23Star change over the last 7 days
  • coze-loop@coze-dev

    Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.

    5,706+7Star change over the last 7 days
  • AutoRAG@Marker-Inc-Korea

    AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.

    5,057+7Star change over the last 7 days
  • Kiln@Kiln-AI

    Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

    5,037+7Star change over the last 7 days
  • lmms-eval@EvolvingLMMs-Lab

    One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

    4,383+11Star change over the last 7 days
  • VLMEvalKit@open-compass

    Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

    4,362+10Star change over the last 7 days
  • evo@MichaelGrupp

    Python package for the evaluation of odometry and SLAM

    4,310+7Star change over the last 7 days
  • langwatch@langwatch

    The platform for LLM evaluations and AI agent testing

    3,517+13Star change over the last 7 days
  • mteb@embeddings-benchmark

    MTEB: State-of-the-art evaluation of embeddings across languages and modalities

    3,409+9Star change over the last 7 days
  • evalscope@modelscope

    A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

    3,331+47Star change over the last 7 days
  • SuperCLUE@CLUEbenchmark

    SuperCLUE: 中文通用大模型综合性基准 | A Benchmark for Foundation Models in Chinese

    3,299-1Star change over the last 7 days
  • lmnr@lmnr-ai

    Laminar - open-source observability platform purpose-built for AI agents. YC S24.

    3,210+24Star change over the last 7 days
  • klipse@viebel

    Klipse is a JavaScript plugin for embedding interactive code snippets in tech blogs.

    3,134+0Star change over the last 7 days
  • ChainForge@ianarawjo

    An open-source visual programming environment for battle-testing prompts to LLMs.

    3,028+1Star change over the last 7 days
  • yao-meta-skill@yaojingang

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    2,557+160Star change over the last 7 days
  • lighteval@huggingface

    Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

    2,530+8Star change over the last 7 days
  • evaluate@huggingface

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    2,480+2Star change over the last 7 days
  • uptrain@uptrain-ai

    UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

    2,364+3Star change over the last 7 days
  • graph-rag-agent@1517005260

    拼好RAG:手搓并融合了GraphRAG、LightRAG、Neo4j-llm-graph-builder进行知识图谱构建以及搜索;整合DeepSearch技术实现私域RAG的推理;自制针对GraphRAG的评估框架| Integrate GraphRAG, LightRAG, and Neo4j-llm-graph-builder for knowledge graph construction and search. Combine DeepSearch for private RAG reasoning. Create a custom evaluation framework for GraphRAG.

    2,330+8Star change over the last 7 days
  • fable-method@Sahir619

    The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

    2,267+23Star change over the last 7 days
  • inspector@MCPJam

    Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

    2,177+23Star change over the last 7 days
  • 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥

    2,165+2Star change over the last 7 days
← Back to topics