evaluation
Tracked open-source repos tagged evaluation, sorted by stars.
- #31
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
★ 2,143+2Star change over the last 7 days - #32★ 2,088+0Star change over the last 7 days
- #33
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
★ 2,037-2Star change over the last 7 days - #34
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
★ 2,012-1Star change over the last 7 days - #35
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
★ 1,846+0Star change over the last 7 days - #36
WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.
★ 1,786+1Star change over the last 7 days - #37
心理健康大模型 (LLM x Mental Health), Pre & Post-training & Dataset & Evaluation & Depoly & RAG, with InternLM / Qwen / Baichuan / DeepSeek / Mixtral / LLama / GLM series models
★ 1,781+1Star change over the last 7 days - #38
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
★ 1,760+12Star change over the last 7 days - #39★ 1,668+0Star change over the last 7 days
- #40
The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
★ 1,608-1Star change over the last 7 days - #41★ 1,506+0Star change over the last 7 days
- #42
Evaluation code for various unsupervised automated metrics for Natural Language Generation.
★ 1,391+0Star change over the last 7 days - #43★ 1,300+0Star change over the last 7 days
- #44★ 1,260+0Star change over the last 7 days
- #45
A framework for comprehensive diagnosis and optimization of agents using simulated, realistic synthetic interactions
★ 1,257+0Star change over the last 7 days - #46
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
★ 1,227+7Star change over the last 7 days - #47
Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.
★ 1,224+0Star change over the last 7 days - #48★ 1,206+0Star change over the last 7 days
- #49
High-fidelity performance metrics for generative models in PyTorch
★ 1,200+1Star change over the last 7 days - #50
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
★ 1,176-22Star change over the last 7 days - #51
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
★ 1,156+2Star change over the last 7 days - #52★ 1,155+7Star change over the last 7 days
- #53
Evaluate your LLM's response with Prometheus and GPT4 💯
★ 1,111+4Star change over the last 7 days - #54
LangSmith Client SDK Implementations
★ 1,049+8Star change over the last 7 days - #55
SemanticKITTI API for visualizing dataset, processing data, and evaluating results.
★ 898+3Star change over the last 7 days - #56
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
★ 885+3Star change over the last 7 days - #57
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
★ 823+0Star change over the last 7 days - #58★ 819+12Star change over the last 7 days
- #59★ 814+1Star change over the last 7 days
- #60
Web Codegen Scorer is a tool for evaluating the quality of web code generated by LLMs.
★ 775+4Star change over the last 7 days