Skip to main content
buildradar
Sign in
Topic · evaluation

evaluation

Tracked open-source repos tagged evaluation, sorted by stars.

80 repos
  • evaluation-guidebook@huggingface

    Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

    2,143+2Star change over the last 7 days
  • avalanche@ContinualAI

    Avalanche: an End-to-End Library for Continual Learning based on PyTorch.

    2,088+0Star change over the last 7 days
  • EvalAI@Cloud-CV

    :cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI

    2,037-2Star change over the last 7 days
  • alpaca_eval@tatsu-lab

    An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.

    2,012-1Star change over the last 7 days
  • AB3DMOT@xinshuoweng

    (IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

    1,846+0Star change over the last 7 days
  • WFGY@onestardao

    WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.

    1,786+1Star change over the last 7 days
  • EmoLLM@SmartFlowAI

    心理健康大模型 (LLM x Mental Health), Pre & Post-training & Dataset & Evaluation & Depoly & RAG, with InternLM / Qwen / Baichuan / DeepSeek / Mixtral / LLama / GLM series models

    1,781+1Star change over the last 7 days
  • trpc-agent-go@trpc-group

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

    1,760+12Star change over the last 7 days
  • ragbits@deepsense-ai

    Building blocks for rapid development of GenAI applications

    1,668+0Star change over the last 7 days
  • The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

    1,608-1Star change over the last 7 days
  • pycm@sepandhaghighi

    Multi-class confusion matrix library in Python

    1,506+0Star change over the last 7 days
  • nlg-eval@Maluuba

    Evaluation code for various unsupervised automated metrics for Natural Language Generation.

    1,391+0Star change over the last 7 days
  • lispy@abo-abo

    Short and sweet LISP editing

    1,300+0Star change over the last 7 days
  • xai@EthicalML

    XAI - An eXplainability toolbox for machine learning

    1,260+0Star change over the last 7 days
  • intellagent@plurai-ai

    A framework for comprehensive diagnosis and optimization of agents using simulated, realistic synthetic interactions

    1,257+0Star change over the last 7 days
  • KernelBench@ScalingIntelligence

    KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

    1,227+7Star change over the last 7 days
  • kubetorch@run-house

    Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.

    1,224+0Star change over the last 7 days
  • fuzzbench@google

    FuzzBench - Fuzzer benchmarking as a service.

    1,206+0Star change over the last 7 days
  • torch-fidelity@toshas

    High-fidelity performance metrics for generative models in PyTorch

    1,200+1Star change over the last 7 days
  • Tracely-ai@Jwuthri

    Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

    1,176-22Star change over the last 7 days
  • ncalc@ncalc

    NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.

    1,156+2Star change over the last 7 days
  • Gym@NVIDIA-NeMo

    Evaluate and improve models and agents using environments

    1,155+7Star change over the last 7 days
  • prometheus-eval@prometheus-eval

    Evaluate your LLM's response with Prometheus and GPT4 💯

    1,111+4Star change over the last 7 days
  • langsmith-sdk@langchain-ai

    LangSmith Client SDK Implementations

    1,049+8Star change over the last 7 days
  • SemanticKITTI API for visualizing dataset, processing data, and evaluating results.

    898+3Star change over the last 7 days
  • deepfabric@nolabs-ai

    Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

    885+3Star change over the last 7 days
  • ClawProBench@suyoumo

    ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.

    823+0Star change over the last 7 days
  • OpenJudge@agentscope-ai

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    819+12Star change over the last 7 days
  • gval@PaesslerAG

    Expression evaluation in golang

    814+1Star change over the last 7 days
  • web-codegen-scorer@angular

    Web Codegen Scorer is a tool for evaluating the quality of web code generated by LLMs.

    775+4Star change over the last 7 days
← Back to topics