Skip to main content
buildradar
Sign in
Topic · benchmark

benchmark

Tracked open-source repos tagged benchmark, sorted by stars.

158 repos
  • RoboTwin@RoboTwin-Platform

    [ICML 2026] RoboTwin 2.0 Offical Code Repo

    2,800+12Star change over the last 7 days
  • java-sec-code@JoyChou93

    Java web common vulnerabilities and security code which is base on springboot and spring security

    2,675+1Star change over the last 7 days
  • CPU-X@TheTumultuousUnicornOfDarkness

    CPU-X is a Free software that gathers information on CPU, motherboard and more

    2,647+0Star change over the last 7 days
  • tsung@processone

    Tsung is a high-performance benchmark framework for various protocols including HTTP, XMPP, LDAP, etc.

    2,629+0Star change over the last 7 days
  • little-coder@itayinbarr

    A harness optimized to smaller LLMs

    2,528+16Star change over the last 7 days
  • mitata@evanwashere

    benchmark tooling that loves you ❤️

    2,525+4Star change over the last 7 days
  • InternVideo@OpenGVLab

    [ECCV2024] Video Foundation Models & Data for Multimodal Understanding

    2,374+5Star change over the last 7 days
  • tinybench@tinylibs

    🔎 A simple, tiny and lightweight benchmarking library!

    2,362+0Star change over the last 7 days
  • ecs@oneclickvirt

    VPS Fusion Monster Server Test GO Version Aiming to be the most comprehensive server testing project, implemented in Go with zero environment dependencies. VPS融合怪服务器测评项目 GO版本 尽量成为最全能的服务器测评项目,使用 Go 实现,无需任何环境依赖。

    2,329+13Star change over the last 7 days
  • beir@beir-cellar

    A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.

    2,283+5Star change over the last 7 days
  • LIBERO@Lifelong-Robot-Learning

    Benchmarking Knowledge Transfer in Lifelong Robot Learning

    2,264+14Star change over the last 7 days
  • cista@felixguendling

    Cista is a simple, high-performance, zero-copy C++ serialization & reflection library.

    2,247+1Star change over the last 7 days
  • 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥

    2,164-1Star change over the last 7 days
  • :zap: Go web framework benchmark

    2,135+1Star change over the last 7 days
  • An objective comparison of multiple frameworks that allow us to "transform" our web apps to desktop applications.

    1,989+0Star change over the last 7 days
  • logparser@logpai

    A machine learning toolkit for log parsing [ICSE'19, DSN'16]

    1,987+0Star change over the last 7 days
  • tapnet@google-deepmind

    Tracking Any Point (TAP)

    1,973+2Star change over the last 7 days
  • tau2-bench@sierra-research

    τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

    1,938+30Star change over the last 7 days
  • less_slow.cpp@ashvardanian

    Playing around "Less Slow" coding practices in C++ 20, C, CUDA, PTX, & Assembly, from numerics & SIMD to coroutines, ranges, exception handling, networking and user-space IO

    1,923-1Star change over the last 7 days
  • evalplus@evalplus

    Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

    1,806+1Star change over the last 7 days
  • training@mlcommons

    Reference implementations of MLPerf® training benchmarks

    1,771-1Star change over the last 7 days
  • VBench@Vchitect

    [CVPR2024 Highlight] VBench - We Evaluate Video Generation

    1,757+7Star change over the last 7 days
  • skillsbench@benchflow-ai

    SkillsBench evaluates how well skills work and how effective agents are at using them.

    1,742+11Star change over the last 7 days
  • nanobench@martinus

    Simple, fast, accurate single-header microbenchmarking functionality for C++11/14/17/20

    1,733+1Star change over the last 7 days
  • hotpath-rs@pawurb

    Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.

    1,707+7Star change over the last 7 days
  • BEHAVIOR-1K@StanfordVL

    BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx

    1,677+8Star change over the last 7 days
  • inference@mlcommons

    Reference implementations of MLPerf® inference benchmarks

    1,624+2Star change over the last 7 days
  • The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

    1,608-1Star change over the last 7 days
  • ADR@uber

    ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

    1,527+21Star change over the last 7 days
  • py-motmetrics@cheind

    :bar_chart: Benchmark multiple object trackers (MOT) in Python

    1,487+0Star change over the last 7 days
← Back to topics