Skip to main content
buildradar
Sign in
Topic · benchmark

benchmark

Tracked open-source repos tagged benchmark, sorted by stars.

158 repos
  • Monocular Depth Estimation Toolbox based on MMSegmentation.

    971+0Star change over the last 7 days
  • blazehttp@chaitin

    BlazeHTTP 是一款简单易用的 WAF 防护效果测试工具。BlazeHTTP stands as a user-friendly WAF protection efficacy evaluation tool.

    952+0Star change over the last 7 days
  • grpc_bench@LesnyRumcajs

    Various gRPC benchmarks

    939+1Star change over the last 7 days
  • agoo@ohler55

    A High Performance HTTP Server for Ruby

    933+0Star change over the last 7 days
  • nvbench@NVIDIA

    CUDA Kernel Benchmarking Library

    925+1Star change over the last 7 days
  • s3-benchmark@dvassallo

    Measure Amazon S3's performance from any location.

    910-1Star change over the last 7 days
  • rl4co@ai4co

    A PyTorch library for all things Reinforcement Learning (RL) for Combinatorial Optimization (CO)

    900+1Star change over the last 7 days
  • bencher@bencherdev

    🐰 Bencher - Continuous Benchmarking

    894+2Star change over the last 7 days
  • Windows, macOS and Android storage (HDD, SSD, RAM) speed testing/performance benchmarking app

    871+1Star change over the last 7 days
  • Celero@DigitalInBlue

    C++ Benchmark Authoring Library/Framework

    863+0Star change over the last 7 days
  • human-learn@koaning

    Natural Intelligence is still a pretty good idea.

    833+0Star change over the last 7 days
  • OpenCUA@xlang-ai

    [NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents

    833+4Star change over the last 7 days
  • 📊 Benchmark Comparison of Packages with Runtime Validation and TypeScript Support

    830+3Star change over the last 7 days
  • deep_research_bench@Ayanami0730

    DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

    826+7Star change over the last 7 days
  • ClawProBench@suyoumo

    ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.

    823+0Star change over the last 7 days
  • warp@minio

    S3 benchmarking tool

    817+2Star change over the last 7 days
  • agentdojo@ethz-spylab

    A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.

    806+27Star change over the last 7 days
  • Yet another implementation of computer language benchmarks game

    800-1Star change over the last 7 days
  • sbt-jmh@sbt

    "Trust no one, bench everything." - sbt plugin for JMH (Java Microbenchmark Harness)

    796+0Star change over the last 7 days
  • mobilegym@Purewhiter

    [EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training

    786+9Star change over the last 7 days
  • robustbench@RobustBench

    RobustBench: a standardized adversarial robustness benchmark [NeurIPS 2021 Benchmarks and Datasets Track]

    782+0Star change over the last 7 days
  • HammerDB@TPC-Council

    HammerDB: The industry standard open-source database benchmark

    782+2Star change over the last 7 days
  • r3f-perf@utsuboco

    Easily monitor your ThreeJS performances.

    780+0Star change over the last 7 days
  • TheAgentCompany@TheAgentCompany

    An agent benchmark with tasks in a simulated software company.

    775+4Star change over the last 7 days
  • http_bench@linkxzhou

    golang HTTP stress testing tool, support single and distributed, http/1, http/2 and http/3.

    771+0Star change over the last 7 days
  • LightCompress@ModelTC

    [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

    747+4Star change over the last 7 days
  • Another benchmark for some python frameworks

    721+0Star change over the last 7 days
  • Raw benchmarks on throughput, latency and transfer of Hello World on popular microservices frameworks

    715+0Star change over the last 7 days
  • caliper@hyperledger-caliper

    A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper

    705+1Star change over the last 7 days
  • AI_Diplomacy@GoodStartLabs

    Frontier Models playing the board game Diplomacy.

    699+1Star change over the last 7 days
← Back to topics