benchmark
Tracked open-source repos tagged benchmark, sorted by stars.
- #91
Monocular Depth Estimation Toolbox based on MMSegmentation.
★ 971+0Star change over the last 7 days - #92
BlazeHTTP 是一款简单易用的 WAF 防护效果测试工具。BlazeHTTP stands as a user-friendly WAF protection efficacy evaluation tool.
★ 952+0Star change over the last 7 days - #93
Various gRPC benchmarks
★ 939+1Star change over the last 7 days - #94★ 933+0Star change over the last 7 days
- #95★ 925+1Star change over the last 7 days
- #96
Measure Amazon S3's performance from any location.
★ 910-1Star change over the last 7 days - #97
A PyTorch library for all things Reinforcement Learning (RL) for Combinatorial Optimization (CO)
★ 900+1Star change over the last 7 days - #98★ 894+2Star change over the last 7 days
- #99
Windows, macOS and Android storage (HDD, SSD, RAM) speed testing/performance benchmarking app
★ 871+1Star change over the last 7 days - #100★ 863+0Star change over the last 7 days
- #101
Natural Intelligence is still a pretty good idea.
★ 833+0Star change over the last 7 days - #102★ 833+4Star change over the last 7 days
- #103
📊 Benchmark Comparison of Packages with Runtime Validation and TypeScript Support
★ 830+3Star change over the last 7 days - #104
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
★ 826+7Star change over the last 7 days - #105
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
★ 823+0Star change over the last 7 days - #106★ 817+2Star change over the last 7 days
- #107★ 806+27Star change over the last 7 days
- #108
Yet another implementation of computer language benchmarks game
★ 800-1Star change over the last 7 days - #109★ 796+0Star change over the last 7 days
- #110
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training
★ 786+9Star change over the last 7 days - #111
RobustBench: a standardized adversarial robustness benchmark [NeurIPS 2021 Benchmarks and Datasets Track]
★ 782+0Star change over the last 7 days - #112★ 782+2Star change over the last 7 days
- #113★ 780+0Star change over the last 7 days
- #114
An agent benchmark with tasks in a simulated software company.
★ 775+4Star change over the last 7 days - #115
golang HTTP stress testing tool, support single and distributed, http/1, http/2 and http/3.
★ 771+0Star change over the last 7 days - #116
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
★ 747+4Star change over the last 7 days - #117
Another benchmark for some python frameworks
★ 721+0Star change over the last 7 days - #118
Raw benchmarks on throughput, latency and transfer of Hello World on popular microservices frameworks
★ 715+0Star change over the last 7 days - #119
A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper
★ 705+1Star change over the last 7 days - #120
Frontier Models playing the board game Diplomacy.
★ 699+1Star change over the last 7 days