benchmark
Tracked open-source repos tagged benchmark, sorted by stars.
- #121
Point based and tiny object detection and localization code set of UCAS-VG
★ 694+1Star change over the last 7 days - #122
Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".
★ 693+2Star change over the last 7 days - #123
A benchmarking framework for the Julia language
★ 683+0Star change over the last 7 days - #124
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
★ 682+2Star change over the last 7 days - #125★ 664+54Star change over the last 7 days
- #126
CAIRI Supervised, Semi- and Self-Supervised Visual Representation Learning Toolbox and Benchmark
★ 658+0Star change over the last 7 days - #127
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
★ 656-2Star change over the last 7 days - #128
BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to quickly compare different MARL algorithms, tasks, and models while being systematically grounded in its two core tenets: reproducibility and standardization.
★ 656+2Star change over the last 7 days - #129
The first-ever vast natural language processing benchmark for Indonesian Language. We provide multiple downstream tasks, pre-trained IndoBERT models, and a starter code! (AACL-IJCNLP 2020)
★ 655+1Star change over the last 7 days - #130
A repository of pretty cool datasets that I collected for network science and machine learning research.
★ 655+0Star change over the last 7 days - #131
A.S.E (AICGSecEval) is a repository-level AI-generated code security evaluation benchmark developed by Tencent Wukong Code Security Team.
★ 655+0Star change over the last 7 days - #132
Kotlin multiplatform benchmarking toolkit
★ 644+0Star change over the last 7 days - #133★ 632+2Star change over the last 7 days
- #134
AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
★ 629+1Star change over the last 7 days - #135
Federated Learning Benchmark - Federated Learning on Non-IID Data Silos: An Experimental Study (ICDE 2022)
★ 618+0Star change over the last 7 days - #136
Performance testing matchers for RSpec
★ 618+0Star change over the last 7 days - #137
A Benchmark of Text Classification in PyTorch
★ 608+0Star change over the last 7 days - #138★ 604+0Star change over the last 7 days
- #139
A global network of probes to run network tests like ping, traceroute and DNS resolve
★ 597+3Star change over the last 7 days - #140
Visual Object Tracking
★ 596+1Star change over the last 7 days - #141
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
★ 594—Star change over the last 7 days - #142★ 588+2Star change over the last 7 days
- #143★ 580+1Star change over the last 7 days
- #144
ParseBench - A Document Parsing Benchmark for AI Agents
★ 565+10Star change over the last 7 days - #145
Serialization library written in C++17 - Pack C++ structs into a compact byte-array without any macros or boilerplate code
★ 560+0Star change over the last 7 days - #146
Naive performance comparison of a few programming languages (JavaScript, Kotlin, Rust, Swift, Nim, Python, Go, Haskell, D, C++, Java, C#, Object Pascal, Ada, Lua, Ruby)
★ 557+1Star change over the last 7 days - #147
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 553+3Star change over the last 7 days - #148
Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
★ 550+3Star change over the last 7 days - #149
Dataset and benchmark for RAG on company internal documents.
★ 545+8Star change over the last 7 days - #150
Performance benchmarking and testing framework for .NET applications :chart_with_upwards_trend:
★ 540+0Star change over the last 7 days