evaluation
Tracked open-source repos tagged evaluation, sorted by stars.
- #61★ 768-1Star change over the last 7 days
- #62
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
★ 747+1Star change over the last 7 days - #63
🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.
★ 725+0Star change over the last 7 days - #64
Python implementation of the IOU Tracker
★ 702+0Star change over the last 7 days - #65★ 697+4Star change over the last 7 days
- #66
Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".
★ 693+2Star change over the last 7 days - #67★ 657+40Star change over the last 7 days
- #68
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
★ 656-2Star change over the last 7 days - #69
AutoPrompt: Automatic Prompt Construction for Masked Language Models.
★ 641+1Star change over the last 7 days - #70★ 632+2Star change over the last 7 days
- #71
A Simple Math and Pseudo C# Expression Evaluator in One C# File. Can also execute small C# like scripts
★ 630+0Star change over the last 7 days - #72
A resource repository for machine unlearning in large language models
★ 623+0Star change over the last 7 days - #73
TCExam is a CBA (Computer-Based Assessment) system (e-exam, CBT - Computer Based Testing) for universities, schools and companies, that enables educators and trainers to author, schedule, deliver, and report on surveys, quizzes, tests and exams.
★ 619+0Star change over the last 7 days - #74
Simple Safe Sandboxed Extensible Expression Evaluator for Python
★ 611+0Star change over the last 7 days - #75
This is a toolbox repository to help evaluate various methods that perform image matching from a pair of images.
★ 594-1Star change over the last 7 days - #76
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
★ 593+1Star change over the last 7 days - #77
A collection of datasets that pair questions with SQL queries.
★ 588+1Star change over the last 7 days - #78
ParseBench - A Document Parsing Benchmark for AI Agents
★ 563+6Star change over the last 7 days - #79
Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
★ 550+3Star change over the last 7 days - #80
Dataset and benchmark for RAG on company internal documents.
★ 542+9Star change over the last 7 days