Skip to main content
buildradar
Sign in
Topic · evaluation

evaluation

Tracked open-source repos tagged evaluation, sorted by stars.

80 repos
  • RAG-FiT@IntelLabs

    Framework for enhancing LLMs for RAG tasks using fine-tuning.

    768-1Star change over the last 7 days
  • LightCompress@ModelTC

    [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

    747+1Star change over the last 7 days
  • opik-openclaw@comet-ml

    🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.

    725+0Star change over the last 7 days
  • iou-tracker@bochinski

    Python implementation of the IOU Tracker

    702+0Star change over the last 7 days
  • ranx@AmenRa

    ⚡️A Blazing-Fast Python Library for Ranking Evaluation, Comparison, and Fusion 🐍

    697+4Star change over the last 7 days
  • long-form-factuality@google-deepmind

    Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".

    693+2Star change over the last 7 days
  • ClawBench@TIGER-AI-Lab

    Open-source benchmark for browser AI agents on daily tasks.

    657+40Star change over the last 7 days
  • Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

    656-2Star change over the last 7 days
  • autoprompt@ucinlp

    AutoPrompt: Automatic Prompt Construction for Masked Language Models.

    641+1Star change over the last 7 days
  • TrustLLM@HowieHwong

    [ICML 2024] TrustLLM: Trustworthiness in Large Language Models

    632+2Star change over the last 7 days
  • A Simple Math and Pseudo C# Expression Evaluator in One C# File. Can also execute small C# like scripts

    630+0Star change over the last 7 days
  • A resource repository for machine unlearning in large language models

    623+0Star change over the last 7 days
  • tcexam@tecnickcom

    TCExam is a CBA (Computer-Based Assessment) system (e-exam, CBT - Computer Based Testing) for universities, schools and companies, that enables educators and trainers to author, schedule, deliver, and report on surveys, quizzes, tests and exams.

    619+0Star change over the last 7 days
  • simpleeval@danthedeckie

    Simple Safe Sandboxed Extensible Expression Evaluator for Python

    611+0Star change over the last 7 days
  • image-matching-toolbox@GrumpyZhou

    This is a toolbox repository to help evaluate various methods that perform image matching from a pair of images.

    594-1Star change over the last 7 days
  • MMMU@MMMU-Benchmark

    This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

    593+1Star change over the last 7 days
  • text2sql-data@jkkummerfeld

    A collection of datasets that pair questions with SQL queries.

    588+1Star change over the last 7 days
  • ParseBench@run-llama

    ParseBench - A Document Parsing Benchmark for AI Agents

    563+6Star change over the last 7 days
  • Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.

    550+3Star change over the last 7 days
  • Dataset and benchmark for RAG on company internal documents.

    542+9Star change over the last 7 days
← Back to topics