Skip to main content
buildradar
Sign in
Topic · evaluation-metrics

evaluation-metrics

Tracked open-source repos tagged evaluation-metrics, sorted by stars.

Repos
12
Total stars
39,801
Avg. stars
3,317
Share
0.00%

Topics that frequently appear alongside evaluation-metrics on the same repo.

Recent risers

Repos created in the last 90 days, tagged evaluation-metrics.

  • eval@frontier-harness-eval

    Public results and task definitions for FrontierHarness Eval

    160
  • deepeval@confident-ai

    The LLM Evaluation Framework

    18,063+110Star change over the last 7 days
  • agentops@AgentOps-AI

    Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI

    5,808+3Star change over the last 7 days
  • tiny-universe@datawhalechina

    《大模型白盒子构建指南》:一个全手搓的Tiny-Universe

    5,040+9Star change over the last 7 days
  • lighteval@huggingface

    Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

    2,533+3Star change over the last 7 days
  • evaluation-guidebook@huggingface

    Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

    2,143+2Star change over the last 7 days
  • AB3DMOT@xinshuoweng

    (IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

    1,846+0Star change over the last 7 days
  • jiwer@jitsi

    Evaluate your speech-to-text system with similarity measures such as word error rate (WER)

    928+1Star change over the last 7 days
  • OCTIS@MIND-Lab

    OCTIS: Comparing Topic Models is Simple! A python package to optimize and evaluate topic models (accepted at EACL2021 demo track)

    804+0Star change over the last 7 days
  • COMET@Unbabel

    A Neural Framework for MT Evaluation

    778+1Star change over the last 7 days
  • ranx@AmenRa

    ⚡️A Blazing-Fast Python Library for Ranking Evaluation, Comparison, and Fusion 🐍

    697+4Star change over the last 7 days
  • :chart_with_upwards_trend: Implementation of eight evaluation metrics to access the similarity between two images. The eight metrics are as follows: RMSE, PSNR, SSIM, ISSM, FSIM, SRE, SAM, and UIQ.

    647+0Star change over the last 7 days
  • continuous-eval@relari-ai

    Data-Driven Evaluation for LLM-Powered Applications

    517+1Star change over the last 7 days
← Back to topics