Skip to main content
buildradar
Sign in
Topic · llm-evaluation

llm-evaluation

Tracked open-source repos tagged llm-evaluation, sorted by stars.

37 repos
  • agent-skills-eval@darkrishabh

    A test runner for agentskills.io-style AI agent skills

    724+16Star change over the last 7 days
  • ClawBench@TIGER-AI-Lab

    Open-source benchmark for browser AI agents on daily tasks.

    664+54Star change over the last 7 days
  • Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

    656-2Star change over the last 7 days
  • Awesome papers involving LLMs in Social Science.

    648+1Star change over the last 7 days
  • SkillCorpus@EverMind-AI

    Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

    594Star change over the last 7 days
  • Dataset and benchmark for RAG on company internal documents.

    545+8Star change over the last 7 days
  • continuous-eval@relari-ai

    Data-Driven Evaluation for LLM-Powered Applications

    517+1Star change over the last 7 days
← Back to topics