Topic · llm-evaluation
llm-evaluation
Tracked open-source repos tagged llm-evaluation, sorted by stars.
37 repos
- #31
A test runner for agentskills.io-style AI agent skills
★ 724+16Star change over the last 7 days - #32★ 664+54Star change over the last 7 days
- #33
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
★ 656-2Star change over the last 7 days - #34
Awesome papers involving LLMs in Social Science.
★ 648+1Star change over the last 7 days - #35
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
★ 594—Star change over the last 7 days - #36
Dataset and benchmark for RAG on company internal documents.
★ 545+8Star change over the last 7 days - #37
Data-Driven Evaluation for LLM-Powered Applications
★ 517+1Star change over the last 7 days