跳到主要內容
buildradar
Sign in
主題 · evaluation-framework

evaluation-framework

標記 evaluation-framework 主題、收錄中的開源專案,依星數排序。

專案數
10
總星數
70,935
平均星數
7,094
佔比
0.00%

常跟 evaluation-framework 一起出現在同一個專案上的主題。

近期新秀

近 90 天內建立、標記 evaluation-framework 主題的專案。

  • eval@frontier-harness-eval

    Public results and task definitions for FrontierHarness Eval

    154
  • promptfoo@promptfoo

    測試你的提示、代理與檢索增強生成,進行 AI 紅隊/滲透測試/漏洞掃描。比較 GPT、Claude、Gemini、DeepSeek 等表現。使用簡潔宣告式設定,支援命令列與 CI/CD 整合。OpenAI 與 Anthropic 均有使用。

    24,767+103近 7 天星數變化
  • deepeval@confident-ai

    LLM 評估框架

    18,063+110近 7 天星數變化
  • lm-evaluation-harness@EleutherAI

    用於少樣本評估語言模型的框架

    13,873+46近 7 天星數變化
  • Kiln@Kiln-AI

    構建、評估與優化 AI 系統。包含評估、RAG、代理、微調、合成數據生成、資料集管理、MCP 等。

    5,042+5近 7 天星數變化
  • lighteval@huggingface

    Lighteval 是您跨多種後端評估 LLM 的全方位工具包

    2,533+3近 7 天星數變化
  • EvalAI@Cloud-CV

    :cloud: :rocket: :bar_chart: :chart_with_upwards_trend: 評估 AI 的最新技術水準

    2,037-2近 7 天星數變化
  • future-agi@future-agi

    開源的端到端平台,用於評估、觀察與改進 LLM 及 AI agent 應用程式。包含追蹤、評估、模擬、資料集、閘道器、安全護欄。支援自行架構。採用 Apache 2.0 授權。

    1,909+42近 7 天星數變化
  • ai-agents-the-definitive-guide@Nicolepcx

    AI Agents 終極指南儲存庫

    1,567+220近 7 天星數變化
  • AgentLab@ServiceNow

    AgentLab:一個用於在各種任務上開發、測試與評測 web agent 的開源框架,專為可擴充性與可重現性而設計。

    629+1近 7 天星數變化
  • continuous-eval@relari-ai

    基於 LLM 應用程式的資料驅動評估

    517+1近 7 天星數變化
← 返回主題列表