llm-evaluation
標記 llm-evaluation 主題、收錄中的開源專案,依星數排序。
相關主題
常跟 llm-evaluation 一起出現在同一個專案上的主題。
近期新秀
近 90 天內建立、標記 llm-evaluation 主題的專案。
- #1
一個精選且無廢話的最佳 AI 代理建構與評估資源庫——論文、部落格、演講、工具、基準測試。由 BenchFlow 維護。
★ 865 - #2
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
★ 646 - #3
Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible epochs to converge on optimization of a goal.
★ 352 - #4
Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
★ 160
- #1
🪢 開源 AI 工程平台:LLM 評估、可觀測性、指標、提示管理、遊樂場、資料集。整合 OpenTelemetry、LangChain、OpenAI SDK、LiteLLM 等。🍊YC W23
★ 34,116+207近 7 天星數變化 - #2★ 27,783+55近 7 天星數變化
- #3
測試你的提示、代理與檢索增強生成,進行 AI 紅隊/滲透測試/漏洞掃描。比較 GPT、Claude、Gemini、DeepSeek 等表現。使用簡潔宣告式設定,支援命令列與 CI/CD 整合。OpenAI 與 Anthropic 均有使用。
★ 24,767+103近 7 天星數變化 - #4★ 21,752+84近 7 天星數變化
- #5★ 18,063+110近 7 天星數變化
- #6★ 12,643+1,247近 7 天星數變化
- #7★ 11,298+55近 7 天星數變化
- #8★ 9,094+23近 7 天星數變化
- #9
非线智能 NoneLinear - ReLE评测:中文 AI 大模型能力评测(持续更新):目前已囊括374个大模型,覆盖 chatgpt、gpt-5.4、谷歌 gemini-3.1-pro、Claude-4.6、文心 ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤 senseChat 等商用模型,以及 step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱 GLM-5.1、MiMo-V2、LongCat、gemma4、mistral 等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。
★ 6,415+6近 7 天星數變化 - #10★ 6,127+11近 7 天星數變化
- #11
全棧 AI Red Teaming 平台,通過 Agent Scan、Skills Scan、MCP Scan、AI Infra Scan 以及 LLM jailbreak 評估來保護 AI 生態系統。
★ 6,117+59近 7 天星數變化 - #12
🐢 開源的 LLM 代理評估與測試函式庫
★ 5,801+9近 7 天星數變化 - #13
Agent OS:代理自行變得更聰明。我們只負責介面:等級指令與預期結果不會寫入成功合約。面試篩選、階段評估、預算演化迴圈。MCP 伺服器,支援 13 種執行環境:Claude Code、Codex CLI、Gemini CLI、OpenCode、Copilot、Kiro 等。
★ 5,753+25近 7 天星數變化 - #14
LLM 實務指南:從基礎到在 AWS 上部署先進 LLM 與 RAG 應用,採用 LLMOps 最佳實踐
★ 5,310+8近 7 天星數變化 - #15★ 5,058+1近 7 天星數變化
- #16★ 4,390+7近 7 天星數變化
- #17
一套科學方法、流程、演算法與系統,用於構建故事與模型
★ 3,676+2近 7 天星數變化 - #18★ 3,530+2近 7 天星數變化
- #19★ 3,219+9近 7 天星數變化
- #20
生成式 AI 的完整資源,包含詳細的路線圖、專案、使用案例、面試準備和程式碼練習。
★ 2,610+1近 7 天星數變化 - #21
Agentic LLM 漏洞掃描器 / AI 紅隊測試套件 🧪
★ 1,983+3近 7 天星數變化 - #22
開源的端到端平台,用於評估、觀察與改進 LLM 及 AI agent 應用程式。包含追蹤、評估、模擬、資料集、閘道器、安全護欄。支援自行架構。採用 Apache 2.0 授權。
★ 1,909+42近 7 天星數變化 - #23★ 1,641-1近 7 天星數變化
- #24★ 1,573+4近 7 天星數變化
- #25
Prompty 讓您能輕鬆為 AI 應用程式建立、管理、除錯及評估 LLM prompt。Prompty 是專為 LLM prompt 設計的資產類別與格式,旨在提升開發者的可觀測性、可理解性與可攜性。
★ 1,255+1近 7 天星數變化 - #26★ 1,194+1近 7 天星數變化
- #27
針對 AI agents 的追蹤原生 CI/CD — 生產環境失敗轉化為阻擋 PR 的迴歸測試。自動偵測、分群、封裝為隔離案例,在 CI 中以 0 美元重播。
★ 1,176-22近 7 天星數變化 - #28★ 1,060+2近 7 天星數變化
- #29
具備高韌性工作流程、證據佐證 RAG、版本化報告與自動化品質評估的 AI 股票研究 Agent。
★ 1,032-1近 7 天星數變化 - #30
一個精選且無廢話的最佳 AI 代理建構與評估資源庫——論文、部落格、演講、工具、基準測試。由 BenchFlow 維護。
★ 865+18近 7 天星數變化