evals
標記 evals 主題、收錄中的開源專案,依星數排序。
相關主題
常跟 evals 一起出現在同一個專案上的主題。
近期新秀
近 90 天內建立、標記 evals 主題的專案。
- #1
Waku Waku!Waku Agent 是一個由您真正擁有的本地優先 AI Agent 框架,包含迴圈、記憶與評估,所有程式碼皆以易讀性為核心構建,隨專案成長仍保持清晰。
★ 1,649 - #2
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
★ 1,087 - #3
一個精選且無廢話的最佳 AI 代理建構與評估資源庫——論文、部落格、演講、工具、基準測試。由 BenchFlow 維護。
★ 865 - #4
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, and customer craft. Runs on the free Groq API.
★ 614 - #5★ 160
- #1★ 27,657+96近 7 天星數變化
- #2★ 11,298+55近 7 天星數變化
- #3
Python SDK,用於 AI 代理監控、LLM 成本追蹤、效能基準測試等,支援多數 LLM 與代理框架,包括 CrewAI、Agno、OpenAI Agents SDK、Langchain、Autogen、AG2、CamelAI
★ 5,808+3近 7 天星數變化 - #4★ 5,042+5近 7 天星數變化
- #5★ 4,884+120近 7 天星數變化
- #6★ 4,450+4近 7 天星數變化
- #7★ 3,530+2近 7 天星數變化
- #8★ 3,219+9近 7 天星數變化
- #9
專為建構生產環境 AI 系統與評估(evals)的工程師設計的 AI 系統設計指南。
★ 2,978+59近 7 天星數變化 - #10
將任何工作流程轉化為可在 17 個平台上安裝的重複使用式 AI 代理技能 — Claude Code、Copilot、Cursor、Windsurf、Codex、Gemini、Kiro 等。一個 SKILL.md,適用所有平台。
★ 2,371+20近 7 天星數變化 - #11★ 2,183+6近 7 天星數變化
- #12
AI 代理框架的可觀測性與執行強制。透過策略強制捕獲每次執行與執行階段可靠性。內建 40 種策略、本地儀表板,無需帳戶並提供大方的免費雲端方案
★ 1,719+203近 7 天星數變化 - #13★ 1,673+3近 7 天星數變化
- #14
Waku Waku!Waku Agent 是一個由您真正擁有的本地優先 AI Agent 框架,包含迴圈、記憶與評估,所有程式碼皆以易讀性為核心構建,隨專案成長仍保持清晰。
★ 1,649+41近 7 天星數變化 - #15★ 1,539+1近 7 天星數變化
- #16★ 1,200+1近 7 天星數變化
- #17
針對 AI agents 的追蹤原生 CI/CD — 生產環境失敗轉化為阻擋 PR 的迴歸測試。自動偵測、分群、封裝為隔離案例,在 CI 中以 0 美元重播。
★ 1,176-22近 7 天星數變化 - #18★ 1,125+63近 7 天星數變化
- #19
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
★ 1,087+181近 7 天星數變化 - #20
一個能證明自身技能運作正常的技能建立器。為 Claude Code 與 Codex 提供證據驅動的技能建立:基準測試生成、各技能迴歸評估、生態系統檢查、跨執行階段編譯,以及選用的主動式建議工具。
★ 888+2近 7 天星數變化 - #21
一個精選且無廢話的最佳 AI 代理建構與評估資源庫——論文、部落格、演講、工具、基準測試。由 BenchFlow 維護。
★ 865+18近 7 天星數變化 - #22
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, and customer craft. Runs on the free Groq API.
★ 614—近 7 天星數變化