evals
标记 evals 主题、收录中的开源项目,按星标数排序。
相关主题
常跟 evals 一起出现在同一个项目上的主题。
近期新秀
近 90 天内创建、标记 evals 主题的项目。
- #1
Waku Waku!Waku Agent 是一个由您真正拥有的本地优先 AI Agent 框架,包含循环、记忆与评估,所有代码均以易读性为核心构建,随项目成长仍保持清晰。
★ 1,649 - #2
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
★ 1,087 - #3
一个精选且无废话的最佳 AI 代理构建与评估资源库——论文、博客、演讲、工具、基准测试。由 BenchFlow 维护。
★ 865 - #4
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, and customer craft. Runs on the free Groq API.
★ 617 - #5★ 160
- #1★ 27,657+96近 7 天星标变化
- #2★ 11,298+55近 7 天星标变化
- #3
Python SDK,用于 AI 代理监控、LLM 成本追踪、基准测试等,集成多数 LLM 和代理框架,包括 CrewAI、Agno、OpenAI Agents SDK、Langchain、Autogen、AG2、CamelAI
★ 5,808+3近 7 天星标变化 - #4★ 5,042+5近 7 天星标变化
- #5★ 4,884+120近 7 天星标变化
- #6★ 4,450+4近 7 天星标变化
- #7★ 3,530+2近 7 天星标变化
- #8★ 3,219+9近 7 天星标变化
- #9
专为构建生产环境 AI 系统与评估(evals)的工程师设计的 AI 系统设计指南。
★ 2,978+59近 7 天星标变化 - #10
将任何工作流程转化为可在 17 个平台上安装的可复用 AI 智能体技能 — Claude Code、Copilot、Cursor、Windsurf、Codex、Gemini、Kiro 等。一个 SKILL.md,适用所有平台。
★ 2,371+20近 7 天星标变化 - #11★ 2,183+6近 7 天星标变化
- #12
AI 代理框架的可观测性与执行强制。通过策略强制捕获每次运行与运行时可靠性。内置 40 种策略、本地仪表板,无需账户并提供慷慨的免费云端方案
★ 1,719+203近 7 天星标变化 - #13★ 1,673+3近 7 天星标变化
- #14
Waku Waku!Waku Agent 是一个由您真正拥有的本地优先 AI Agent 框架,包含循环、记忆与评估,所有代码均以易读性为核心构建,随项目成长仍保持清晰。
★ 1,649+41近 7 天星标变化 - #15★ 1,539+1近 7 天星标变化
- #16★ 1,200+1近 7 天星标变化
- #17
针对 AI agents 的追踪原生 CI/CD — 生产环境失败转化为阻挡 PR 的回归测试。自动检测、分群、封装为隔离案例,在 CI 中以 0 美元重放。
★ 1,176-22近 7 天星标变化 - #18★ 1,125+75近 7 天星标变化
- #19
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
★ 1,087+190近 7 天星标变化 - #20
一个能证明自身技能运作正常的技能构建器。为 Claude Code 与 Codex 提供证据驱动的技能构建:基准测试生成、各技能回归评估、生态系统检查、跨运行时编译,以及选用的主动式建议工具。
★ 888+3近 7 天星标变化 - #21
一个精选且无废话的最佳 AI 代理构建与评估资源库——论文、博客、演讲、工具、基准测试。由 BenchFlow 维护。
★ 865+18近 7 天星标变化 - #22
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, agents (loop from scratch, tool design, guardrails, MCP, Skills), fine-tuning vs LoRA, prompt-injection/security, LLMOps, and customer craft. Runs on the free Groq API.
★ 617—近 7 天星标变化