ai-safety
Tracked open-source repos tagged ai-safety, sorted by stars.
Related topics
Topics that frequently appear alongside ai-safety on the same repo.
Recent risers
Repos created in the last 90 days, tagged ai-safety.
- #1
Reference architecture for AI agents that propose actions but cannot authorize them — immutable intent capture, an independent Decionis policy verdict (ALLOW/ESCALATE/BLOCK), verified human approval, and a SafeExecutor that consumes a single-use intent-bound grant.
★ 533 - #2
The spend meter and budget gate for AI voice agents. Meters STT + TTS + LLM + telephony per call, out of the box (Pipecat, LiveKit — Python & TypeScript). Hard-stops the next turn before it crosses your ceiling. Local, no account, no telemetry. Built by Floe — cost controls for Voice AI.
★ 434
- #1
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
★ 12,643+1,247Star change over the last 7 days - #2
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
★ 6,180+32Star change over the last 7 days - #3
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
★ 6,117+59Star change over the last 7 days - #4
A curated list of awesome responsible machine learning resources.
★ 4,065+2Star change over the last 7 days - #5★ 1,921+48Star change over the last 7 days
- #6
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
★ 1,613+1Star change over the last 7 days - #7
A pre-execution guard for AI coding agents. It blocks destructive Git and file system commands, plus common attempts to access sensitive files, before a tool call runs. Supports Amp Code, Antigravity CLI, Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot CLI, Grok Build, Hermes Agent, Kimi Code, OpenClaw, OpenCode, and Pi.
★ 1,521+9Star change over the last 7 days - #8★ 1,430+0Star change over the last 7 days
- #9
[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
★ 1,194+1Star change over the last 7 days - #10
Evidence-backed use cases, prompts, integrations, evaluations, and safety notes for OpenAI GPT-6 Astra.
★ 1,049+7Star change over the last 7 days - #11
Agentlens is a trusted agent trading platform. Here, you can quickly find the Agent that meets your needs, and you can also publish your own Agent to turn it into your digital asset. We encourage everyone to transform their areas of expertise into Agents and turn them into digital assets, allowing others to see your unique strengths.
★ 1,035+0Star change over the last 7 days - #12
Fast, security-first headless browser for AI agents. Pure Rust, no Chromium or V8. Built for scraping, web testing, and red teaming, with direct HTTP traffic control and auditable sessions.
★ 623+48Star change over the last 7 days - #13
A resource repository for machine unlearning in large language models
★ 623+0Star change over the last 7 days - #14★ 559+0Star change over the last 7 days
- #15
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
★ 548+9Star change over the last 7 days - #16
Reference architecture for AI agents that propose actions but cannot authorize them — immutable intent capture, an independent Decionis policy verdict (ALLOW/ESCALATE/BLOCK), verified human approval, and a SafeExecutor that consumes a single-use intent-bound grant.
★ 533+3Star change over the last 7 days - #17
PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022
★ 522+2Star change over the last 7 days - #18
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
★ 502—Star change over the last 7 days