Skip to main content
buildradar
Sign in
Topic · ai-safety

ai-safety

Tracked open-source repos tagged ai-safety, sorted by stars.

Repos
18
Total stars
41,106
Avg. stars
2,284
Share
0.00%

Topics that frequently appear alongside ai-safety on the same repo.

Recent risers

Repos created in the last 90 days, tagged ai-safety.

  • agent-safe-pipeline@decionis

    Reference architecture for AI agents that propose actions but cannot authorize them — immutable intent capture, an independent Decionis policy verdict (ALLOW/ESCALATE/BLOCK), verified human approval, and a SafeExecutor that consumes a single-use intent-bound grant.

    533
  • floe-guard@Floe-Labs

    The spend meter and budget gate for AI voice agents. Meters STT + TTS + LLM + telephony per call, out of the box (Pipecat, LiveKit — Python & TypeScript). Hard-stops the next turn before it crosses your ceiling. Local, no account, no telemetry. Built by Floe — cost controls for Voice AI.

    434
  • iFixAi@ifixai-ai

    Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

    12,643+1,247Star change over the last 7 days
  • AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.

    6,180+32Star change over the last 7 days
  • AI-Infra-Guard@Tencent

    A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

    6,117+59Star change over the last 7 days
  • A curated list of awesome responsible machine learning resources.

    4,065+2Star change over the last 7 days
  • valqore@valqore

    Safety-first guardrails for AI-driven cloud and Kubernetes operations

    1,921+48Star change over the last 7 days
  • safe-rlhf@PKU-Alignment

    Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

    1,613+1Star change over the last 7 days
  • cc-safety-net@kenryu42

    A pre-execution guard for AI coding agents. It blocks destructive Git and file system commands, plus common attempts to access sensitive files, before a tool call runs. Supports Amp Code, Antigravity CLI, Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot CLI, Grok Build, Hermes Agent, Kimi Code, OpenClaw, OpenCode, and Pi.

    1,521+9Star change over the last 7 days
  • MOSS-RLHF@OpenLMLab

    Secrets of RLHF in Large Language Models Part I: PPO

    1,430+0Star change over the last 7 days
  • uqlm@cvs-health

    [JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"

    1,194+1Star change over the last 7 days
  • Evidence-backed use cases, prompts, integrations, evaluations, and safety notes for OpenAI GPT-6 Astra.

    1,049+7Star change over the last 7 days
  • AgentLens@ZhangJinHaHaHa

    Agentlens is a trusted agent trading platform. Here, you can quickly find the Agent that meets your needs, and you can also publish your own Agent to turn it into your digital asset. We encourage everyone to transform their areas of expertise into Agents and turn them into digital assets, allowing others to see your unique strengths.

    1,035+0Star change over the last 7 days
  • h5i@h5i-dev

    Fast, security-first headless browser for AI agents. Pure Rust, no Chromium or V8. Built for scraping, web testing, and red teaming, with direct HTTP traffic control and auditable sessions.

    623+48Star change over the last 7 days
  • A resource repository for machine unlearning in large language models

    623+0Star change over the last 7 days
  • langtest@PacificAI

    Deliver safe & effective language models

    559+0Star change over the last 7 days
  • PostTrainBench@aisa-group

    Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

    548+9Star change over the last 7 days
  • agent-safe-pipeline@decionis

    Reference architecture for AI agents that propose actions but cannot authorize them — immutable intent capture, an independent Decionis policy verdict (ALLOW/ESCALATE/BLOCK), verified human approval, and a SafeExecutor that consumes a single-use intent-bound grant.

    533+3Star change over the last 7 days
  • PromptInject@agencyenterprise

    PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022

    522+2Star change over the last 7 days
  • cordum@cordum-io

    The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.

    502Star change over the last 7 days
← Back to topics