OpenAPPA is a deterministic AI guardrail system designed to prevent data exfiltration from prompt injection and model hallucination without breaking agent behavior. It’s open, vendor‑agnostic, MIT‑licensed, and reports benchmark results showing 89% task completion with zero successful attacks (versus Claude Auto mode at 90% completion and 10% attacks, and FIDES at 41%/31%). The core claim is that probabilistic second‑model judges and simple regex or rule blacklists fail in practice because they cannot track data flow across tool calls, are themselves prompt‑injectable, or become so restrictive they disrupt agents.
Instead of matching patterns, OpenAPPA tracks labeled data flows across an agent’s execution graph using a declarative configuration (appa.toml) that specifies data sources, audiences, trust levels, and authorities. Every trajectory carries an audience×trust label that only tightens, and decisions are derived algebraically so prompts cannot subvert policy. Enforcement returns machine‑readable remedy plans - sanitizers to redact secrets or PII, one‑off authorities to approve specific actions, or disposable subagents to isolate risky reads - allowing agents to continue useful work. The system is pluggable into existing agent loops, auditable in CI/CD, and built to scale across many agents.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.