The threat model, without the sales filter.
This page summarizes the engine’s own docs/threat-model.md — the full
document publishes with the source at release. The rule is the site’s
rule: a security claim you have to dig for is marketing.
Containment, graded
What actually contains an agent today.
The boundary: per-run Docker containers
The engine’s real containment layer: on the docker runtime, a run gets an ephemeral container — its filesystem staged copy-in/copy-out, its network a private lane whose only peer is the gateway. When the threat model says “boundary”, this is the layer it means. Stated plainly: the shipped fleet opts into the cli tier today, so this boundary is wired on main but not yet the shipped configuration; container hardening (resource limits, cap-drop) is in progress.
cli 'workspace-write' is a trusted tier
It maps to the harness’s skip-permissions mode. The agent can leave the
task workdir and do anything its uid can — read ~/.ssh, touch the
operator’s other checkouts, reach the network. PID-namespace
confinement scopes process visibility and signals only. The threat
model forbids calling it a sandbox; so does this page.
cli 'restricted' is a harness flag
A read-only-world hint to the harness — weaker than the harness’s own restricted mode, and explicitly not a security boundary. Graded as convenience, not containment.
Unconfined by omission
A cli manifest without a sandbox: level runs unconfined. The registry
ships developer.yaml as workspace-write by design — with the grade
above written next to it, in public, where reviewers look first.
Known gaps, named
What the gate does not cover — today.
Native harness tools bypass the gate
The reflex tool gate judges MCP calls routed through the agentgateway.
A harness’s native tools — Claude Code’s built-in Bash, Edit, Write —
never cross the gateway, so they are not gated. A real incident
(pkill -9 on host fixtures) proved the hole; the native-tools policy
that closes it is tracked work.
Opt-in auth is opt-in
Role-based NATS accounts, bus TLS, API TLS: the hardened profile exists,
is contract-tested, and is not the default. The default dev profile
is loopback and unauthenticated. /metrics stays plaintext either way.
Ad-hoc A2A sends are ungated by default
The reflex gate on ad-hoc sends defaults to none. The send subjects
must be gated before they face untrusted callers — the README’s known
limits say it; this page repeats it.
JIT credentials, scoped
Per-run mint and revoke covers gateway-backed provider runs. The LLM plane through the gateway is opt-in today; Claude Code subscription sessions inherit credentials instead. The full secrets broker is a parked track — claimed nowhere on this site.
The discipline
What holds independent of any layer.
- The event log is the runtime — every decision, approval and refusal
is an event;
aioctl replayre-verifies a run against its definition hash. An audit is a query. - Deny-by-default tools — a manifest’s MCP allowlist is rendered into the run’s config and enforced at the gateway; an undeclared server does not exist on the wire.
- Secrets move through env only — read at the last responsible moment, never logged, never on envelopes; a CI guard test proves tracked files carry neither secrets nor operator paths.
- Scanned continuously — gitleaks on every code event; the full
hermetic ladder (e2e, govulncheck, pip-audit, smoke against a live
NATS stack) on every push to
mainand every merge-ready PR. - The gate fails closed — reflex unreachable at dispatch means every gated step holds, never proceeds ungoverned. The reflex gate itself is triage, graded as such: a small non-generative model scoring typed questions, with deterministic threshold rails doing the deciding.
Reporting
Found something? Tell us.
Security issues are welcome — coordinated disclosure preferred, public
issues fine for anything not exploitable. The machine-readable contact
is in security.txt; the short version:
Contact: johannes.girard@gmail.com · Preferred languages: English, French · Scope: the engine, this site, and anything carrying the AiOverload name.
We will acknowledge within 72 hours and keep you posted on the fix. No legal threats for good-faith research — the threat model exists because we want the holes found.