Agents are about to hold the keys. Who governs them?

Every framework answers “how do I build an agent?” Almost nobody answers the question that matters once the demo ends: how do I govern one?

Why now

The missing half isn't smarter actors.

Agents are leaving the chat window. They take keys, push branches, message customers, spend money. An autonomous actor without an audit trail isn’t automation — it’s a liability with an API token. The industry keeps shipping smarter actors; the layer that makes an autonomous actor accountable is the empty one.

The answer, in one sentence: named intents, deny-by-default, every decision in the event log. Here is what a governed step actually looks like:

  1. Named intent

    what this is for, plus machine-checkable acceptance

  2. Reflex gate

    typed questions: route · verdict · risk · escalate

  3. AgentGateway

    JIT credentials, default-deny MCP, wire-level

  4. Agent runtime

    ephemeral sidecar, argv-only, sandboxed

  5. Judge

    did the output meet the declared goal?

go — dispatch hold — a human decides, words on the log block — the run fails loudly

Judge says "not there yet"? The step loops with feedback — or holds for a human, or blocks the run.

every decision lands in the event log — aioctl replay re-verifies the whole run
Framework + glue AiOverload
Access whatever the agent’s env happened to contain agentgateway — JIT credentials minted per run, revoked after
Audit trail logs you grep afterwards the event log is the runtime; aioctl replay re-verifies it
MCP access whatever was configured default-deny: only registry-allowlisted servers, wire-level enforcement
Risky step runs, or crashes, or “helpfully” retries parks as step.held with a reason and a risk score
Reproducibility prompt + hope decisions are a pure function of definition hash + ordered events

Principles

Six rules the engine refuses to break.

Named intents, or nothing runs

Every task declares what it wants and how done is verified — a required intent: and machine-checkable acceptance:. Work without a name doesn’t enter the system, so everything that runs can be judged.

Deny by default

Agents get no ambient access. Credentials are minted per run and revoked at terminus; only registry-allowlisted tools exist on the wire; manifests carry an env allowlist and a sandbox level. Granting is a registry edit, reviewed like code.

The event log is the truth

State is a fold over an append-only log — replayable, reindexable, auditable. If a screen and the log disagree, the log wins.

Loud failures over silent degradation

A held run exits held. A failed step fails the run. Nothing hangs, nothing guesses, nothing “helpfully” retries behind your back — every refusal names itself.

The engine owns the harness boundary

Manifests declare intent — provider, model, prompt mode, sandbox. The engine materializes the reality: settings, system prompt, per-run tool allowlist. Agents can’t reach past what the engine serves them.

Initiative is not impunity

Agents may propose — an always-on manager files tasks for approval; nothing auto-executes without a gate or a human. Initiative is resident; execution stays governed.

Honesty

What we do not claim.

Claude is the battle-tested path — today

Wired: Claude, Codex, Kimi, local models (Qwen, Gemma), any HTTP LLM. Battle-tested in CI today: Claude. Harnesses are added by evidence — nothing gets the “proven” label until its smoke suite passes.

The gate is not the sandbox

The reflex gate is triage — fast judgment over typed questions. The hard execution boundary is the sidecar’s isolation. Anyone selling you a language model as a security boundary is selling you something.

Temporal wins raw durability

If you need decades-proven workflow durability across data centers, use Temporal. AiOverload keeps retry timers in memory; a boot sweep heals interrupted runs, but that’s an honest trade, not a substitute.

SaaS wins zero-setup

If you want something running in five minutes with no infrastructure, a hosted coding agent will beat a self-hosted control plane. AiOverload is for people who need the keys to live in their own registry.

Frameworks win prototype speed

CrewAI or LangGraph will get you a demo faster. AiOverload is for when the demo starts touching real repos, real keys, and real consequences.

Known limits, stated up front — not buried:

  • Retry timers are in memory today; a boot recovery sweep re-drives interrupted judgments, holds, and cascades (durable timers are tracked work).
  • A2A tasks are single-shot: no cancel, no multi-turn yet.
  • The http runtime is not A2A-discoverable yet.
  • The reflex gate is CPU-local and small by design — it is the crew’s reflexes, not its brain.

Loud failures, no silent fallback. A held run exits held. A failed step fails the run. Nothing hangs, nothing guesses, nothing “helpfully” retries behind your back — the console mirrors the engine rule.