AGPL-3.0 · self-hosted · public release soon

Agents can read repos, deploy code, call APIs. Who governs them?

Agents are probabilistic. Execution shouldn’t be.

Your agents, one governed door. AiOverload runs agents as deterministic, event-sourced workflows on your infrastructure — behind an agentgateway where the declared tool plane is deny-by-default and credentials are per-run, scored step by step by a small reflex gate, with every decision replayable from the event log.

New, small, and auditable. One maintainer, no traction claims, no “enterprise-grade” adjectives — judge it by the code and by the log.

Get notified at release How it works The vision

aioctl — a risky step parks, a human decides, the log remembers
$ aioctl run registry/workflows/dev-review-loop-reflex.yaml
run.started 9c4e71a8f2b64d1e8a5c3f7b2d1904e dev-review-loop-reflex 8 steps

step.held review_2
risk=0.81 reason=review-approves-a-push-the-gate-cannot-vouch-for

$ aioctl approve 9c4e71a8f2b64d1e8a5c3f7b2d1904e review_2 --input "review ok — do not push"
run.running review_2 re-dispatched with operator input on the log

step.completed report verdict: PASS (machine-readable, on the log)
run.completed 8/8 steps · 1 human approval recorded · 0 silent fallbacks
# illustrative — real capture pending

tool.decide

Every tool call gets the same treatment.

The terminal above shows a step held at dispatch. The gate also sits on the tool path: every MCP tool call routed through the gateway is judged before it happens — and the decision lands on the same log. The named gap, stated rather than buried: a harness’s native tools (Claude Code’s built-in Bash, Edit, Write) do not transit the gateway, so they do not pass the tool gate today — closing that hole is tracked work, and the threat model says so.

tool.decided — scored, on the log
tool.decided  server=github  tool=create_pull_request
             action=hold  risk=0.62
             reason=unattributed-write-to-protected-branch

Every MCP call routed through the gateway gets the same treatment: scored, logged, replayable. Deny is a refusal at the wire — the call never happens; it is not a warning after it.

Decisions are labelled against outcomes. Every gateway tool decision is paired with the observed outcome — an allow is later marked confirmed or negated, a hold justified or a false_positive — and the labels accumulate as training data for a future gate. (No fine-tuning exists today; the labelling loop does.) Static allowlists match tool names; the reflex gate scores the kind of call — from server, tool and argument shape — and fails closed. The loop is built and allow/hold precision is measured per tool; the data accumulates from release, and the published calibration stays re-measurable on your own decision log.

The short version

Six ideas, no magic.

The event log is the runtime

Runs are event-sourced: the log is the state, so any run can be replayed and re-verified decision by decision — determinism is a tested property, not a promise.

Git is the source of truth

The fleet is declared in the repo; tasks arrive from your forge. A manifest is a diff you review like code — what the engine serves is what you merged, and drift is loud in both directions.

A gate, not a chatbot

Every gated step is scored by a small local model that answers typed questions — route, verdict, risk — and fails closed. It never generates text, so it cannot answer off-schema. Calibration is published, not promised.

Boundaries are rendered, not promised

For provider-backed runs through the gateway, credentials are minted when the run starts and revoked when it ends. Only registry-allowlisted tools get through, enforced at the wire — config you review like code. (The subscription-session exception is named on How it works.)

Agents are addressable peers

Agents talk to the engine (and each other) over NATS — every interaction is an event, no direct calls.

Risky steps park for humans

A step the gate cannot vouch for is held with a reason and a risk score — never silently retried. Your approval, with your exact words, lands on the log before anything re-dispatches.

How it works, end to end →

Who it's for

For anyone industrializing work with agents they must control.

Forward deployed engineers, consulting companies, and businesses of every size — anyone who wants their processes run by AI agents and workflows they control, not a black box they rent. The work may be software; it may be any external tool: an API or a webhook through the ingress port, browser automation or a legacy terminal application through a declared MCP target. Wherever a harness can reach through a governed door, the deal is the same — judged steps, held side effects, a replayable log.

Forward Deployed Engineers

You embed in the customer’s business and own the loop from “what should we build” to “it’s live.” AiOverload carries the four categories of your toolkit — orchestration, guardrails, evaluation, observability — as one self-hosted control plane declared in the client’s repo, and still running when you roll off.

Consulting companies

You implement AI for clients and stay answerable after the delivery. Each engagement gets a governed, self-hosted control plane declared in the client’s own repo — agents under a gate, every action replayable, and the audit answer ready before the client asks for it.

IT & compliance

Platform teams, IT departments, CISOs: when an agent pushes a branch or messages a customer, someone must be able to say what happened, who approved it, and prove it. Approvals land on the log with the operator’s exact words; an audit is a query, not an archaeology project.

Regulated sectors

The threat model grades its own containment layers and names its gaps in public — including the ones most vendors bury. Judge the security posture the way your reviewers will.

Operations on business apps

Inbound mail or a CRM record becomes a governed run through the ingress port; the agent acts on the app only through gateway tool targets, and every outbound side effect — often irreversible — parks for a human before it happens.

Developers running several agents

Your agents live in a repo you already review; tasks arrive through the forge you already use; and every run — including the one that went wrong — replays on request. Self-hosted, AGPL-3.0, no seat counter.

Not the target? If you want five-minute setup with no infrastructure, a hosted coding agent will beat a self-hosted control plane — that trade is stated on the honesty list, not hidden in the FAQ.

Operator console

Built. Shipping with the release.

Reading the log alone, you would think AiOverload is CLI-only. It is not: the operator console is built. An approvals queue that parks steps and tool calls, approve/reject with your reason on the log, run detail with the gate-decisions panel, a gate-quality page measuring allow/hold precision per tool, and fleet and project views over the directory read model.

aio-console — approvals
runwhat is parkedrisk
9c4e71a8 step review_2 hold · 0.81 approve · reject
f207bbc1 tool github/create_pull_request hold · 0.62 approve · reject

reason=unattributed-write-to-protected-branch

Your words land on the log with the decision — aioctl replay re-verifies the whole run later.

Operator console: built, shipping with the release. The preview above is a static mock of the real screens — no live data, no embellishment; screenshots publish with the source.

Before the release

Decided on paper, in the repo, before the code.

The roadmap isn’t slideware: significant architecture decisions land as written decision documents — with rejected alternatives on the record — before the implementation slices start. The latest: a resident agent runtime (an always-on manager that watches drift and proposes tasks, plus a console assistant — initiative is resident, execution stays governed). The engine repo is still private before the public release; those documents publish with the source. It’s the first item on the roadmap.

Meanwhile the engine is measured, not asserted: 95 Go packages at the latest measured commit, and a hermetic CI ladder — build, vet, contract schemas, dead-code, secrets, and smoke against a live NATS stack — on every push to main and every PR labelled merge-ready (the fast basic tier runs earlier, unlabelled). Every number a page on this site shows should be checkable — the status page is where they get stamped.

Honest by design

What we do not claim.

One battle-tested harness — so far

Wired today: the Claude Code harness, pointed at the Mistral API or a self-hosted serving box by the provider registry, plus an HTTP runtime for any LLM. Battle-tested in CI today: Claude Code. Codex, Kimi and Gemma harnesses are roadmap, added by evidence — nothing gets the “proven” label until its smoke suite passes.

The gate is not the sandbox

The reflex gate is triage — fast judgment over typed questions. The engine’s containment boundary is the per-run Docker container — and the shipped fleet doesn’t run there yet: today’s agents opt into the cli tier, a trusted tier with PID-namespace confinement at most (none without a declared sandbox level). Making the container runtime the shipped default is in progress. A language model is not a security boundary — and neither is a flag.

Not for every job

Prototype fast with a framework; run zero-setup with a SaaS. AiOverload is for when agents touch real repos, real keys, and real consequences — and the keys live in your registry.

Built mainly with AI — disclosed

Written mainly with Kimi (2.8 Preview, Kimi 3) plus Claude Code, Z.ai and Mistral on review — under one engineer’s direction, and judged by a hermetic CI ladder, not by anyone’s confidence. How it’s built →

The full honesty list →

Why not X?

The honest map.

No single tool does all of these: execute governed, hold a human door on an event log, replay every decision, and apply a rule-based hold/escalate decision on reflex verdicts at fixed thresholds (93.3/90.0/75.6). The neighbors, honestly:

What it does What it does not
Sandboxes · micro-VMs (E2B, Firecracker) isolate execution hold a step on a scored verdict
Workflow engines (Temporal) replayable orchestration write your guardrails for you
LLM gateways · proxies filter the network hold a human door on an event log
Frameworks (LangGraph, CrewAI) orchestrate agents govern beyond configuration
External integration engines deterministic plumbing between systems judgment, a human door, a log of record
Agent control planes (the Paperclip class) run agent teams with approvals and budgets rebuild state from an event log; use your forge as the source of truth

None of them is bad at its job. The closest neighbours are single-process platforms with their own ticket system; AiOverload bets on four structural choices instead: the run is an event-sourced log you can replay and prove, git and the forge are the source of truth, execution is distributed (separate control plane, workers and sidecars), and the license is AGPL-3.0 — nobody can close the engine into a competing hosted service. When the jobs meet at one door, someone has to keep the log.

Keep reading

The overview is the trailer. The pages are the film.

How it works

The event-sourced pipeline, the reflex gate, per-run credentials, the sidecar boundary — every claim drawn, every step on the log.

The engine, end to end →

Platform

Agent manifests and personas, A2A-native agents, the governed task plane, GitOps fleet directory, knowledge plane, forge integration, and the ingress port that puts business apps behind the same gate.

The full surface →

Vision

Agents are about to hold the keys — who governs them? The case for a governed-execution layer, stated as principles, not marketing.

Why this exists →

About

One senior platform engineer from Normandy, building in the open with AI pair-programmers and a hermetic CI ladder. No team, no pitch deck.

Who builds this →

Public release soon.

Self-hosted, AGPL-3.0 — source and docs publish with the release. Leave your address: one email at release. That’s it. Every claim on this site is checkable in the meantime — the status page says what’s shipped, partial, or roadmap, stamped with the engine commit.

Get notified at release See the roadmap