The platform around the engine.

Where agents, tasks, knowledge, and your repositories become one governed system — all of it declared as contracts, all of it auditable.

Agents are configuration

The manifest: persona, boundaries, budget — one reviewed file.

An agent is a typed, schema-pinned entity in registry/agents/. The five-piece contract set — Go types + loader, YAML, JSON schema, contract tests, docs — means a manifest that doesn’t validate doesn’t exist:

registry/agents/developer.yaml — simplified
name: developer
runtime: cli # cli · docker · http
provider: local-dev # provider registry reference
model: qwen3-coder-30b # must exist on the provider
persona: implementer # engine-owned base prompt + persona
prompt_mode: minimal
goal: ship the reviewed change
sandbox: workspace-write # restricted · workspace-write
mcp_servers: [aio-tools] # default-deny: only these cross the wire
knowledge:
scopes: [project, fleet] # wiki read scope, enforced at the tool
env_allowlist: [HOME, PATH] # everything else is absent
progress_timeout: 10m # silence past this = named kill
  • Persona is mandatory. The engine renders the system prompt from its own base prompt plus the manifest’s persona — the agent brings skills, never instructions.
  • Granting is a registry edit, reviewed like code — runtime, model, sandbox, env, tools, knowledge scopes are all diffs.
  • Contracts sweep every cross-reference: manifest→provider→model must resolve, and declared capacity slots must agree with the provider’s parallelism — disagreement fails CI, not production.

A2A first

Agents are citizens, not subprocesses.

Every agent advertises an A2A AgentCard (name, skills, capabilities) published to a shared KV store at worker boot, and answers on governed subjects — aio.a2a.discovery, aio.a2a.card.<agent>, aio.a2a.send.<agent>, aio.a2a.tasks.get.<task-id>. A peer — or the console — can discover the fleet and hand work to a named agent over the bus, with no bespoke integration per harness.

A2A speaks HTTP by specification — AiOverload keeps the dialect but not the transport state. Inbound sends are wrapped into ordinary runs: an A2A message becomes run.started, step.dispatched, gate.decided — dispatched, gated, and judged as events on the log, with task status mirrored back to the caller. The wire stays compatible with the protocol; the state is event-sourced. That is what makes high availability a deployment detail instead of a rewrite: when the fleet outgrows one box, redundancy and failover ride JetStream — not HTTP connection affinity.

Honest scope, again: ad-hoc A2A sends become single-turn runs today — same gate as workflow steps, input-required on hold, rejected on block — while cancel, multi-turn, and message dedup are tracked work. The resident runtime’s assistant persona will ride this wire with session-continuous memory on top.

Define once

Provider, agent, platform — three declarations, then reality.

Everything the engine serves is declared, reviewed like code, and validated by contract tests. The walkthrough, in three files:

1. The provider — what serves the model, and on what terms:

registry/providers/local-dev.yaml — simplified
id: local-dev
protocol: openai # wire dialect
hosting: self-hosted # self-hosted · cloud
management: external # the platform loads models on demand
models:
- id: qwen3-coder-30b
context_length: 131072 # capacity is a contract
parallelism: 2 # declared concurrency slots
single_model: true # switching exclusivity, enforced

2. The agent — persona, access, tools (the manifest above): who it is, what it may touch, what it may call.

3. The platform — where it runs. Today: subprocess (cli) and Docker, per the platform-profiles decision — one container contract, two drivers. Kubernetes is the main goal: the same contract served by a K8s driver (pods, PVC-backed workspaces, no Docker-in-Docker), engaged deliberately, not bolted on.

1 · provider.yaml what serves the model capacity · concurrency cloud or your GPU box 2 · agent.yaml persona · sandbox · env tools (deny-by-default) knowledge scopes 3 · runtime cli subprocess · docker kubernetes — the goal one contract, N drivers the engine renders the reality — settings, prompt, allowlist, gateway key
  1. provider.yaml — what serves the model: capacity, concurrency, cloud or your GPU box.
  2. agent.yaml — who the agent is: persona, sandbox, env allowlist, tools (deny-by-default), knowledge scopes.
  3. runtime — cli subprocess or docker today; kubernetes is the goal. One contract, N drivers.

Intent in git, reality at the boundary. The manifest declares; the sidecar materializes; the agent can reach nothing the registry didn't name.

The task plane

Work with a name, judged by machine.

Tasks are managed work items — event-sourced like runs, with a lifecycle on the task-events stream and a folded state in KV. Two fields make them governable:

  • intent: — required, human-auditable. What this is for, in prose an operator can judge.
  • acceptance: — required, machine-checkable. Verdicts on named steps, non-empty outputs, existing artifacts, green test commands. At least one machine gate, because prose criteria are input to a review step, never the done signal.

Around them: scope declarations (repos, paths_in, paths_out), a priority for queue admission, dependencies between tasks, and epics that roll up their children. The console’s task plane is the operator surface — a picker with intent/scope/acceptance previews, typed forms for the workflow’s declared vars, and lifecycle controls (start, pause, resume, cancel) with every mutation audited as task.* events.

An agent — or, soon, the resident manager — files a task with POST /api/tasks; a human starts it. Filing is audited; execution is governed; nothing auto-dispatches without a gate or an approval.

The console is real

A frontend, honestly — young, imperfect, and first-class.

The operator console is not an afterthought admin page stapled to the API. It is a Vite + Preact + TypeScript SPA served by the engine, with an Orval-generated typed client built from the OpenAPI contract — the same contract the CLI speaks — and live updates over SSE, not polling. The committed dist/ is rebuilt only through a pinned, reviewable container build, and the legacy dashboard was retired only after the new console reached parity with it.

Said with the usual honesty: it is young, parts are still being rebuilt, and it is the surface the roadmap invests in first — compose, watch, govern, learn. But it exists, it is tested at the transport level, and it is governed by the same event log as everything else.

GitOps

The fleet directory: git is the source of truth.

The served fleet converges from a fleet directory — a master repo plus per-project overlays — so the registry you review in git is the registry the engine serves:

The problem: the fleet drifts from the repo — someone hot-fixes a manifest on the box, and three weeks later nobody can say which agents the platform actually runs. The answer: the directory converges the served fleet from git, and drift is loud in both directions.

your repositories GitHub · GitLab webhooks + polling forge ingress signature-verified sink freshness loud reconciler: forge events trigger convergence push · merge · wiki.published → sync + converge → drift re-checked fleet directory master repo (aio-admin/) + per-project .aio/ overlays add-only · closed world var defaults per project served fleet the agents · workflows the engine actually resolves drift, loud in both directions — named, recorded, surfaced
  1. Your repositories — GitHub or GitLab, pushing webhooks or polled where inbound can't reach.
  2. Forge ingress — one signature-verified sink; freshness recorded loudly either way.
  3. Fleet directory — master repo plus per-project .aio/ overlays: add-only, closed world.
  4. Served fleet — the agents and workflows the engine actually resolves, converged from the directory.
  5. Drift, loud both ways — actual-vs-declared is named and surfaced; the reconciler re-converges on forge events.

Specialization without redefinition: overlays are add-only — a colliding entity is a named overlay_collision, never a silent override. Project var defaults consume declared workflow vars only (unknown_project_var is refused). The overlay is closed-world: nothing executable lives in .aio/.

an issue becomes work

GitHub issue #142 — “wire the codex harness” — becomes a governed task: intent: from the body, machine acceptance:, lifecycle writes flowing back to the thread.

a project specializes

A project overlay adds deploy-check.yaml plus var defaults; it’s served at the next sync. A colliding name is a named overlay_collision — never a silent override.

a hand-edit is caught

Someone edits the served registry on the box by hand: the drift is recorded, named, surfaced in the console — and converged back at the next forge event.

Issues first

Work enters through the forge. Decisions live in git.

The task plane’s front door is not a web form — it is your issue tracker. A GitHub issue or a GitLab work item becomes a governed task: the issue body carries the intent:, the acceptance bundle names the machine gates, and lifecycle writes flow back to the issue so the thread your team already reads stays the record it always was.

And the data posture follows the same rule — almost nothing lives only inside the app:

  • The durable artifacts are in git: the fleet directory, manifests, workflows, the wiki. Reviewed, versioned, forkable.
  • The engine’s own memory is the event log — replayable, and the only record of what happened.
  • Everything the console shows — run indexes, approvals queues, usage rollups — is a projection: a rebuildable cache over the log and the directory. Delete it, and the engine re-derives it. If a screen and the sources disagree, the sources win.
  • Your approval words are recorded on the log and surfaced back to the issue — the decision is trackable where the work was asked for.

This is the anti-“black box” architecture: governance you can replay — every decision an event, every artifact in git.

Knowledge plane

A fleet wiki agents may read — within scope, with provenance.

Long-term memory is not a prompt appendix; it is a governed read model over a knowledge base that lives in your repos — parsed in the open Knowledge Format (OKF v0.2): one concept per page, typed frontmatter — howto, reference, decision, pitfall, convention, index — strict for our bundle, tolerant for foreign ones, every breach named.

Agents reach it through exactly two scoped tools — knowledge/search (metadata and pointers) and knowledge/get (one page) — over request/reply on the bus:

  • access is manifest-declared and opt-in: knowledge: scopes: [project, fleet] — no block, no access (knowledge_scopes_not_declared);
  • a request beyond the declared scope is refused at the tool boundary (knowledge_scope_refused) — never an empty result pretending nothing exists;
  • scope is enforced at the fold too: a project page in the master wiki is a named scope_containment_violation, recorded and skipped;
  • every page carries provenance, so “where did the agent learn that?” has an answer — and sessions distill into knowledge, so what the fleet learns survives the conversation that found it.

The same read model powers the console’s wiki surface — humans and agents read the same pages, through gates sized for each. And because the pages are files in the master wiki/ tree (or a project’s declared wiki_path), the knowledge base is reviewed, versioned, and forked like everything else — the wiki is the repo.

Forge-native

GitHub and GitLab, first-class — git stays the truth.

The engine meets your repos where they are, both directions:

  • Inbound, two producers one sink: signature-verified webhooks where the internet reaches, outbound polling (GitHub issue-events with ETag conditional requests; GitLab events API with watermarks) where it doesn’t. Same normalized event vocabulary, same ingress, freshness loudly recorded either way.
  • Outbound: review loops open PRs and MRs through the gateway; issues become governed tasks with intent: and acceptance:; CI signals feed cascade triggers.
  • The loop closes in git: forge events trigger directory syncs, the reconciler converges the served fleet, and drift is recorded named and loud. There is one source of truth and it is the repository — the engine is its executor, never its rival.

Telemetry

If you can't observe it, you can't govern it.

Every layer emits — the engine, the reflex gate, the agents. The LGTM stack ships in the compose file, provisioned and pre-wired with shared trace IDs.

Metrics

Prometheus end to end: runs, holds, gate verdicts, latencies — plus provisioned alert rules (a hold/block surge pages you) with a debugging runbook. Grafana dashboards are provisioned, not screenshots — aioverload-overview loads on first boot.

Logs

Loki, structured and queryable per run and per step. “What did step.3 actually see?” is a query, not an archaeology project.

Traces

Tempo + OpenTelemetry: one trace ID follows a step across the engine, the reflex gate, and the agent sidecar — no correlation guesswork.

Decision telemetry

Every pre-dispatch gate decision is recorded as a gate.decided event and landed in reflex/decisions.jsonl. A control-plane labeler pairs each decision with the run’s outcome — justified, false_positive, confirmed, negated (blocks are honestly unobservable: a veto causes the failure, nothing can refute it) — and aio_gate_* gauges report veto precision, escalation-justified rate, and routing accuracy per workflow — and, on the tool path, allow/hold precision per tool on the console’s gate-quality page. Gate quality is measured, not asserted — and the decision log doubles as fine-tuning data for the next gate.

Deep-dive docs publish with the source at release.

Golden workflows

Recipes that ship in the registry.

Workflows are commented YAML in the registry — the procedure is immutable, the agent bindings are governed, and every skip is on the log. The ones worth reading first:

judge-loop

A step runs, a judge scores it against its declared goal, the loop decides: iterate, hold, or block.

dev-review-loop

Code change → review → machine-readable verdict. The reflex variant lets the gate hold risky dispatches.

code-review-cascade

A failed review fans out governed child reviews; failures cascade into a postmortem run.

tier-routed-dev

Work is routed across model tiers — cheap local drafts, stronger models for review — by declared policy.

research-fanout

A question fans out to parallel researchers over an http runtime and consolidates — no local model required.

prospecting-crew

The non-code proof: the gate is expected to hold the outreach decision until a human approves it.

manager-epic decomposes an epic into governed children, subworkflow-demo runs a workflow inside a workflow, and echo-demo needs no AI at all — it’s the first run you should execute once the repository is public.