Reliability & ops · inside AJO

Systems.

The part of AJO that prevents and resolves systemwide malfunctions.

Fleet health
nominal
daemon
launcher
alive
daemon
agent_runner
alive
daemon
slack_mirror
alive
dead-letters
0
in backlog
token quota
61% of window
kill switch
clear
armed, one flag away
kill switch quota breaker stuck-claim sweep watchdog dead-letter
all systems nominal · monitor runs every 2h
The monitor sweeps every two hours.

Take the watching layer away and a fleet that "runs itself" only ever runs itself on camera. Build it as its own non-negotiable discipline instead, and the same fleet becomes something you can actually walk away from — because a kill switch stops the fleet from making a single new write, a quota breaker trips before a runaway process gets near the budget, a stuck claim gets swept and re-delivered rather than left stalling in silence, and a poisoned message gets dead-lettered instead of looping forever. None of that is a suggestion in a prompt. It fires whether or not anyone is watching it fire.

01How it works

Four roles, one job: uptime.

  • Monitor — every two hours it looks squarely at the whole fleet's errors, its dead-letter backlog, and whether every daemon is genuinely still alive, and it's the one that decides whether something deserves a concern or an ultra.
  • Ops — carries out restarts, deploys, and config edits, and can hot-reload the registry and every prompt without ever needing a restart to do it.
  • Incident handler — coordinates the response to anything ultra-tier across every project at once, and stays on it, broadcasting status to everyone affected until it's genuinely closed.
  • Lead — owns the architecture of the message bus itself, the foundation the other three roles operate inside of and never get to question.
02The guards

Structural, not a prompt.

Autonomy here is only ever as safe as the guards enforcing it, so none of them are a suggestion in a prompt — every one fires on its own, in code.

Kill switch
a single flag freezes every write across the entire fleet at once, with a hard variant that refuses to auto-clear itself.
Quota breaker
per-role and system-wide spend caps that trip before a runaway process gets anywhere near burning the budget.
Stuck-claim sweep
anything left in-flight past 30 minutes gets swept and re-delivered automatically.
Poison-message breaker
a message that keeps failing is dead-lettered outright rather than left looping forever.
03Where the human still stands

Honest about its limits.

A monitoring layer that overstates what it's actually capable of is worse than having none at all, so a few things stay with a person entirely on purpose. The runner genuinely cannot restart itself, so ops tells me exactly which process needs a human hand on it, and a true ultra-tier incident is designed to surface to a person rather than get quietly swallowed by the system that caused it. The guards exist to buy time honestly — never to pretend they're a substitute for one.

Runs inside AJO. It watches the daemons every other project depends on. Code is private; this page is the record.

← All work