Agent tooling · inside AJO

Agent Dev.

The part of AJO that builds and repairs the agents themselves.

Forging a role — act, validate, revise
ready
looping
Role contract
scope · output shape · escalation
• Output schema deterministic
• Required fields deterministic
• Citation integrity deterministic
• Quote length deterministic
• Quality rubric rubric
✓ Every check cleared.
The runtime acts, checks its own output against everything the role is required to satisfy, and feeds any failure straight back into another revision — looping for as long as it takes to clear every check, or escalating honestly with a structured reason when it can't. Nothing that failed a gate is ever allowed to ship.

An agent's reliability was never a wording problem, so it stops being treated as one: scope, required output, and the checks a role has to clear are declared up front, and the runtime keeps acting and revising against those checks until every one is genuinely satisfied, or it escalates honestly instead of pretending it's done. And the loop doesn't stop at first deployment — when a role actually fails on a live turn, that real failure is the thing that gets fixed, not a hypothetical stand-in for it. The fleet gets better at itself by learning from what really broke.

01How it works

Standing up a new role.

  • The lead decides the shape — what a project is genuinely missing, what a new role is allowed to own, and just as important, everything it must refuse.
  • The coder wires it in — the role joins the registry with its own inbox and outbox already built, with no plumbing left for a human to finish by hand.
  • The prompt writer authors it — the role's instructions get written as a tight, enforceable contract, never as a personality to inhabit.
  • The evaluator tests it — the role runs against a real packet of work before it's ever trusted live, so belief is never the reason it goes into service.
  • Hot-reload, no restart — the registry and every prompt are re-read on each turn, so a role change or a model swap takes hold in about two seconds, never a redeploy.
02What makes it different

A contract, not a persona.

Every role ships as a contract with validators, some of them deterministic. Deterministic checks enforce rules that admit no interpretation — the output actually matches its schema, every required field is genuinely present, each citation truly supports the claim it's attached to, no quote is allowed to run past its limit. A rubric check lets the model grade the work against named criteria, but a rubric is never trusted alone: the framework simply refuses to let a role exist without at least one deterministic check standing behind it.

When a role can't pass inside its iteration budget, it doesn't fabricate a pass to get out of the loop — it returns a structured escalation that names the last attempt honestly and every check that failed it. A role that can't clear its own gate is built to say so, not to paper over it.

03Forge from failure

A real failure is the input.

When a role fails on a live turn, that exact failure is what gets fixed, not a hypothetical stand-in for it. The actual task, the actual tool-call trace, and the actual reason it failed all go straight to the improver, which rewrites the role's instructions against that real evidence and then runs the offline loop again to confirm the new version genuinely converges before it's trusted back online. The fleet gets better at its own roles by learning from what really broke, never from a guess at what might.

Runs inside AJO. It's the project that builds and hardens every other role, including the ones behind Tool Engineering and Discovery. Code is private; this page is the record.

← All work