Fieldstone.
A support agent. Escalating isn't the system failing; it's the system working exactly as designed.
Try the live demo ↗Every action Fieldstone takes runs through a typed tool, never a prompt it's left to improvise around. Retrieval is grounded in the store's own data, not the model's memory. And confidence is never a single number — it's two independent signals. Escalation isn't a fallback for the cases nothing else catches; it's a first-class tool, because knowing the edge of what it's sure of is the entire discipline — checked against a golden-set evaluation, not just claimed.
Answer what it's sure of. Escalate the rest.
Nothing here runs on inference the model was left to improvise. Every action goes through a typed tool built for exactly what it does:
An autonomous support agent's single most valuable move is knowing the edge of its own knowledge, completely and without exception. That's the whole reason escalation is built as a tool with full standing, never left as a footnote for whatever the rest of the system fails to catch.
An Opus-judged golden set.
Knowing it works means measuring it the same way every time: an 18-case RAG golden set, judged by Opus on four rubric dimensions — tool routing, grounding, escalation, response quality. Stability is verdict agreement across runs, not a single lucky pass.
One residual miss is documented in the postmortem and deliberately left unpatched. Tuning the agent's language to clear a judge rubric is the first step of Goodhart drift, not a fix — and saying plainly what a metric does and doesn't prove, every time, is the whole discipline, not a caveat added at the end.
The disagreement is the signal.
Every answer gets two fully independent reads, never one number standing in for the whole judgment: the model's self-rating of its own reply, and a retrieval-derived score from the cosine similarity of the chunks it used. When they disagree — a confident retrieval under a model that rates its own reply thin — that gap tells you more than either number alone ever could. It's instrumented. It doesn't gate replies yet, and won't until calibration is measured first: a number only earns the right to block a reply once you know what it actually proves.
SDK direct, no framework.
No LangChain, no LlamaIndex, no agent framework. Full visibility into what the model sees and does at every step isn't a byproduct of that choice — it's the entire point of the build.