AI support agent · evals · Live demo

Fieldstone.

A support agent. Escalating isn't the system failing; it's the system working exactly as designed.

Try the live demo ↗
One agent turn — route, ground, or escalate
routing
cust
Where is my order FG-100001?
Typed tool selected
lookup_order check_return_eligibility initiate_return search_help_center escalate_to_human
model self-rating95%
retrieval grounding92%
answered · grounded grounded in the store’s order data
Every action runs through a typed tool — never an improvised prompt. When a request sits outside its scope, escalation is the answer, never a guess.

Every action Fieldstone takes runs through a typed tool, never a prompt it's left to improvise around. Retrieval is grounded in the store's own data, not the model's memory. And confidence is never a single number — it's two independent signals. Escalation isn't a fallback for the cases nothing else catches; it's a first-class tool, because knowing the edge of what it's sure of is the entire discipline — checked against a golden-set evaluation, not just claimed.

01The design

Answer what it's sure of. Escalate the rest.

Nothing here runs on inference the model was left to improvise. Every action goes through a typed tool built for exactly what it does:

lookup_order · check_return_eligibility · initiate_return
order actions as typed tools, so answers come from the store’s data, never inference.
search_help_center
retrieval from a 50-article help center built from the store’s own data, for anything the order tools can’t answer.
escalate_to_human
a first-class tool with an eleven-value typed reason_code enum (payment_dispute, policy_exception, account_access, out_of_scope…) a contact-center platform would actually route on.

An autonomous support agent's single most valuable move is knowing the edge of its own knowledge, completely and without exception. That's the whole reason escalation is built as a tool with full standing, never left as a footnote for whatever the rest of the system fails to catch.

02How I know it works

An Opus-judged golden set.

Knowing it works means measuring it the same way every time: an 18-case RAG golden set, judged by Opus on four rubric dimensions — tool routing, grounding, escalation, response quality. Stability is verdict agreement across runs, not a single lucky pass.

Pass
cases · runs
pass rate
stable
Before three fixes
18 · 36
78%
16/18
After fixes
18 · 54
98.1%
17/18

One residual miss is documented in the postmortem and deliberately left unpatched. Tuning the agent's language to clear a judge rubric is the first step of Goodhart drift, not a fix — and saying plainly what a metric does and doesn't prove, every time, is the whole discipline, not a caveat added at the end.

03Dual-signal confidence

The disagreement is the signal.

Every answer gets two fully independent reads, never one number standing in for the whole judgment: the model's self-rating of its own reply, and a retrieval-derived score from the cosine similarity of the chunks it used. When they disagree — a confident retrieval under a model that rates its own reply thin — that gap tells you more than either number alone ever could. It's instrumented. It doesn't gate replies yet, and won't until calibration is measured first: a number only earns the right to block a reply once you know what it actually proves.

04Architecture

SDK direct, no framework.

Runtime
Next.js App Router, Anthropic SDK direct — full visibility into every step.
Retrieval
50-article corpus generated offline against store.json, embedded once on voyage-4-lite. Runtime cosine ~2 ms.
State
In-memory everything — session state and corpus from JSON. No database, no vector store. On Vercel.
Model tiers
Sonnet for the agent turn, Haiku for the confidence side-call, Opus for the eval judge.

No LangChain, no LlamaIndex, no agent framework. Full visibility into what the model sees and does at every step isn't a byproduct of that choice — it's the entire point of the build.

← All work