Three systems, each built on the same pattern: the agent does the routine work, the rules are code, a person keeps the exceptions. All three are live below. The records are sample data; the rules, the code, and the approval step are the real thing.
Refund Approvals · Online Store
Churn & LTV · Subscriptions
Invoice Follow-Up · Services Firm
Transaction Review · Payments
Reorder Ranking · Online Grocery
Concept preview. The screen below is a working mock: open a request, read the evidence, approve it, and watch the decision get logged. The rules behind it are the ones running in the Approval Desk build.
Each dial is a real configuration choice: model tier, retry budget, retrieval depth. Each maps to a number your team will live with. Move one dial and the other two move too.
Seconds per decision, p50 and p95. Set by model tier and retry budget.
Verdict and citation validity against an agreed golden set.
Dollars per case, itemized by step. You set the ceiling; this dial reports what the system spent against it.
These are failure modes, not certainties. Each one shows up when one dial gets set without measuring the other two.
The cheapest model, no review passes. A decision can go out citing a policy section that does not apply. Fast, cheap, and wrong. Where the work is regulated, wrong is measured in customer harm.
Top-tier models and unlimited review loops. A routine approval can take ninety seconds and forty times the tokens. At volume the queue backs up and staff start overriding the system.
The verification step is the first expense cut, and it is the check that catches invented criteria. The savings can hold until the first audit finding.
The engagement work is choosing the setting deliberately, with the trade-offs on paper, before production.
A bench that runs the same banking questions through several hosted models and reports accuracy, cost, and speed side by side, on real data. Full Banking77 test set, 3,080 messages, three hosted Bedrock models.
Open Model Bench