Agents that do the routine work and ask you before anything leaves. Giggit AI designs, builds, and tests them for small and mid-size businesses where mistakes are costly, tuned on three dials: speed, quality, and cost. They log every step, score every decision, and stay within the limits you set.
Operations Assistant · delivered for a client
Delivered for a client's receivables list. The assistant reads the invoices and the customer history, drafts the next action, and holds it for a person.
Every answer carries a receipt of what it read. Nothing is sent without a person's approval.
AREA · delivered for a client
A customized AREA was delivered for a fintech platform. It answers questions from a database with a citation on every number and a check on every answer.
The public version runs on CMS Open Payments. Eight reference questions are re-asked every six hours.
Model Selection · Model Bench
Local classifier 89.4% at $0 (no provider fee; hosting excluded); best hosted model 72.2% at $0.40 per 1,000 messages, on 3,080 banking messages.
Benchmark · public data · live comparison
Transaction Review · Fraud Radar
Reviewing the riskiest 1% catches 26.2% of the fraud at 90.1% precision, on 118,108 held-out transactions.
Reference system · public data · live scoring endpoint
Reorder Ranking · Reorder Radar
NDCG@10 0.549 against 0.518 for the stronger baseline, on 12,231 held-out shoppers. Offline ranking quality, not revenue lift.
Reference system · public data · live ranking endpoint
Customer Value · Buyer Value Radar
Six- and twelve-month spend forecasts on real wholesale data. WAPE 68.8% at twelve months on 1,295 holdout customers, GBP.
Reference system · public data · live scoring endpoint
How the systems are performing
Five systems, checked every six hours on fixed reference inputs. Live Checks
Engagements run in short loops with a checkpoint at every step. Scope is fixed, and the success measure is agreed in writing before work starts.
1 · Scope
Deciding what an agent does on its own, and what it hands to a person, sets everything downstream. We scope those boundaries, design the tools and orchestration, and build the simplest loop that holds up on real inputs.
What you get: a working agent workflow plus a record of every autonomy decision
2 · Measure
We build the measurement layer first: calibrated scoring, trajectory-level evaluation that grades how the agent reached an answer, and regression suites on every change. Then the brakes: budget ceilings, step caps, circuit breakers, escalation paths.
What you get: an evaluation suite in your repo, baseline scores, and runtime controls with drill results. Our own systems, checked every six hours →
3 · Cite Sources
Retrieval pipelines measured end to end, from ingestion to reranking. Every answer traces to a source. Every decision cites its source: policy sections, registry records, documents. Unsourced answers do not ship.
What you get: retrieval quality you can measure, and answers that cite where they came from. See it live: Research Agent →
4 · Size to the Team
A five-person company needs three workflows that stop eating the founder's week, and a platform is rarely the answer. The expensive model runs only on steps that need it, routine steps run on cheaper ones, and the measurements decide which is which.
What you get: a decision you can defend with the trade-offs on paper, and a runbook a non-engineer can follow. See the model-selection numbers →
How It Is Wired
A short call is usually enough to tell whether we are a fit. You leave with a clear read on what agents can and cannot do for you.
Start a Conversation