ai & agent infrastructure, designed and built

Governed Agents for Decisions That Matter

Agents that do the routine work and ask you before anything leaves. Giggit AI designs, builds, and tests them for small and mid-size businesses where mistakes are costly, tuned on three dials: speed, quality, and cost. They log every step, score every decision, and stay within the limits you set.

operations assistantsample records
Invoice Follow-Up With Evidence and Approval
"What is the story with the disputed invoice?"
INV-2044 · Cobalt Retail Co · $12,500Issued Jul 15 26 · due Aug 14 26 · disputed, 23 days lateAP note: "Invoice has no PO number. We cannot process without PO-88213."
Customer record · Cobalt Retail CoContact Marcus Lee · Net 30 · pays on time when a PO number is on the invoice. AP rejects anything without one.
What the assistant foundINV-2044 for $12,500 is not a late payer problem. It is our mistake. Cobalt's AP rejected it because the PO number was missing, and their history says they pay on time when the PO is on the invoice. A reminder would make it worse. The right move is a corrected invoice with PO-88213 and a short apology.
Draft · to Marcus Lee · Corrected invoice INV-2044 with PO-88213Hi Marcus Lee, thank you for flagging the missing PO number on INV-2044. You were right: it was issued without it. Attached is the corrected invoice for $12,500 with PO-88213 on it. Could you confirm it is now in your queue? Apologies for the extra step on your side.Held. Approval stays off until the corrected invoice is attached.
Run this example
Wired with Claude Codeagent in GitHub Actions Anthropic APIClaude models AWSLambda · S3 · Bedrock GitHub Actionstests · deploys · checks Vercelsite and functions Postgreson Neon
how the work is delivered

How the Work Is Delivered

Engagements run in short loops with a checkpoint at every step. Scope is fixed, and the success measure is agreed in writing before work starts.

1 · Scope

What Should the Agent Own?

Deciding what an agent does on its own, and what it hands to a person, sets everything downstream. We scope those boundaries, design the tools and orchestration, and build the simplest loop that holds up on real inputs.

What you get: a working agent workflow plus a record of every autonomy decision

2 · Measure

How Do You Know It Works?

We build the measurement layer first: calibrated scoring, trajectory-level evaluation that grades how the agent reached an answer, and regression suites on every change. Then the brakes: budget ceilings, step caps, circuit breakers, escalation paths.

What you get: an evaluation suite in your repo, baseline scores, and runtime controls with drill results. Our own systems, checked every six hours →

3 · Cite Sources

Can It Show Its Sources?

Retrieval pipelines measured end to end, from ingestion to reranking. Every answer traces to a source. Every decision cites its source: policy sections, registry records, documents. Unsourced answers do not ship.

What you get: retrieval quality you can measure, and answers that cite where they came from. See it live: Research Agent →

4 · Size to the Team

What Can a Small Team Actually Run?

A five-person company needs three workflows that stop eating the founder's week, and a platform is rarely the answer. The expensive model runs only on steps that need it, routine steps run on cheaper ones, and the measurements decide which is which.

What you get: a decision you can defend with the trade-offs on paper, and a runbook a non-engineer can follow. See the model-selection numbers →

How It Is Wired

An LLM in the Loop, With the Rails Around It

1 · GitHubEvery system is a repository with tests in CI. A Claude Code agent runs in GitHub Actions on labeled issues, writes the change, and pushes a branch. A person reviews and merges.
2 · DeployA merge to main deploys: Vercel for this site and its functions, AWS Lambda behind API Gateway for the scoring models, with the model files on S3.
3 · The Model CallThe Operations Assistant plans with a Claude model through the Anthropic API. AREA and Model Bench call hosted models through AWS Bedrock. Every call is logged with tokens and cost.
4 · The DataRead-only tools over the records. AREA reads Postgres on Neon through allow-listed SQL, one query per question, with a citation check before the answer is shown.
5 · The ChecksEvery six hours GitHub Actions sends each system its reference inputs and writes the score to a public ledger. Live Checks
01The agent proposes. You approve what matters.The agent acts only within an approved list of actions, every claim is checked against sources, and it hands off to a person when it is unsure.
02Every decision cites its source.Every claim traces to a source: policy sections, registry records, documents. Unsourced answers do not ship.
03Measure everything we ship.Each system ships with its own tests: known-answer sets, failure tracking, and cost per decision.
contact

Have a Workflow in Mind? Let's Look at It Together

A short call is usually enough to tell whether we are a fit. You leave with a clear read on what agents can and cannot do for you.

Start a Conversation