03 · Agent Experience Audit  ·  5-Day Diagnostic

Your agent is already in market. Now find out what trust it's actually building.

A fixed-scope, fixed-timeline diagnostic of one shipped AI product against the AUX framework. You walk away with a Trust Scorecard, a named gap list, and a 90-day prioritized fix roadmap — in five business days.

Timeline
5 business days
Scope
One shipped AI product
Starting price
$14,500 flat
Output
Scorecard + roadmap

A 5-business-day diagnostic of one shipped AI product against the AUX framework. Fixed scope, fixed timeline, fixed price.

Most AI products ship without ever being evaluated against how they build — or break — user trust. The Agent Experience Audit closes that gap. We examine your product across the full AUX trust architecture, score it against the 10 AUX heuristics, identify named failure modes with evidence, and hand you a prioritized roadmap your team can act on immediately.

One product. One audit. A "product" means one user-facing AI surface — a copilot, an assistant, an agentic workflow. Multi-product or platform-level audits are a separate engagement. This constraint is what makes delivery in five days possible.

Five dimensions. Everything examined, nothing assumed.

Every audit covers the same five dimensions in the same sequence. This is what makes the methodology consistent and the scorecard comparable across products.

Dimension
What we examine
Output
T01–T04 Trust Architecture
How the product behaves across Functional → Contextual → Judgment → Advocacy stages
Stage-by-stage maturity score
H01–H10 10 AUX Heuristics
Each heuristic scored 0–3 against observed behavior across ~20 scripted scenarios
Heatmap with pass/fail flags
tg.* Trust Gap Taxonomy
Named failure modes detected in your product — classified, evidenced, and staged
Gap list with severity + evidence links
MEM Memory & Boundaries
Memory scopes, retention policy, escalation paths, refusal triggers, consent signals
Trust Contract gap analysis
BENCH Vendor / Peer Benchmark
How your product compares to 2 named comparables we've previously evaluated
Comparative scorecard
Not in scope — so you can say no without negotiation
  • Custom code review or model evaluation
  • Fine-tuning, dataset audits, RLHF assessment
  • Implementation of fixes (that's Blueprint Sprint)
  • Multi-product or multi-team rollout
  • Live ride-alongs with end users
  • Compliance certification (SOC2, HIPAA, etc.)
  • Anything taking us past the 5-day window
  • Any scope added after kickoff packet is signed off
Required Inputs · Kickoff Packet

Five things you send us 3 business days before Day 1.

If the packet isn't complete by Day −3, kickoff slips a week. No exceptions — this is what stops the engagement running into 12-day chaos.

  • Product access — sandbox account, demo URL, or recorded walkthrough
  • Agent spec — populated agent-spec.schema.yaml template (we provide the blank)
  • 5–10 representative transcripts or session recordings
  • Memory/persistence policy doc if one exists — or we note its absence
  • One named contact with authority to answer follow-up questions within 24h

⟶ We send the blank agent-spec template and a packet checklist on booking confirmation.

Same sequence, every time. No scope creep, no surprises.

The structure is fixed by design. Every audit follows the same five-day cycle — which is what makes it deliverable in five days and comparable across clients.

Day
What happens
Artifact produced
Pre-Day · T−3
Kickoff Packet
Kickoff packet received, reviewed, gap-questions sent back. Any missing inputs trigger a one-week slip — we send a clear notice within 24h of review.
Internal Reviewed inputs + gap questions
Day 1
Discovery
60-min kickoff call. Walk through the agent spec, the product team's intended behavior, edge case history, and known gaps. Founder hears the product team's own theory of how it works — before we probe for where that theory breaks.
Internal Discovery memo
Day 2
Behavioral Analysis
We run the product against ~20 scripted scenarios, probing each H01–H10 systematically. All interactions are transcribed. Initial heatmap scored. Anomalies flagged for deeper probing.
Working Annotated transcripts + draft heatmap
Day 3
Gap Classification
Apply the trust-gap taxonomy to observed anomalies. Each detected gap receives a tg.family.name classification, a severity rating, a stage-collapse description, and an evidence link.
Working Named gap list
Day 4
Roadmap Synthesis
Translate scores and gaps into a three-tier 90-day roadmap: must-fix before next release, fix this quarter, and architectural debt. Prioritization is by user trust impact, not engineering effort.
Draft Prioritized 90-day roadmap
Day 5
Readout
90-min live readout call. Walk through the scorecard, named gaps, roadmap, and the next step recommendation. Session recorded. Q&A with your product and engineering leads. Written deliverable handed over same day.
Final Full report — PDF + Markdown

Three deliverables your team can act on immediately.

Every audit produces the same three primary deliverables. No slide decks, no generic recommendations. Each document is specific to your product — built from your transcripts, your agent spec, and your users' observed experience.

D1

Trust Scorecard

Your product's maturity across the four trust architecture stages (T01–T04) and all ten AUX heuristics (H01–H10), scored 0–3 with a visual heatmap and an overall grade.

  • Stage-by-stage maturity assessment
  • H01–H10 scores with pass/fail flags
  • Peer benchmark comparison
  • Overall trust grade
D2

Named Gap List

Every trust failure mode detected, classified with its tg.* taxonomy code, severity rating, and linked to the specific transcript evidence that surfaced it.

  • tg.* taxonomy classification
  • Severity: high / medium / low
  • Stage collapse mapping
  • Evidence transcript links
D3

90-Day Fix Roadmap

A prioritized three-tier action plan: must-fix before next release, fix this quarter, and architectural debt. Each item links to a named gap and a recommended fix pattern from the AUX library.

  • Must-fix / this quarter / architectural
  • Fix patterns from AUX library
  • Blueprint Sprint scope stub
  • Acceptance criteria for each fix

The Trust Scorecard looks like this. Names and scores are illustrative — your product's results will vary.

Agent Experience Audit · Anonymized Co. · 5-day engagement · auxfirst
Trust Architecture
T01 Functional  
T02 Contextual  
T03 Judgment   
T04 Advocacy   
Stage reached: T02  ·  Fix 2 heuristics before next release
AUX Heuristics (0 – 3)
H01 Intent visibility       2.4
H02 Progressive transp. 1.8
H03 User steering        2.6
H04 Dynamic trust       0.9
H05 Boundaries          3.0
H06 Uncertainty         0.7  ⚠ FAIL
H07 Assertiveness       0.8  ⚠ FAIL
H08 Context efficiency   2.2
H09 Multi-agent         1.5
H10 Consistency         2.4
Overall grade: C+

Two tiers. Same 5-day cycle. Same rigor.

Both tiers cover the full audit methodology. The difference is scope of stakeholders, product complexity, and readout format.

Intensive
$24,500
Flat fee · no hourly overruns

Complex product or stakeholders

Same five-day cycle, extended for products with multiple locales, regulated environments, or stakeholders who need a board-level readout.

  • Multi-locale or regulated product
  • Product with >100k MAU
  • Multi-stakeholder readout (board, legal, eng)
  • Additional stakeholder interviews (up to 4)
  • Executive summary formatted for board
  • Blueprint Sprint scope stub included
Payment: 50% on signature · 50% on readout

When this engagement is the right move.

Strong fit
  • You have a shipped AI product and genuine uncertainty about whether it's building trust or eroding it
  • A feature release, fundraise, or enterprise contract is coming and you need third-party validation
  • Your team built it fast and never stepped back to evaluate the full experience against a framework
  • You want a prioritized fix list before committing engineering cycles to the wrong improvements
  • You're considering a Blueprint Sprint and want to scope it against real, diagnosed gaps — not assumptions
Not the right fit
  • You're still deciding whether to build — start with the Executive Seminar
  • You're designing a new agent from scratch — start with the Blueprint Sprint
  • You want us to fix the product, not evaluate it — fixes are a separate engagement
  • You can't provide product access or representative transcripts within the kickoff window

The mirror image: can machines parse your brand?

This audit asks whether humans can trust your agent. Its mirror is whether machines can understand your brand — because in 2026 an AI answer engine is often the first thing that decides whether you're in the consideration set at all.

The Entity Legibility add-on scores how cleanly AI systems can parse, attribute, and cite you: semantic-triple coverage (subject → predicate → object), Schema.org and DefinedTerm markup, an answer-first AI Info Page, machine-readable discovery files (llms.txt, capabilities.json), and off-page entity consistency. The same discipline as the audit — making a system legible — pointed at the brand instead of the agent.

It folds into this engagement or the AI-readiness audit as a discrete module. Ask about the Entity Legibility add-on →

Bring us a shipped product.
Walk away knowing exactly what to fix.

A 30-minute scoping call to confirm the product is in scope, identify your contact, and agree a kickoff date. If the audit isn't the right engagement for where you are, we'll tell you that on the call.

Book a scoping call

Tell us which product you want audited and what's driving the timing. We'll confirm fit and come back with a kickoff date within one business day.

View All Engagements