Most AI teams can make an agent work. The harder problem is making people willing to delegate real work to it. auxfirst is an agentic experience agency: we design the relationship layer — what the agent is allowed to do, when it must ask, what it remembers, how it explains itself and how a person takes back control.
Not sure whether your problem is capability or trust? That is exactly what the first session is for.
Book an Agent Experience Working SessionNew to the category? Start with What is Agentic User Experience?
Agents that technically work still fail in the hands of real users — and they fail in patterns we see repeatedly:
Nobody can say what the agent is permitted to do without asking. So people either over-trust it or refuse to use it.
The agent remembers things users cannot see, edit or delete. Every surprise costs trust that took weeks to build.
The agent presents a confident answer and a wrong answer identically. Users cannot calibrate.
When the agent hits an edge case, there is no clean path back to a human — so the whole flow is abandoned.
A single visible error ends adoption, because there is no way to correct, undo or teach the agent.
The pilot impressed the steering committee. Six weeks later, usage is flat.
These are design problems, not model problems. More capable models do not fix them.
An agentic experience agency designs the behavior and trust layer of AI products — the decisions that sit between a working model and a person willing to rely on it. In practice, that means six things:
What the agent does, in what order, with what tone and what refusal conditions.
Which actions are autonomous, which need approval, which are off-limits — and how that changes as trust grows.
How confidence, uncertainty, evidence and provenance are shown at the moment of decision.
What is remembered, for how long, visible to whom, and editable by whom.
How the agent hands work to a person, and how the person hands it back.
How you test whether people can actually understand, direct and recover control from the agent before you scale it.
This is distinct from building the agent. We produce the specification your engineering team — or your build partner — implements against. For the full discipline behind this work, see our guide to Agentic User Experience, or the process framework in Agent-First Design.
The right moment
The right conditions
We would rather tell you now:
Not our work
Wrong fit
"Agency" covers very different jobs in AI right now. Here is where each type is genuinely the right call — and where we fit.
| Provider type | Primary question | Typical output | Gap | When it is right |
|---|---|---|---|---|
| UX/UI agency | How should the interface work? | Flows, screens, design system | May not define agent authority or behavior | Human-operated products and interface redesigns |
| AI automation agency | How can the workflow be automated? | Integrations, automations, deployed agents | May optimize completion without designing trust | Clear, bounded process automation |
| AI product studio | How do we build the product? | Product strategy, prototype, software | Broad remit; the agent relationship may be one workstream | End-to-end product delivery |
| MCP/infrastructure consultancy | How should tools and runtimes connect? | Architecture, servers, tool schemas | Does not own human adoption | Agent infrastructure and interoperability |
| auxfirst | What should the agent own, and how will people trust it? | Behavior, autonomy, memory, escalation, trust and validation spec | Requires an internal product/business owner | Teams building or fixing consequential agents |
↔ scroll table on mobile
A leadership session that establishes shared language on agentic AI, autonomy and trust, so the organization stops arguing past itself.
A hands-on working session with the product, design and engineering people who will own the agent, ending in agreed autonomy and escalation decisions.
A structured evaluation of an existing agent against the AUX heuristics, producing a trust scorecard and a prioritized fix list.
Agent behavior, autonomy map, memory strategy, escalation model and interaction blueprint — specified for build.
Testing an agent experience with real users before scale: can they understand it, direct it, and recover control when it is wrong?
Ongoing senior input for teams shipping agentic products continuously, rather than in one project.
Adjacent fixed-price diagnostics: the Agent Operability Audit (is this workflow ready for an agent at all?) and the Agentic Shelf Audit (what AI models already believe about your brand).
Not sure which one fits? Book a working session and we will tell you honestly — including if the answer is none of them.
Book an Agent Experience Working SessionWhere the current or planned experience stands against the AUX heuristics, with the gaps ranked by consequence.
The specified experience: what the agent does, says, shows and refuses, across the main paths and the failure paths.
Every action the agent can take, classified by autonomy level, with the conditions for moving between levels.
What is stored, for how long, who can see it, who can edit it, and what the user is told.
Approval gates, handoff paths, audit requirements and accountability.
How to validate the experience with real users before scaling, and what would constitute failure.
These are specification artifacts. Your team — or your build partner — implements against them.
Discover → Map → Design → Validate → Hand over
What we need from you: a product or business owner who can make decisions, access to the people who will use or supervise the agent, and whoever owns risk, compliance or legal if the domain requires it.
By product type
Shopping, service, onboarding and support agents where a mistake is visible to the customer.
Agents acting inside operational processes, where the supervisor is a colleague rather than a customer.
Tools, APIs and environments that other agents operate, where the user is partly software.
By industry
Shopping, merchandising, pricing and service agents. See our gallery of AI agents in retail.
High-consequence flows where auditability and evidence are not optional.
Agentic workflows across planning, production and measurement.
Products adding agentic features to an existing user relationship.
We work in the open. The frameworks below are published and maintained by auxfirst, and every engagement runs through them.
Framework
Four stages an agent earns: Functional → Contextual → Judgment → Advocacy. You design for the stage the task actually needs.
Framework
Rate every action by consequence, then match the friction — silent, confirmed, or human-gated — to the heat. See the ladder →
Framework
Capture, Reference, Credit, Stamp — the four moves that make an agent's actions traceable and accountable.
Framework
A legible contract for each agent: scope, capability, and hard limits, written for the people who rely on it.
The evaluation lens is the ten AUX heuristics; the interaction vocabulary is the pattern library. Both ship openly through TrustKit (the toolkit) and Agent Process Design (the method).
Regulatory hooks we design against: EU AI Act · Singapore IMDA Model AI Governance Framework · NIST AI risk guidance.
Published analysis
Field note · Repository review
A bank's open-source AI lab inventoried in full — then read by its defaults, which is where a team's real trust model shows.
Field guide · Free
Why enterprise agent pilots stall between demo and adoption, and the five-layer whole product that carries them across.
Open source · MIT + CC BY 4.0
The canon as editable YAML: heuristics, trust architecture, the trust-gap taxonomy, agent specs and memory policies.
Figma Community · Free
The one-page artifact we use in workshops to map an agent's trust topology, autonomy boundaries and handoffs.
Agent Experience (AX) is about whether a software agent can operate your system. Agentic User Experience (AUX) is about whether a person can safely delegate to that agent. AX asks whether the machine can use your product; AUX asks whether a human will let it act on their behalf.
We design it and specify it. Your engineering team or build partner implements. We work alongside them and stay available through implementation when that helps.
It depends on the format. An Executive Seminar runs half a day to a full day and a Team Workshop one to three days. The diagnostic and design engagements — the Agent Experience Audit and the Blueprint Sprint — are scoped in days to weeks, not months. An Advisory Retainer runs quarterly or annually.
At minimum a product or business owner with decision rights, plus the people who will use or supervise the agent. In regulated domains, add risk, compliance or legal early rather than late.
Yes. We work remotely with teams across Europe, and run sessions on-site where that is the better format.
Because agentic products introduce design decisions most UX practices have never had to make: autonomy levels, memory visibility, escalation design and trust calibration. Often the best outcome is that your team leaves with the framework and does this themselves next time.
Book an Agent Experience Working Session. If we are not the right fit, we will say so on that call.
Ready when you are. Tell us the agent, the workflow and who owns it — we will tell you whether this is a trust problem worth designing for.
Book an Agent Experience Working SessionThat's the gap we design for. Start with an Agent Experience Audit and we'll show you exactly where the trust breaks — and what to fix first.
Book an audit