Method · v1.1 · August 2026

How we score agent supervision

A published method for assessing whether the person accountable for a piece of work can see, approve, and undo what an AI agent did to it.

Version
v1.1
Axes
2
Criteria
12
Scale
0–2
Applied to
11 platforms

This is the scoring method behind the Agent Supervision Register — an assessment of twelve enterprise software platforms against two axes: what their AI agents can do, and what a human can do about it afterwards.

It is published separately from the results so that the method can be checked, argued with, reused, and applied to platforms we did not assess. If you disagree with a score in the register, this page is where the disagreement should start.

The short version

Most published comparisons of AI agents in enterprise software measure one thing: capability. How much can the agent do, how autonomously, across how many systems.

We score a second axis alongside it: supervision — whether the human who remains accountable for the work can tell what the agent did, approve it before it happens, and reverse it afterwards.

Both axes use six criteria, scored 0–2, for a maximum of 12. Both are assessed against public vendor documentation only. The two scores are reported separately and never combined into a single number, because the gap between them is the finding.

Scope

What is assessed: public product documentation, admin and governance documentation, developer documentation, release notes, and official product descriptions, published by the vendor, fetched and dated during the assessment window.

What is not assessed: hands-on product testing, private beta access, customer interviews, analyst reports, third-party reviews, or vendor briefings. We did not use the products, and we did not contact the vendors. This is a documentation audit, and the scores describe what a buyer can verify before signing anything — which is deliberately the same information a buyer has.

What “not documented” means: exactly that. A zero records that we could not find the capability described in public vendor documentation after searching from at least two distinct angles. It does not mean the capability is absent from the product. Vendors with thin public documentation will score lower than vendors with the same product and better documentation. We think that is a defensible thing to measure — buyers, auditors and answer engines all read the documentation, not the roadmap — but it is a measurement of documentation, and we say so.

Products, not vendors. Scores attach to a named product at a named date, not to a company. Where a vendor has several agent products, the register names which one was scored. Where a capability exists in a vendor’s service product but not its work-management product, it is not scored — it is noted.

Why product identity is part of the method

In a category renaming itself this fast, getting the product name wrong is not a cosmetic error — it attaches twelve scored cells to the wrong thing. Before publication, every product name in the register was re-verified against the vendor’s own documentation. That pass changed five entries and withdrew one from scoring entirely. The companion AI layers name index is the output of the same check.

Axis one: capability

What the agent can do without a person present.

Criteria 1–6

Capability

Criterion012
Context reachOwn product onlyOwn product plus a fixed set of integrationsOwn product plus arbitrary external systems — connectors, MCP, web
Write authorityDrafts and suggests onlyWrites within narrow bounds, or produces drafts a human must publishCreates, updates and deletes records and artefacts directly
Unattended executionRuns only when invokedScheduled runs onlyEvent triggers, schedules, and continuous monitoring with no human present
Multi-step reasoningFixed rulesMultiple steps along a predefined pathPlans, evaluates, adapts, and recovers from its own failures
Delegation and compositionSingle agentMultiple agents, no delegation between themDocumented subagents, parallel execution, or third-party agents in one flow
ExtensibilityFixed prebuilt agentsConfigurable prebuilt agentsNon-developers can build custom agents, tools and connections

Axis two: supervision

What the accountable human can do about it.

Criteria 7–12

Supervision

CriterionThe question it answers012
Agent identity and disclosureCan a human tell which agent acted, on whose authority, and that it was not a person?Not documentedPartial, or visible only to administratorsDocumented and visible to the person affected
Action receiptIs there a per-action record a non-engineer can read afterwards, and can the person affected read it?Not documentedAdmin-only, lifecycle-only, permissioned to the agent’s owner rather than the affected party, or behind a separate licensePer-run detail available to the accountable user
ReversalCan an agent’s action be undone, and how far back does that reach?Not documentedConfiguration rollback only, or named as a recommendationDocumented reversal of the work the agent performed
Consequence-scaled approvalCan approval requirements be set by how costly the action is, rather than on or off per agent?Not documentedBlanket or per-agent approval onlyApproval gates keyed to the consequence of the action
Escalation handoffIs there a documented route to a human when the agent is out of its depth, carrying context across?Not documentedFallback behavior without human handoffDocumented handoff to a named human with context preserved
Provenance at the decision pointAre sources and uncertainty shown where the human decides, rather than buried in a log?Not documentedAvailable after the factSurfaced at the moment of review
These six are derived from auxfirst’s Action Heat Ladder work, restated here as criteria that can be scored from documentation by someone who has never read it.

The scoring rule that matters most

Rule one A recommendation is not a control.

Several major vendors publish excellent guidance on agent supervision — approval gates before destructive actions, undo flows, clear agent identity, inspectable state — addressed to the developers building agents on their platform. That guidance is real, it is well-written, and it is optional.

Where a vendor documents a supervision behavior as something the platform does, it scores as documented. Where a vendor documents it as something a builder should do, it scores 1 at most, and the distinction is stated in the row. A platform that publishes the specification and leaves implementation to third parties has done something genuinely valuable and has not shipped a control.

This rule produces the register’s most counterintuitive result, and it is the single most important thing to understand before disputing a score. In the August 2026 edition it cost one platform three points and no other platform more than one — an asymmetry that is a measurement of how advisory that platform’s supervision story is, not a penalty applied to it.

The second rule

Rule two A receipt has an audience.

The action receipt criterion asks two questions, and both must be answered for a 2:

  1. Does a per-action record exist that a non-engineer can read?
  2. Can the person whose work was changed read it?

A run log permissioned to the agent’s owner, an administrator, a security team, or a separate governance product satisfies the first and not the second. Those cases score 1.

This distinction is deliberate and is the reason the method exists. Supervision built for the person who deployed the agent is a different product from supervision built for the person accountable for the work, and the two are routinely described in the same language. Where a cell scores 1 on this criterion, the register states which of the two questions failed.

Evidence standard

Every score traces to a specific passage on a specific page, fetched on a specific date. Search-result snippets do not count as evidence; pages were retrieved in full. Where only a snippet was available, the cell is marked partially verified and the reason given.

Each cell carries the claim, the vendor page it came from, the date we fetched it and the page’s own last-updated date where published, and a confidence marker: verified, partially verified, or unverified.

Forward-looking vendor statements — “expected to be generally available in [month]” — are never scored as shipped. They are recorded as pending, with the date of the statement, and re-checked at the next revision.

Marketing language is not evidence of a mechanism. A phrase repeated verbatim across a vendor’s release notes, product pages and event listings is positioning; we score the documented surface, not the slogan, and we say when we could not find one.

An entry that cannot meet this standard is withdrawn rather than published with a caveat. In the August 2026 edition one platform was withdrawn on exactly this basis — its cells could not be tied to fetched pages for the product they were attached to, so it carries no score and is excluded from the medians.

Known limitations

We would rather publish these than have them found.

  1. The rubric penalizes deliberate restraint. A platform that never executes without approval has nothing to reverse, and takes a zero on reversal for a gap it does not structurally have. Read the two axes together, not the supervision total alone. This was found by applying the method to a platform designed around staged review, and it is a real flaw rather than a hypothetical one.
  2. Documentation quality is a confound. A well-supervised product with poor public documentation will score below a worse product with better documentation.
  3. Both axes are assessed at a moment. This is a category where product names, brands and capabilities change monthly. Four product names and one umbrella brand changed in the months around the assessment window, and two vendors shipped new tiers above existing products. Every row is dated for this reason.
  4. Six criteria are not the whole of supervision. Cost controls, rate limiting, data-loss prevention, testing and evaluation surfaces all matter and are not scored. They were excluded to keep the axis focused on the accountable individual rather than the administrator, which is the distinction this method exists to draw.
  5. Weighting is uniform and that is a choice. All six criteria count equally. A reasonable person could argue reversal matters more than provenance. We publish the per-cell scores so anyone who thinks so can reweight them.
  6. No hands-on testing. A documented control may work poorly. An undocumented one may work well. This method cannot distinguish either case.

Corrections

We expect to be wrong in individual cells and we would like to know where.

To contest a score, cite the public documentation page that supports a different reading. We will re-check against that page and either amend the cell or explain why it does not change the score. Both outcomes are logged with a date in the revision history, and amendments are shown alongside what they replaced rather than silently overwriting them.

We do not remove scores on request, and we do not accept private or unpublished evidence — if it is not publicly documented, it does not change a score, because the score measures what is publicly documented.

Send a correction through the contact form.

Reusing this method

The criteria on this page are published so they can be applied to platforms we did not assess. If you use them, we ask two things: score against public documentation rather than impressions, and publish your evidence per cell. Attribution to auxfirst is welcome and not required.

Revision history

VersionDateChange
v1.121 Aug 2026First publication. Adds product-identity verification as a prerequisite step, states the withdrawal rule for entries that cannot meet the evidence standard, and records limitation 1 as observed rather than anticipated.
v1.020 Aug 2026Internal draft. Two axes, twelve criteria, both scoring rules. Not published.

Emil Krzemiński is the founder of auxfirst, the agentic experience design agency — helping product, developer and business teams design AI systems that remember, adapt, and earn the right to act. Start with the register, the agent development lifecycle, or a conversation. For how machines read auxfirst, see the AI info page.