How we score agent supervision
A published method for assessing whether the person accountable for a piece of work can see, approve, and undo what an AI agent did to it.
This is the scoring method behind the Agent Supervision Register — an assessment of twelve enterprise software platforms against two axes: what their AI agents can do, and what a human can do about it afterwards.
It is published separately from the results so that the method can be checked, argued with, reused, and applied to platforms we did not assess. If you disagree with a score in the register, this page is where the disagreement should start.
The short version
Most published comparisons of AI agents in enterprise software measure one thing: capability. How much can the agent do, how autonomously, across how many systems.
We score a second axis alongside it: supervision — whether the human who remains accountable for the work can tell what the agent did, approve it before it happens, and reverse it afterwards.
Both axes use six criteria, scored 0–2, for a maximum of 12. Both are assessed against public vendor documentation only. The two scores are reported separately and never combined into a single number, because the gap between them is the finding.
Scope
What is assessed: public product documentation, admin and governance documentation, developer documentation, release notes, and official product descriptions, published by the vendor, fetched and dated during the assessment window.
What is not assessed: hands-on product testing, private beta access, customer interviews, analyst reports, third-party reviews, or vendor briefings. We did not use the products, and we did not contact the vendors. This is a documentation audit, and the scores describe what a buyer can verify before signing anything — which is deliberately the same information a buyer has.
What “not documented” means: exactly that. A zero records that we could not find the capability described in public vendor documentation after searching from at least two distinct angles. It does not mean the capability is absent from the product. Vendors with thin public documentation will score lower than vendors with the same product and better documentation. We think that is a defensible thing to measure — buyers, auditors and answer engines all read the documentation, not the roadmap — but it is a measurement of documentation, and we say so.
Products, not vendors. Scores attach to a named product at a named date, not to a company. Where a vendor has several agent products, the register names which one was scored. Where a capability exists in a vendor’s service product but not its work-management product, it is not scored — it is noted.
In a category renaming itself this fast, getting the product name wrong is not a cosmetic error — it attaches twelve scored cells to the wrong thing. Before publication, every product name in the register was re-verified against the vendor’s own documentation. That pass changed five entries and withdrew one from scoring entirely. The companion AI layers name index is the output of the same check.
Axis one: capability
What the agent can do without a person present.
Capability
| Criterion | 0 | 1 | 2 |
|---|---|---|---|
| Context reach | Own product only | Own product plus a fixed set of integrations | Own product plus arbitrary external systems — connectors, MCP, web |
| Write authority | Drafts and suggests only | Writes within narrow bounds, or produces drafts a human must publish | Creates, updates and deletes records and artefacts directly |
| Unattended execution | Runs only when invoked | Scheduled runs only | Event triggers, schedules, and continuous monitoring with no human present |
| Multi-step reasoning | Fixed rules | Multiple steps along a predefined path | Plans, evaluates, adapts, and recovers from its own failures |
| Delegation and composition | Single agent | Multiple agents, no delegation between them | Documented subagents, parallel execution, or third-party agents in one flow |
| Extensibility | Fixed prebuilt agents | Configurable prebuilt agents | Non-developers can build custom agents, tools and connections |
Axis two: supervision
What the accountable human can do about it.
Supervision
| Criterion | The question it answers | 0 | 1 | 2 |
|---|---|---|---|---|
| Agent identity and disclosure | Can a human tell which agent acted, on whose authority, and that it was not a person? | Not documented | Partial, or visible only to administrators | Documented and visible to the person affected |
| Action receipt | Is there a per-action record a non-engineer can read afterwards, and can the person affected read it? | Not documented | Admin-only, lifecycle-only, permissioned to the agent’s owner rather than the affected party, or behind a separate license | Per-run detail available to the accountable user |
| Reversal | Can an agent’s action be undone, and how far back does that reach? | Not documented | Configuration rollback only, or named as a recommendation | Documented reversal of the work the agent performed |
| Consequence-scaled approval | Can approval requirements be set by how costly the action is, rather than on or off per agent? | Not documented | Blanket or per-agent approval only | Approval gates keyed to the consequence of the action |
| Escalation handoff | Is there a documented route to a human when the agent is out of its depth, carrying context across? | Not documented | Fallback behavior without human handoff | Documented handoff to a named human with context preserved |
| Provenance at the decision point | Are sources and uncertainty shown where the human decides, rather than buried in a log? | Not documented | Available after the fact | Surfaced at the moment of review |
The scoring rule that matters most
Several major vendors publish excellent guidance on agent supervision — approval gates before destructive actions, undo flows, clear agent identity, inspectable state — addressed to the developers building agents on their platform. That guidance is real, it is well-written, and it is optional.
Where a vendor documents a supervision behavior as something the platform does, it scores as documented. Where a vendor documents it as something a builder should do, it scores 1 at most, and the distinction is stated in the row. A platform that publishes the specification and leaves implementation to third parties has done something genuinely valuable and has not shipped a control.
This rule produces the register’s most counterintuitive result, and it is the single most important thing to understand before disputing a score. In the August 2026 edition it cost one platform three points and no other platform more than one — an asymmetry that is a measurement of how advisory that platform’s supervision story is, not a penalty applied to it.
The second rule
The action receipt criterion asks two questions, and both must be answered for a 2:
- Does a per-action record exist that a non-engineer can read?
- Can the person whose work was changed read it?
A run log permissioned to the agent’s owner, an administrator, a security team, or a separate governance product satisfies the first and not the second. Those cases score 1.
This distinction is deliberate and is the reason the method exists. Supervision built for the person who deployed the agent is a different product from supervision built for the person accountable for the work, and the two are routinely described in the same language. Where a cell scores 1 on this criterion, the register states which of the two questions failed.
Evidence standard
Every score traces to a specific passage on a specific page, fetched on a specific date. Search-result snippets do not count as evidence; pages were retrieved in full. Where only a snippet was available, the cell is marked partially verified and the reason given.
Each cell carries the claim, the vendor page it came from, the date we fetched it and the page’s own last-updated date where published, and a confidence marker: verified, partially verified, or unverified.
Forward-looking vendor statements — “expected to be generally available in [month]” — are never scored as shipped. They are recorded as pending, with the date of the statement, and re-checked at the next revision.
Marketing language is not evidence of a mechanism. A phrase repeated verbatim across a vendor’s release notes, product pages and event listings is positioning; we score the documented surface, not the slogan, and we say when we could not find one.
An entry that cannot meet this standard is withdrawn rather than published with a caveat. In the August 2026 edition one platform was withdrawn on exactly this basis — its cells could not be tied to fetched pages for the product they were attached to, so it carries no score and is excluded from the medians.
Known limitations
We would rather publish these than have them found.
- The rubric penalizes deliberate restraint. A platform that never executes without approval has nothing to reverse, and takes a zero on reversal for a gap it does not structurally have. Read the two axes together, not the supervision total alone. This was found by applying the method to a platform designed around staged review, and it is a real flaw rather than a hypothetical one.
- Documentation quality is a confound. A well-supervised product with poor public documentation will score below a worse product with better documentation.
- Both axes are assessed at a moment. This is a category where product names, brands and capabilities change monthly. Four product names and one umbrella brand changed in the months around the assessment window, and two vendors shipped new tiers above existing products. Every row is dated for this reason.
- Six criteria are not the whole of supervision. Cost controls, rate limiting, data-loss prevention, testing and evaluation surfaces all matter and are not scored. They were excluded to keep the axis focused on the accountable individual rather than the administrator, which is the distinction this method exists to draw.
- Weighting is uniform and that is a choice. All six criteria count equally. A reasonable person could argue reversal matters more than provenance. We publish the per-cell scores so anyone who thinks so can reweight them.
- No hands-on testing. A documented control may work poorly. An undocumented one may work well. This method cannot distinguish either case.
Corrections
We expect to be wrong in individual cells and we would like to know where.
To contest a score, cite the public documentation page that supports a different reading. We will re-check against that page and either amend the cell or explain why it does not change the score. Both outcomes are logged with a date in the revision history, and amendments are shown alongside what they replaced rather than silently overwriting them.
We do not remove scores on request, and we do not accept private or unpublished evidence — if it is not publicly documented, it does not change a score, because the score measures what is publicly documented.
Send a correction through the contact form.
Reusing this method
The criteria on this page are published so they can be applied to platforms we did not assess. If you use them, we ask two things: score against public documentation rather than impressions, and publish your evidence per cell. Attribution to auxfirst is welcome and not required.
Revision history
| Version | Date | Change |
|---|---|---|
| v1.1 | 21 Aug 2026 | First publication. Adds product-identity verification as a prerequisite step, states the withdrawal rule for entries that cannot meet the evidence standard, and records limitation 1 as observed rather than anticipated. |
| v1.0 | 20 Aug 2026 | Internal draft. Two axes, twelve criteria, both scoring rules. Not published. |