The Agent Supervision Register
Twelve enterprise platforms, scored on two axes: what their AI agents can do, and what a human can do about it afterwards.
Twelve platforms were assessed in August 2026 against two six-part scales: capability — what the agent can do without a person present — and supervision — whether the person accountable for the work can see, approve and reverse what the agent did. Each criterion scores 0 to 2, for a maximum of 12 per axis.
All scores come from public vendor documentation, retrieved and dated during the assessment window. No products were tested. The two scores are reported separately and never combined, because the difference between them is what the register exists to show.
Eleven platforms are scored across 132 cells. One — Klaviyo — is recorded as assessed but unverified, and is excluded from the medians. The reasons are stated in its entry rather than hidden.
That gap is not an accident of this particular sample. It is what happens when a category is sold on autonomy and bought on trust, and the two are specified by different teams. Every vendor here can tell you what their agent is capable of. Fewer than half can tell you what the person whose work it changed will see afterwards.
How to read a score
- Capability — context reach, write authority, unattended execution, multi-step reasoning, delegation, extensibility.
- Supervision — agent identity and disclosure, action receipt, reversal, consequence-scaled approval, escalation handoff, provenance at the decision point.
- 0 means not found in public documentation. It does not mean the product cannot do it.
- 1 means partial, administrator-only, or documented as advice to a builder rather than platform behavior.
- 2 means documented and available to the person accountable for the work.
Two rules do most of the work, and both are stated in full on the method page. A recommendation is not a control — guidance published to developers about what their agents should do is not a thing the platform does. And a receipt has an audience — a run log a non-engineer can read scores 1 if the only people permissioned to read it are the agent's owner, an administrator, or a security team.
What the agent can do, against what you can do about it
It is not a claim that a platform is unsafe, and it is not a recommendation against buying. It records that a specific control could not be found in public documentation, which is the same information a buyer has before signing. A vendor with a strong internal control and thin public documentation scores below a weaker vendor who writes things down. That is a measurement of documentation, and this register says so rather than obscuring it.
Summary
Sorted alphabetically, not ranked. ? marks a score resting on partial evidence; the entry says which cells and why.
Eleven scored platforms, and one recorded as unverified
| Platform | Product assessed | Capability | Supervision | Gap | Assessed |
|---|---|---|---|---|---|
| Adobe | CX Enterprise Coworker | 11 | 5? | 6 | 21 Aug 2026 |
| Atlassian | Rovo Agents | 11 | 5 | 6 | 20 Aug 2026 |
| ClickUp | Super Agents | 11 | 6 | 5 | 20 Aug 2026 |
| Google Workspace | Workspace Studio | 11 | 7 | 4 | 21 Aug 2026 |
| HubSpot | Agent Hub | 11? | 8? | 3 | 21 Aug 2026 |
| Microsoft | Copilot Studio agents | 12 | 4 | 8 | 21 Aug 2026 |
| monday.com | monday agents | 12 | 9 | 3 | 20 Aug 2026 |
| Notion | Custom Agents | 11 | 5 | 6 | 21 Aug 2026 |
| Salesforce | Agentforce | 11 | 9 | 2 | 20 Aug 2026 |
| ServiceNow | ServiceNow AI Agents | 12 | 5 | 7 | 21 Aug 2026 |
| Slack | Agents on Slack | 11 | 7 | 4 | 20 Aug 2026 |
| Klaviyo | Composer | — | — | — | Unverified |
The entries
Each entry gives twelve scored cells with the finding behind each one, the notes that qualify them, the pages consulted, and a confidence marker. Where a cell rests on weaker evidence, the entry says so in the cell rather than in a footnote.
Adobe
Product assessed: CX Enterprise Coworker — a new tier above Adobe Experience Platform Agents, which remain in production and are not scored here
The phrase “with human oversight” appears identically across Adobe release notes, product pages and event listings. It is positioning; no documented per-action review surface was found behind it. Scored on the documented mechanism, not the phrase — Adobe may have the surface and simply not document it publicly.
Corrected 21 Aug 2026. An earlier draft of this entry described CX Enterprise Coworker as the successor to Experience Platform Agents, and dated it to a product description effective 30 July 2026. Both were wrong. Coworker is a new tier that became generally available on 10 June 2026; the ten-plus Experience Platform Agents introduced at Summit 2025 are still shipping and still in production underneath it. Adobe’s own launch material does not describe them as replaced. The scored cells describe Coworker and are unaffected; what changed is what the score is attached to. Experience Platform Agents are not scored in this edition and are candidates for the next.
Partially verifiedSix of twelve cells rest on search-level evidence. This entry requires a full documentation pass before it is treated as settled.
Atlassian
Product assessed: Rovo Agents
Rovo has the most architecturally developed agent design in the register and the most restricted receipt. The audit trail is scoped to administration rather than work, and gated behind a separate security product.
Parallel subagent delegation, claimed in some third-party coverage, was not found in Atlassian documentation and is not scored.
Partially verifiedVerified on the audit-log and permission cells. Identity and provenance cells rest on documentation titles and summaries rather than fetched pages.
ClickUp
Product assessed: Super Agents
ClickUp has the best-built action receipt in the register. Its localized documentation URLs refer to the feature as the Super Agent audit log.
The product name “Brain²”, which circulates in several published roundups, does not appear in ClickUp’s documentation. The products are ClickUp Brain, Brain MAX and the AI Hub.
VerifiedAll twelve cells trace to fetched pages.
Google Workspace
Product assessed: Workspace Studio
The strongest single release in the register, and the newest — rollout to Rapid Release domains began 20 August 2026, the day of assessment. Feature availability will vary by domain and release track for some weeks.
Google states that at Studio’s original launch it could assist with tasks such as drafting an email but could not execute them autonomously. Any capability claim about Workspace Studio dated before August 2026 describes a different product.
The receipt propagating into the artefact’s own audit trail rather than a separate agent console is, at time of assessment, unique in this register. Agent identities run least-privileged with unique auditable identifiers.
VerifiedSeveral controls are described as rolling out or in beta; re-check at the next revision.
HubSpot
Product assessed: Agent Hub (the agents, renamed from Breeze Agents) and Agent Builder (renamed from Breeze studio)
The most provisional scored entry here. Three supervision cells rest on documentation pages retrieved through search rather than fetched, and the underlying product has churned heavily: custom assistants were retired on 13 July 2026 and a further set of agents on 23 July 2026. Agent Builder was in beta at assessment.
Corrected 21 Aug 2026. This entry previously used the names “Breeze agents” and “the agent builder”. HubSpot made two renames in this area and they are easily transposed: the agents became Agent Hub, and the builder became Agent Builder. Breeze remains the umbrella brand over HubSpot’s AI generally. No scored cell changed; the products were correctly identified and incorrectly named.
Partially verifiedTreat every HubSpot cell as provisional pending a full pass against current documentation.
Klaviyo
Product assessed: Composer — a new agent launched March 2026, shipping alongside the K:AI Marketing Agent rather than replacing it
This entry previously carried a capability score of 6 and a supervision score of 7, and described Composer as the K:AI Marketing Agent renamed. Both the relationship and the confidence were wrong.
Composer is a new product, launched March 2026 and in public beta from 30 June 2026, sitting alongside the earlier K:AI Marketing Agent. K:AI is the umbrella over both, plus the separate Customer Agent. That matters here more than it would elsewhere: this entry’s own confidence line recorded that no page was fetched in full during the assessment pass. If cells were scored from Marketing Agent documentation and then attached to Composer, they describe the wrong product — and correcting the prose would not correct the numbers.
Rather than publish scores that cannot be defended, the entry is withdrawn from the summary table and excluded from both medians until it can be re-fetched and rescored. The observations below are retained as recorded, unscored, so the correction is visible rather than silent.
What was observed, unscored. Composer builds complete campaigns, flows and forms as drafts and cannot publish them; execution requires a person. AI-generated content is flagged to the operator with an instruction to review before use. Output is staged as ordinary campaign objects in the account, with no distinct agent run log documented. Quality and brand evaluations run before output reaches the reviewer.
The idea worth keeping. Klaviyo was the only platform assessed where supervision exceeded capability, and it produced the clearest demonstration in this register that the two axes trade against each other by design rather than by accident. It also exposed a genuine flaw in the rubric: a platform that never executes without approval has nothing to reverse, and takes a zero on reversal for a gap it does not structurally have. That limitation is now stated on the method page and survives this entry’s withdrawal.
UnverifiedNot scored. Re-fetch required before this entry returns to the register.
Microsoft
Product assessed: Autonomous agents in Copilot Studio, with Agent 365 as the governance layer
The widest gap in the register: full marks on capability, four on supervision. The mechanism exists in quantity and is aimed at administrators. Microsoft’s own guidance documents the correct supervision design — keep a human in the loop for high-stakes tasks, maintain detailed logs of triggers received, decisions made and actions taken — as advice to the person building the agent. Under this register’s scoring rule, published guidance is not a platform control.
Licensing, corrected 21 Aug 2026. Agent 365 has been generally available for the Commercial segment since 1 May 2026, licensed per user. An earlier draft described Microsoft E5 as the recommended prerequisite. It is required for new purchases from 1 June 2026, and the requirement varies by segment: E5 for enterprise, F5-level Defender and Purview for frontline, Business Premium for SMB. Microsoft 365 E7 bundles E5, Agent 365, Copilot and the Entra Suite. No scored cell changed — licensing is not a supervision control — but the distinction between recommended and required is the difference between a preference and a budget line.
VerifiedAll twelve cells trace to fetched pages.
monday.com
Product assessed: monday agents
The highest combined position in the register, and the only vendor gating write access by default rather than by configuration.
A documentation contradiction worth flagging. The same monday support article states both that undo covers configuration actions only and, in a later section, that an agent’s actions can be viewed and undone from the activity log. These cannot both be true as written. The cell is scored on the more restrictive statement pending clarification.
Agents were in gradual release at assessment; agent creation requires admin permissions. Agent Factory is a separate standalone monday product outside monday.com and is not assessed here.
VerifiedAll twelve cells trace to a fetched page, with one internal contradiction noted above.
Notion
Product assessed: Custom Agents
Notion’s public framing describes agent runs as reviewable and reversible. Both are true of the agent’s configuration. Neither is documented for the work the agent performed.
The receipt cell is the clearest instance in the register of supervision permissioned to the agent’s owner rather than the affected party.
Custom Agents require a Business or Enterprise plan. Notion Agent, the on-demand assistant, is a separate product and is not scored here.
VerifiedAll twelve cells trace to a fetched page.
Salesforce
Product assessed: Agentforce
The only platform in the register scoring full marks on escalation, and the only one that documents disclosure and handoff as patterns the vendor implements rather than recommends. Both are inherited from customer-service software, where handoff has been a designed problem for decades.
“Subagents”, claimed in some third-party coverage, is not the documented primitive. The architecture is agent type → topics → topic instructions → actions.
VerifiedAll twelve cells trace to a fetched page.
ServiceNow
Product assessed: ServiceNow AI Agents, with AI Control Tower as the governance layer
A naming trap worth stating plainly. ServiceNow’s Intelligent Approvals is AI performing approvals — reading policies in plain language, evaluating incoming requests and handling routine decisions instantly — not humans approving AI. Any assessment that scores it as human oversight has it backwards.
Corrected 21 Aug 2026. An earlier draft described Otto as a straight rename of Now Assist and stated that the rename was cosmetic, with technical identifiers, table names and field names unchanged. The rename is real — announced 5 May 2026 at Knowledge 2026 — but it is broader than described: Otto consolidates three previously separate brands, Now Assist, Moveworks and AI Experience, into one assistant surface, and carries orchestration responsibilities alongside assistance. The “cosmetic” claim could not be verified against ServiceNow documentation and has been withdrawn along with the source cited for it. No scored cell changed: Otto is the assistant layer, and this entry scores ServiceNow AI Agents.
Pending verification. ServiceNow stated in May 2026 that AI Control Tower enhancements would enter general availability in August 2026. That was a forward-looking statement and is recorded as pending, not scored as shipped. Re-check at the next revision.
Partially verifiedVerified on governance and naming. Two capability cells rest on evidence gathered for the supervision axis.
Slack
Product assessed: Agents on Slack
Slack publishes the most complete articulation of agent supervision design found anywhere in this register — identity, approval gates, undo flows, inspectable state, progressive authority — and publishes almost all of it as guidance to third-party developers rather than as platform behavior.
Under this register’s scoring rule, a recommendation is not a control. The rule was defined before scoring and applied identically to all platforms; it reduced Slack’s score by three points and reduced no other platform’s by more than one. That asymmetry is a measurement, not a penalty — Slack is the platform here whose supervision story is most heavily advisory.
Worth reading before you approve an agent into a Slack workspace: the agents reachable there are frequently not Slack’s, and what they do about identity and undo is the third-party builder’s decision.
VerifiedAll twelve cells trace to a fetched page.
Not assessed
Fifteen platforms were considered and excluded from this edition to keep the assessment to a depth that could be evidenced per cell: Asana, Airtable, Smartsheet, Miro, Guru, Zoom, Zoho, Zendesk, Intercom, Freshworks, SAP, Oracle, Workday, Rippling, ActiveCampaign.
Two products inside assessed vendors are also unscored, and both are worth naming because customers are using them today: Adobe Experience Platform Agents, which are still in production underneath CX Enterprise Coworker, and Klaviyo’s K:AI Marketing Agent, which still ships alongside Composer. Both are candidates for the next edition.
Absence is not a judgment.
Method, in short
Public product documentation, admin and governance documentation, developer documentation, release notes and official product descriptions, published by the vendor, retrieved and dated in August 2026. No hands-on testing, no analyst reports, no third-party reviews, no vendor briefings, and no vendor contacted. This is a documentation audit, and the scores describe what a buyer can verify before signing anything — deliberately the same information a buyer has.
A zero records that the capability was not found in public vendor documentation after searching from at least two distinct angles. It does not mean the capability is absent from the product.
The full method — criteria definitions, both scoring rules, the evidence standard and six known limitations including the one this edition’s Klaviyo entry exposed — is published separately so it can be checked, argued with, and applied to platforms not assessed here.
Corrections
To contest a score, cite the public documentation page supporting a different reading. We re-check against that page and either amend the cell or explain why it does not change the score. Both outcomes are logged with a date, and amendments are shown alongside what they replaced rather than silently overwriting them — as they are in the Adobe, HubSpot, Klaviyo, Microsoft and ServiceNow entries above.
We do not remove scores on request, and we do not accept private or unpublished evidence — if it is not publicly documented, it does not change a score, because the score measures what is publicly documented.
Send a correction through the contact form.
Revision history
| Date | Change |
|---|---|
| 21 Aug 2026 | First publication. Method v1.1. Eleven platforms scored across 132 cells; Klaviyo recorded as unverified and excluded from the medians after re-verification found its evidence could not be tied to the product it was attached to. Product identity corrected on five entries: Adobe (Coworker is a new tier, not a successor; GA 10 June 2026), HubSpot (Agent Hub and Agent Builder, both renamed), ServiceNow (Otto consolidates three brands; the “cosmetic rename” claim withdrawn as unverifiable), Microsoft (E5 required, not recommended), Klaviyo (Composer is new, not renamed). No scored cell moved as a result of these corrections; what moved is what the scores are attached to. |
It is not a buying recommendation, a ranking, or a safety assessment. It is a record of what twelve vendors have written down about supervision, scored consistently, dated, and published with its own errors visible. If you need to choose between two of these platforms, the scores are an input to that conversation and not the answer to it — the agent buyer’s map is the ten-dimension version of the wider question.
Where this goes next
The pattern across eleven platforms is consistent enough to state plainly: the industry has shipped autonomy and deferred accountability. Every vendor here can tell you what their agent can do. Most cannot tell you what the person whose work it changed will see afterwards, and none documents how to undo it.
If you are choosing between these platforms, start with the naming problem — the AI layers name index establishes what each product is actually called, including four renames and two new tiers widely misreported as renames. If you are past choosing and into rollout, the agent development lifecycle is the long-form field guide, and the Action Heat Ladder is the model behind this register’s consequence-scaled approval criterion.
And if the gap in Figure 1 describes a workflow you are already running, that is the thing worth auditing before it is the thing worth explaining. The Agent Operability Audit is one workflow, three weeks, fixed scope.