# The Agent Supervision Register

**Twelve enterprise platforms, scored on two axes: what their AI agents can do, and what a human can do about it afterwards. Assessed against public vendor documentation, August 2026.**

Last updated: 21 August 2026 · Method v1.1 · auxfirst
Canonical: https://auxfirst.com/news/agent-supervision-register.html
Method: https://auxfirst.com/news/agent-supervision-method.html

| | |
|---|---|
| Platforms scored | 11 of 12 |
| Scored cells | 132 |
| Median capability | 11 / 12 |
| Median supervision | 6 / 12 |

---

## What this covers

Twelve platforms were assessed in August 2026 against two six-part scales: **capability** — what the agent can do without a person present — and **supervision** — whether the person accountable for the work can see, approve and reverse what the agent did. Each criterion scores 0–2, for a maximum of 12 per axis.

All scores come from public vendor documentation, retrieved and dated during the assessment window. No products were tested. **The two scores are reported separately and never combined**, because the difference between them is what the register exists to show.

Eleven platforms are scored across 132 cells. One — Klaviyo — is recorded as assessed but unverified, and is excluded from the medians. The reasons are stated in its entry rather than hidden.

> Median capability, 11 of 12. Median supervision, 6. Not one of the eleven documents a way to reverse what an agent did to your work.

That gap is not an accident of this particular sample. It is what happens when a category is sold on autonomy and bought on trust, and the two are specified by different teams.

## How to read a score

- **Capability** — context reach, write authority, unattended execution, multi-step reasoning, delegation, extensibility.
- **Supervision** — agent identity and disclosure, action receipt, reversal, consequence-scaled approval, escalation handoff, provenance at the decision point.
- **0** means not found in public documentation. It does not mean the product cannot do it.
- **1** means partial, administrator-only, or documented as advice to a builder rather than platform behavior.
- **2** means documented and available to the person accountable for the work.
- **?** marks a cell resting on partial evidence. **!** marks a cell resting on secondary sources.

Two rules do most of the work. **A recommendation is not a control** — guidance published to developers about what their agents should do is not a thing the platform does. And **a receipt has an audience** — a run log a non-engineer can read scores 1 if the only people permissioned to read it are the agent's owner, an administrator, or a security team.

## Summary

| Platform | Product assessed | Capability | Supervision | Gap | Assessed |
|---|---|---|---|---|---|
| Adobe | CX Enterprise Coworker | 11 | 5 ? | 6 | 21 Aug 2026 |
| Atlassian | Rovo Agents | 11 | 5 | 6 | 20 Aug 2026 |
| ClickUp | Super Agents | 11 | 6 | 5 | 20 Aug 2026 |
| Google Workspace | Workspace Studio | 11 | 7 | 4 | 21 Aug 2026 |
| HubSpot | Agent Hub | 11 ? | 8 ? | 3 | 21 Aug 2026 |
| Microsoft | Copilot Studio agents | 12 | 4 | 8 | 21 Aug 2026 |
| monday.com | monday agents | 12 | 9 | 3 | 20 Aug 2026 |
| Notion | Custom Agents | 11 | 5 | 6 | 21 Aug 2026 |
| Salesforce | Agentforce | 11 | 9 | 2 | 20 Aug 2026 |
| ServiceNow | ServiceNow AI Agents | 12 | 5 | 7 | 21 Aug 2026 |
| Slack | Agents on Slack | 11 | 7 | 4 | 20 Aug 2026 |
| *Klaviyo* | *Composer* | *—* | *—* | *—* | *Unverified* |

Sorted alphabetically, not ranked. Median capability 11. Median supervision 6.

Ranked by gap, largest first: Microsoft 8 · ServiceNow 7 · Adobe 6 · Atlassian 6 · Notion 6 · ClickUp 5 · Google Workspace 4 · Slack 4 · HubSpot 3 · monday.com 3 · Salesforce 2.

---

# The entries

## Adobe

**Product assessed:** CX Enterprise Coworker — a new tier above Adobe Experience Platform Agents, which remain in production and are not scored here

**Capability 11/12 · Supervision 5/12 · Gap 6**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Operates across Experience Platform and **Adobe CX Enterprise** applications — the brand renamed from Experience Cloud at Summit 2026 — grounded in customer data. |
| Write authority | 2 | Adobe describes Coworker as planning, executing and validating work, then returning finished output. |
| Unattended execution | 1 | Goal-driven execution documented; scheduled or event-triggered running without a person not found. |
| Multi-step reasoning | 2 | Plan–execute–validate loop described explicitly. |
| Delegation | 2 | Agent Orchestrator selects and coordinates specialist agents. |
| Extensibility | 2 | Agent Composer supports customer-defined grounding, guardrails and bring-your-own agents. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 1 | Agents are individually named; no documented requirement to disclose agent authorship to the person affected. |
| Action receipt | 1 ? | Described as auditable with version management; no per-run detail surface located in public documentation. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 1 | Approval is documented at job completion — work is returned for approval — rather than gated per action by consequence. |
| Escalation handoff | 0 | Not documented. |
| Provenance at decision point | 2 | Adobe documents source citation and grounding verification in agent responses, with explanations returned alongside answers and access controls respected. |

The phrase “with human oversight” appears identically across Adobe release notes, product pages and event listings. It is positioning; no documented per-action review surface was found behind it. Scored on the documented mechanism, not the phrase — Adobe may have the surface and simply not document it publicly.

**Corrected 21 Aug 2026.** An earlier draft of this entry described CX Enterprise Coworker as the *successor* to Experience Platform Agents, and dated it to a product description effective 30 July 2026. Both were wrong. Coworker is a new tier that became generally available on **10 June 2026**; the ten-plus Experience Platform Agents introduced at Summit 2025 are still shipping and still in production underneath it. Adobe’s own launch material does not describe them as replaced. The scored cells describe Coworker and are unaffected; what changed is what the score is attached to. Experience Platform Agents are not scored in this edition and are candidates for the next.

**Sources consulted:** **Sources consulted:** Adobe Experience Platform agent documentation; Adobe newsroom, CX Enterprise and Coworker announcements; Adobe security and trust overview. Individual page URLs were not captured during the assessment pass and are recorded as outstanding.
**Confidence:** Partially verifiedSix of twelve cells rest on search-level evidence. This entry requires a full documentation pass before it is treated as settled.

## Atlassian

**Product assessed:** Rovo Agents

**Capability 11/12 · Supervision 5/12 · Gap 6**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Atlassian products plus connected third-party sources. |
| Write authority | 2 | Agents act on Jira and Confluence content within the invoking user’s permissions. |
| Unattended execution | 1 | Invocable from automation rules, but an agent used in an automation flow is limited to a text response and cannot use its own tools. |
| Multi-step reasoning | 2 | Agents plan across tools and knowledge. |
| Delegation | 2 | Parent agents delegate to subagents. |
| Extensibility | 2 | Rovo Studio is available to everyone in an organization by default; creation rights can be restricted. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Agent identity — name and description — is a required configuration field determining how the agent presents to users. |
| Action receipt | 1 | The audit log records agent lifecycle events — chat started, agent created, updated or deleted, connector changes — **not the work the agent performed**. It also requires Atlassian Guard Standard for admin actions and Atlassian Guard Premium for user actions, both separate paid products. |
| Reversal | 0 | Not documented as an agent mechanism. |
| Consequence-scaled approval | 1 | Permission inheritance means an agent can only do what the invoking user can do. That is access control, not consequence scaling. |
| Escalation handoff | 0 | Not documented in Rovo. Human routing exists in Jira Service Management, a separate product, and is not scored here. |
| Provenance at decision point | 1 | Debugging and live-conversation review surfaces are documented; provenance shown at the point of human review was not found. |

Rovo has the most architecturally developed agent design in the register and the most restricted receipt. The audit trail is scoped to administration rather than work, and gated behind a separate security product.

Parallel subagent delegation, claimed in some third-party coverage, was not found in Atlassian documentation and is not scored.

**Sources consulted:** **Sources consulted:** [Rovo data usage and privacy guide](https://www.atlassian.com/software/rovo/guides/admin-guide/rovo-data-usage-privacy) (fetched 20 Aug 2026); Rovo agent and Rovo Studio support documentation.
**Confidence:** Partially verifiedVerified on the audit-log and permission cells. Identity and provenance cells rest on documentation titles and summaries rather than fetched pages.

## ClickUp

**Product assessed:** Super Agents

**Capability 11/12 · Supervision 6/12 · Gap 5**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Tasks, Docs, Chat, wikis, connected apps and web. |
| Write authority | 2 | Creates and updates tasks, docs, lists and fields directly. |
| Unattended execution | 2 | Schedules and triggers supported; agents run without a person present. |
| Multi-step reasoning | 2 | Multi-step execution with persistent memory. |
| Delegation | 1 | Multiple agent types coexist; delegation between agents not documented. |
| Extensibility | 2 | Super Agent Builder configures agents through natural language, with a prebuilt catalog. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Agents have manageable profiles and are identified by name in the activity record. |
| Action receipt | 2 | The Super Agents Activity view records time run, agent name, location, trigger, who ran it and status, expandable to show which tools ran and when, with the agent’s own explanation of each action. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 1 | Documentation states agents seek human approval for critical decisions; no documented configuration of what counts as critical. |
| Escalation handoff | 0 | Not documented as a mechanism. Escalation appears as a use case, not a control. |
| Provenance at decision point | 1 | The agent’s reasoning is available after the run, not at the point of review. |

ClickUp has the best-built action receipt in the register. Its localized documentation URLs refer to the feature as the Super Agent audit log.

The product name “Brain²”, which circulates in several published roundups, does not appear in ClickUp’s documentation. The products are ClickUp Brain, Brain MAX and the AI Hub.

**Sources consulted:** **Sources consulted:** [What are Super Agents?](https://help.clickup.com/hc/en-us/articles/31010910371991-What-are-Super-Agents) and [Super Agents Activity](https://help.clickup.com/hc/en-us/articles/36455020912919-Super-Agents-Activity) (both fetched 20 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to fetched pages.

## Google Workspace

**Product assessed:** Workspace Studio

**Capability 11/12 · Supervision 7/12 · Gap 4**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Gmail, Drive, Sheets and third-party applications, with webhook integration. |
| Write authority | 2 | Since August 2026, flows execute tasks such as sending email rather than only drafting them. |
| Unattended execution | 2 | Flows run on triggers, including across users. |
| Multi-step reasoning | 2 | Multi-step flows built in natural language. |
| Delegation | 1 | Multi-step flows documented; agent-to-agent delegation not found. |
| Extensibility | 2 | No-code natural-language builder available to end users. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Identity attribution controls whether flow actions show the owner’s identity or are attributed to the flow with owner information visible. Default is attribution to the flow. |
| Action receipt | 2 | Studio audit events cover configuration and execution, and **downstream audit events** — a Drive edit, an email sent — carry flow context including a unique flow identifier and the owner’s information. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 2 | End-user confirmation can be enforced for steps that share data externally, and data-loss prevention supports blocking or enforcing end-user review based on the sourced data, the data used, and the visibility of the output. |
| Escalation handoff | 0 | Not documented. |
| Provenance at decision point | 1 | The approval prompt surfaces the pending action; source provenance and confidence at the point of review not documented. |

The strongest single release in the register, and the newest — rollout to Rapid Release domains began 20 August 2026, the day of assessment. Feature availability will vary by domain and release track for some weeks.

Google states that at Studio’s original launch it could assist with tasks such as drafting an email but could not execute them autonomously. **Any capability claim about Workspace Studio dated before August 2026 describes a different product.**

The receipt propagating into the artefact’s own audit trail rather than a separate agent console is, at time of assessment, unique in this register. Agent identities run least-privileged with unique auditable identifiers.

**Sources consulted:** **Sources consulted:** Google Workspace Updates, [New enterprise security controls for Workspace Studio](https://workspaceupdates.googleblog.com/2026/08/new-enterprise-security-controls-for-Workspace-Studio-enable-expanded-collaboration-use-cases.html) (published 17 Aug 2026, fetched 20 Aug 2026; re-verified 21 Aug 2026).
**Confidence:** VerifiedSeveral controls are described as rolling out or in beta; re-check at the next revision.

## HubSpot

**Product assessed:** Agent Hub (the agents, renamed from Breeze Agents) and Agent Builder (renamed from Breeze studio)

**Capability 11/12 · Supervision 8/12 · Gap 3**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | CRM data plus knowledge sources and external system connections. |
| Write authority | 2 | Agents write to CRM records. |
| Unattended execution | 2 | Agents run on defined triggers and inputs. |
| Multi-step reasoning | 2 | Multi-action agents with instructions and knowledge. |
| Delegation | 1 | Multiple agents; delegation between them not documented. |
| Extensibility | 2 | Agents defined in natural language in Agent Builder. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 1 | Agents are named objects with access controls; no documented disclosure requirement to the affected party. |
| Action receipt | 2 ? | Run history and an agent inbox with per-action detail, including the sources used for each action. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 2 ? | A review-before-running control on tools that write to the CRM. |
| Escalation handoff | 1 ! | Routing of complex cases to a human with conversation history; secondary sources only. |
| Provenance at decision point | 2 ? | Sources shown per action in the run detail — the most granular provenance surface located in this register. |

The most provisional scored entry here. Three supervision cells rest on documentation pages retrieved through search rather than fetched, and the underlying product has churned heavily: custom assistants were retired on 13 July 2026 and a further set of agents on 23 July 2026. Agent Builder was in beta at assessment.

**Corrected 21 Aug 2026.** This entry previously used the names “Breeze agents” and “the agent builder”. HubSpot made *two* renames in this area and they are easily transposed: the agents became **Agent Hub**, and the builder became **Agent Builder**. Breeze remains the umbrella brand over HubSpot’s AI generally. No scored cell changed; the products were correctly identified and incorrectly named.

**Sources consulted:** **Sources consulted:** [Create and customize agents in the agent builder](https://knowledge.hubspot.com/ai/create-and-customize-agents-in-the-agent-builder) (page last updated 23 July 2026, fetched 20 Aug 2026; names re-verified 21 Aug 2026); earlier agent-output documentation via search.
**Confidence:** Partially verifiedTreat every HubSpot cell as provisional pending a full pass against current documentation.

## Klaviyo

**Product assessed:** Composer — a new agent launched March 2026, shipping alongside the K:AI Marketing Agent rather than replacing it

> **Withdrawn from scoring — 21 August 2026**
> This entry previously carried a capability score of 6 and a supervision score of 7, and described Composer as the K:AI Marketing Agent renamed. Both the relationship and the confidence were wrong.
>
> Composer is a **new product**, launched March 2026 and in public beta from 30 June 2026, sitting alongside the earlier K:AI Marketing Agent. K:AI is the umbrella over both, plus the separate Customer Agent. That matters here more than it would elsewhere: this entry’s own confidence line recorded that *no page was fetched in full during the assessment pass*. If cells were scored from Marketing Agent documentation and then attached to Composer, they describe the wrong product — and correcting the prose would not correct the numbers.
>
> Rather than publish scores that cannot be defended, the entry is withdrawn from the summary table and excluded from both medians until it can be re-fetched and rescored. The observations below are retained as recorded, unscored, so the correction is visible rather than silent.
>

**What was observed, unscored.** Composer builds complete campaigns, flows and forms as drafts and cannot publish them; execution requires a person. AI-generated content is flagged to the operator with an instruction to review before use. Output is staged as ordinary campaign objects in the account, with no distinct agent run log documented. Quality and brand evaluations run before output reaches the reviewer.

**The idea worth keeping.** Klaviyo was the only platform assessed where supervision exceeded capability, and it produced the clearest demonstration in this register that the two axes trade against each other by design rather than by accident. It also exposed a genuine flaw in the rubric: *a platform that never executes without approval has nothing to reverse, and takes a zero on reversal for a gap it does not structurally have.* That limitation is now stated on the [method page](agent-supervision-method.html#limits) and survives this entry’s withdrawal.

**Sources consulted:** **Sources consulted:** Klaviyo Composer and K:AI product pages; Klaviyo newsroom, March and June 2026. No page fetched in full during this pass.
**Confidence:** UnverifiedNot scored. Re-fetch required before this entry returns to the register.

## Microsoft

**Product assessed:** Autonomous agents in Copilot Studio, with Agent 365 as the governance layer

**Capability 12/12 · Supervision 4/12 · Gap 8**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Microsoft Graph, hundreds of connectors, and MCP-based tools. |
| Write authority | 2 | Agents execute tasks against connected systems. |
| Unattended execution | 2 | Microsoft’s own description: agents perceive events, make decisions and execute tasks independently using triggers, instructions and guardrails, operating continuously in the background — monitoring data, reacting to conditions and running workflows at scale — rather than responding only in conversations. |
| Multi-step reasoning | 2 | Planning and adaptive execution documented. |
| Delegation | 2 | Multi-agent patterns and agent-to-agent orchestration documented. |
| Extensibility | 2 | Copilot Studio is a low-code builder. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Agents receive directory identities through Microsoft Entra Agent ID and appear in an agent registry. |
| Action receipt | 1 | Agent 365’s observe pillar and its Purview and Defender integration provide **tenant-level telemetry to administrators**. No receipt surface for the person whose work an agent changed. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 1 | Microsoft’s guidance recommends configuring agents to request approval before sensitive actions, and administrators approve agents at deployment. The per-action gate is something the maker builds, not something the platform enforces. |
| Escalation handoff | 0 | Not documented. |
| Provenance at decision point | 0 | Not documented. |

**The widest gap in the register:** full marks on capability, four on supervision. The mechanism exists in quantity and is aimed at administrators. Microsoft’s own guidance documents the correct supervision design — keep a human in the loop for high-stakes tasks, maintain detailed logs of triggers received, decisions made and actions taken — as advice to the person building the agent. Under this register’s scoring rule, published guidance is not a platform control.

**Licensing, corrected 21 Aug 2026.** Agent 365 has been generally available for the Commercial segment since 1 May 2026, licensed per user. An earlier draft described Microsoft E5 as the *recommended* prerequisite. It is **required** for new purchases from 1 June 2026, and the requirement varies by segment: E5 for enterprise, F5-level Defender and Purview for frontline, Business Premium for SMB. Microsoft 365 E7 bundles E5, Agent 365, Copilot and the Entra Suite. No scored cell changed — licensing is not a supervision control — but the distinction between recommended and required is the difference between a preference and a budget line.

**Sources consulted:** **Sources consulted:** [Microsoft Agent 365 overview](https://learn.microsoft.com/en-us/microsoft-agent-365/overview) and [Autonomous agents guidance, Copilot Studio](https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/autonomous-agents) (both fetched 20 Aug 2026; licensing re-verified 21 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to fetched pages.

## monday.com

**Product assessed:** monday agents

**Capability 12/12 · Supervision 9/12 · Gap 3**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Boards, docs, business context and connected third-party applications. |
| Write authority | 2 | Agents create and update items, columns and workflows. |
| Unattended execution | 2 | Triggered and scheduled execution; agents monitor activity. |
| Multi-step reasoning | 2 | Agents apply judgment within guardrails across multi-step work. |
| Delegation | 2 | Multiple agents plus external and MCP-connected agents. |
| Extensibility | 2 | In-product agent builder configured through chat and settings. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | In live agent mode the agent’s avatar moves through the board as a collaborator would; the activity record identifies which agent acted. |
| Action receipt | 2 | The Activity tab records run date, apps used, credits consumed, and what the agent did and why, with a dedicated panel explaining failures. |
| Reversal | 1 | Undo covers **configuration** actions only, not the content the agent produced — for example an update it wrote on an item. |
| Consequence-scaled approval | 2 | Read tools activate on connection; **write tools are added inactive** and require explicit activation after review. Per-tool guardrails restrict scope — posting only in a named channel, emailing only named contacts. |
| Escalation handoff | 1 | Agents can be instructed to ask before acting; documented human handoff with context is not present in the work-management product. |
| Provenance at decision point | 1 | The agent confirms what it configured before proceeding; run reasoning is available afterwards. |

The highest combined position in the register, and the only vendor gating write access by default rather than by configuration.

**A documentation contradiction worth flagging.** The same monday support article states both that undo covers configuration actions only and, in a later section, that an agent’s actions can be viewed and undone from the activity log. These cannot both be true as written. The cell is scored on the more restrictive statement pending clarification.

Agents were in gradual release at assessment; agent creation requires admin permissions. Agent Factory is a separate standalone monday product outside monday.com and is not assessed here.

**Sources consulted:** **Sources consulted:** [AI Agents on monday.com](https://support.monday.com/hc/en-us/articles/33347027353746-AI-Agents-on-monday-com) (page last modified 18 Aug 2026, fetched 20 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to a fetched page, with one internal contradiction noted above.

## Notion

**Product assessed:** Custom Agents

**Capability 11/12 · Supervision 5/12 · Gap 6**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Workspace content plus connected applications. |
| Write authority | 2 | Creates and edits pages and databases. |
| Unattended execution | 2 | Runs on schedules and workspace events — page creation, page updates, property changes, comments, and Slack triggers. |
| Multi-step reasoning | 2 | Multi-step background workflows. |
| Delegation | 1 | Multiple agents; delegation not documented. |
| Extensibility | 2 | Agents defined by users in natural language. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Agents are named workspace objects, visible in the sidebar and identifiable where they post. |
| Action receipt | 1 | A run log records what triggered each run, the actions taken, errors, and the agent’s reasoning at each step — but it is **permissioned to users with Full Access or edit rights over the agent**, not to the person whose work was changed. |
| Reversal | 1 | Version history restores past versions of the agent’s **configuration** — who changed what, and when — not the content the agent produced. |
| Consequence-scaled approval | 0 | Not documented. |
| Escalation handoff | 0 | Not documented. |
| Provenance at decision point | 1 | Reasoning available after the run. |

Notion’s public framing describes agent runs as reviewable and reversible. Both are true of the agent’s *configuration*. Neither is documented for the work the agent performed.

The receipt cell is the clearest instance in the register of supervision permissioned to the agent’s owner rather than the affected party.

Custom Agents require a Business or Enterprise plan. Notion Agent, the on-demand assistant, is a separate product and is not scored here.

**Sources consulted:** **Sources consulted:** [Custom Agents](https://www.notion.com/help/custom-agents) (fetched 20 Aug 2026, re-verified 21 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to a fetched page.

## Salesforce

**Product assessed:** Agentforce

**Capability 11/12 · Supervision 9/12 · Gap 2**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | CRM, Data Cloud, Flow, Apex and MuleSoft-connected systems. |
| Write authority | 2 | Agents execute actions against CRM objects. |
| Unattended execution | 2 | Agents engage and act on defined conditions without a person present. |
| Multi-step reasoning | 2 | Topic-driven planning across actions. |
| Delegation | 1 | Structured as agent types with topics and actions rather than delegating subagents. |
| Extensibility | 2 | Agents configured through Agentforce Builder. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Documented as a design pattern Salesforce applies: transparency by design, with AI-generated content directly and transparently disclosed. |
| Action receipt | 2 | AI interactions are captured in event logs giving visibility into the results of each user interaction, with an audit trail feature in the Einstein Trust Layer for detailed insight into actions and outcomes. |
| Reversal | 0 | Not documented. |
| Consequence-scaled approval | 2 | Salesforce’s acceptable use policy states AI cannot make legal or important decisions without a human making the final decision. Engagement rules constrain when and how agents may act. |
| Escalation handoff | 2 | The Agentforce Service Agent uses topic instructions to determine when to escalate a conversation to a human representative. Handoff is documented as a design pattern, with examples including copying a manager on AI-generated email and providing a dashboard for human oversight. |
| Provenance at decision point | 1 | Grounding and toxicity detection documented; provenance surfaced at the point of human review not found. |

The only platform in the register scoring full marks on escalation, and the only one that documents disclosure and handoff as patterns *the vendor implements* rather than recommends. Both are inherited from customer-service software, where handoff has been a designed problem for decades.

“Subagents”, claimed in some third-party coverage, is not the documented primitive. The architecture is agent type → topics → topic instructions → actions.

**Sources consulted:** **Sources consulted:** [Explore Agentforce Guardrails and Trust Patterns](https://trailhead.salesforce.com/content/learn/modules/trusted-agentic-ai/explore-agentforce-guardrails-and-trust-patterns), Salesforce Trailhead (fetched 20 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to a fetched page.

## ServiceNow

**Product assessed:** ServiceNow AI Agents, with AI Control Tower as the governance layer

**Capability 12/12 · Supervision 5/12 · Gap 7**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Enterprise workflows plus third-party systems through AI Agent Fabric. |
| Write authority | 2 | Agents act and resolve against workflow records. |
| Unattended execution | 2 | Agents run against enterprise processes without a person present. |
| Multi-step reasoning | 2 | Agents reason, plan and execute; runtime observability covers how they reason and where they decide. |
| Delegation | 2 | AI Agent Orchestrator coordinates teams of agents. |
| Extensibility | 2 | AI Agent Studio builds and customizes agents in natural language. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | Agents are registered as non-human identities with scoped permissions and access mapping. |
| Action receipt | 1 | Runtime observability into agent reasoning and decisions sits in AI Control Tower, aimed at AI stewards and risk and compliance teams. **Administrator-facing**, not available to the person whose work an agent changed. |
| Reversal | 0 | A kill switch stops agents in real time. **Stopping is not reversing.** |
| Consequence-scaled approval | 1 | Governance workflows route approvals for agent lifecycle decisions. Not per-action, and not keyed to consequence. |
| Escalation handoff | 0 | Not documented at the action level. |
| Provenance at decision point | 1 | Observability is monitoring, available after the fact. |

**A naming trap worth stating plainly.** ServiceNow’s Intelligent Approvals is *AI performing approvals* — reading policies in plain language, evaluating incoming requests and handling routine decisions instantly — not humans approving AI. Any assessment that scores it as human oversight has it backwards.

**Corrected 21 Aug 2026.** An earlier draft described Otto as a straight rename of Now Assist and stated that the rename was cosmetic, with technical identifiers, table names and field names unchanged. The rename is real — announced 5 May 2026 at Knowledge 2026 — but it is broader than described: Otto consolidates **three** previously separate brands, Now Assist, Moveworks and AI Experience, into one assistant surface, and carries orchestration responsibilities alongside assistance. The “cosmetic” claim could not be verified against ServiceNow documentation and has been withdrawn along with the source cited for it. No scored cell changed: Otto is the assistant layer, and this entry scores ServiceNow AI Agents.

**Pending verification.** ServiceNow stated in May 2026 that AI Control Tower enhancements would enter general availability in August 2026. That was a forward-looking statement and is recorded as pending, not scored as shipped. Re-check at the next revision.

**Sources consulted:** **Sources consulted:** ServiceNow newsroom, AI Control Tower and Knowledge 2026 announcements; [ServiceNow product documentation](https://www.servicenow.com/docs/) (fetched 20 Aug 2026; Otto naming re-verified 21 Aug 2026).
**Confidence:** Partially verifiedVerified on governance and naming. Two capability cells rest on evidence gathered for the supervision axis.

## Slack

**Product assessed:** Agents on Slack

**Capability 11/12 · Supervision 7/12 · Gap 4**

### Capability

| Criterion | Score | Finding |
|---|---|---|
| Context reach | 2 | Messages, files, channels and connected applications. |
| Write authority | 2 | Agents post, act and operate on connected systems. |
| Unattended execution | 2 | Agents respond to events and run background work. |
| Multi-step reasoning | 2 | Multi-step agent work with visible task state. |
| Delegation | 2 | Multiple third-party agents operate in one workspace. |
| Extensibility | 1 | A developer platform and APIs rather than a no-code builder for end users. |

### Supervision

| Criterion | Score | Finding |
|---|---|---|
| Identity and disclosure | 2 | **Platform-enforced.** Slack’s guidance states an agent should be clearly distinguishable from a human at all times, and that an agent masquerading as a human user breaks trust and complicates auditability. |
| Action receipt | 1 | The Audit Logs API records workspace-level events. Per-action agent logging is published as a recommended metric schema *for builders*, not a platform surface. |
| Reversal | 1 | Slack instructs builders to provide undo flows for meaningful operations and to surface recovery paths clearly, with examples such as re-run with different parameters, undo, and show original. Named as a requirement for builders; not a platform primitive. |
| Consequence-scaled approval | 1 | Slack instructs builders to require confirmation for high-impact actions and to place approval gates before an agent creates, sends or deletes anything, with progressive authority moving from confirm-every-time to always-allow per action class. The components exist; the requirement does not. |
| Escalation handoff | 1 | Explicit fallback behavior is required of builders; human handoff is not documented. |
| Provenance at decision point | 1 | Builders are encouraged to have agents cite or reference sources so users can verify accuracy. Encouraged, not enforced. |

Slack publishes the most complete articulation of agent supervision design found anywhere in this register — identity, approval gates, undo flows, inspectable state, progressive authority — and publishes almost all of it as *guidance to third-party developers* rather than as platform behavior.

Under this register’s scoring rule, a recommendation is not a control. The rule was defined before scoring and applied identically to all platforms; it reduced Slack’s score by three points and reduced no other platform’s by more than one. **That asymmetry is a measurement, not a penalty** — Slack is the platform here whose supervision story is most heavily advisory.

Worth reading before you approve an agent into a Slack workspace: the agents reachable there are frequently not Slack’s, and what they do about identity and undo is the third-party builder’s decision.

**Sources consulted:** **Sources consulted:** [Agent governance](https://docs.slack.dev/ai/agent-governance/), Slack developer documentation (fetched 20 Aug 2026).
**Confidence:** VerifiedAll twelve cells trace to a fetched page.

---

# Not assessed

Fifteen platforms were considered and excluded from this edition to keep the assessment to a depth that could be evidenced per cell: Asana, Airtable, Smartsheet, Miro, Guru, Zoom, Zoho, Zendesk, Intercom, Freshworks, SAP, Oracle, Workday, Rippling, ActiveCampaign.

Two products *inside* assessed vendors are also unscored, and both are worth naming because customers are using them today: **Adobe Experience Platform Agents**, still in production underneath CX Enterprise Coworker, and **Klaviyo's K:AI Marketing Agent**, which still ships alongside Composer. Both are candidates for the next edition.

Absence is not a judgment.

# Method, in short

Public product documentation, admin and governance documentation, developer documentation, release notes and official product descriptions, published by the vendor, retrieved and dated in August 2026. No hands-on testing, no analyst reports, no third-party reviews, no vendor briefings, and no vendor contacted. This is a documentation audit, and the scores describe what a buyer can verify before signing anything — deliberately the same information a buyer has.

A zero records that the capability was not found in public vendor documentation after searching from at least two distinct angles. It does not mean the capability is absent from the product.

Full method, criteria definitions, both scoring rules, the evidence standard and six known limitations: https://auxfirst.com/news/agent-supervision-method.html

# Corrections

To contest a score, cite the public documentation page supporting a different reading. We re-check against that page and either amend the cell or explain why it does not change the score. Both outcomes are logged with a date, and **amendments are shown alongside what they replaced rather than silently overwriting them** — as they are in the Adobe, HubSpot, Klaviyo, Microsoft and ServiceNow entries above.

We do not remove scores on request, and we do not accept private or unpublished evidence — if it is not publicly documented, it does not change a score, because the score measures what is publicly documented.

Send a correction: https://auxfirst.com/index.html#contact

## Revision history

| Date | Change |
|---|---|
| 21 August 2026 | **First publication. Method v1.1.** Eleven platforms scored across 132 cells; Klaviyo recorded as unverified and excluded from the medians after re-verification found its evidence could not be tied to the product it was attached to. Product identity corrected on five entries: Adobe (Coworker is a new tier, not a successor; GA 10 June 2026), HubSpot (Agent Hub and Agent Builder, both renamed), ServiceNow (Otto consolidates three brands; the "cosmetic rename" claim withdrawn as unverifiable), Microsoft (E5 required, not recommended), Klaviyo (Composer is new, not renamed). No scored cell moved as a result of these corrections; what moved is what the scores are attached to. |

**What this register is not.** It is not a buying recommendation, a ranking, or a safety assessment. It is a record of what twelve vendors have written down about supervision, scored consistently, dated, and published with its own errors visible.

# Related

- How we score agent supervision — the full method. https://auxfirst.com/news/agent-supervision-method.html
- AI layers name index — what each product is called. https://auxfirst.com/news/ai-layers-name-index.html
- The Action Heat Ladder. https://auxfirst.com/action-heat-ladder.html
- The agent buyer's map. https://auxfirst.com/news/agent-buyers-map.html
- The agent development lifecycle. https://auxfirst.com/agent-development-lifecycle.html
- The Agent Operability Audit. https://auxfirst.com/agent-operability-audit.html

---

*Emil Krzemiński is the founder of auxfirst (https://auxfirst.com/), the agentic experience design agency. For how machines read auxfirst, see https://auxfirst.com/ai-info.html.*
