Every AI agent dashboard offers you activity by the bucket: sessions held, messages exchanged, minutes engaged. All of it is motion, none of it is work. AI employees should be measured the way employees are, on the outcomes their role exists to produce, and the discipline of insisting on that is the difference between buying results and buying noise.
Why activity metrics seduce, and what they hide
Activity is abundant, easy to chart, and always up and to the right, deploy anything conversational and interactions will climb. The seduction is that activity feels like adoption. What it hides: an agent can hold a thousand pleasant sessions that finish nothing, deflect users who needed resolution, and generate engagement that correlates with frustration as easily as value. The tell is that activity metrics never embarrass anyone, and metrics that can't embarrass can't inform. An agent measured on sessions behaves like a widget; one measured on outcomes behaves like a colleague, and vendors reveal themselves by which dashboard they lead with.
Outcome metrics, role by role
Each role has the number a human in that job would carry. Demo-side: qualified meetings created, demo-to-meeting rate, show rate with context, pipeline influenced. Onboarding: time-to-value, activation rate, coverage (the share of signups getting real sessions), early-cohort retention. Support: resolutions, verified on the user's screen, not deflections; time-to-resolution; reopen rate. Training: users trained, meaning observed performing the workflow unassisted, not completions; downstream, tickets avoided and adoption moved. Note the pattern: every real metric is checkable against a state of the world, the meeting exists, the account activated, the fix held, the skill performed. Activity needs no world at all.
The measurement disciplines that keep you honest
Baseline before deploying, or improvement claims are astrology. Compare cohorts, covered versus uncovered, trained versus untrained, rather than before/after alone, because seasons and product changes confound. Watch the counter-metric alongside every headline (resolution rate with reopen rate; activation speed with support load) so gaming one number surfaces in another. And read the leading indicators, escalation quality, question themes, weekly, but decide on the outcome numbers quarterly, at cohort scale. None of this is exotic; it's what you'd do for any hire whose impact you cared about.
The procurement version
One question sorts vendors fast: which metric do you propose we hold you to, and will you show it computed on our data during the pilot? Outcome-confident vendors answer instantly, because their product was designed backward from the number. (Disclosure: Skippr's agents are measured on outcomes by design, meetings, time-to-value, resolutions, users trained, and "outcome language beats activity language" is house doctrine we're happy to be audited against.)
Questions buyers actually ask
Are activity metrics useless?
As diagnostics, no, session depth and question themes guide tuning. As success measures, yes: they can't distinguish value from noise.
What if outcomes take months to show?
Use staged evidence: coverage and leading indicators in weeks, cohort outcomes in a quarter. Just never let the leading indicators become the verdict.
How do we attribute outcomes to the agent?
Cohort comparison does most of it: covered versus uncovered segments, same period. Perfect attribution is a research project; decision-grade attribution is a spreadsheet.
See what a live agent actually does
The category is easier to watch than to define. Fifteen minutes is enough to see where the mechanism differs from everything it gets confused with.