To measure training, climb the metric ladder: activity (enrollments, completions), learning (quiz scores), performance (the workflow executed live, unassisted), and business impact (tickets, adoption, ramp, renewals). Most programs report the bottom two rungs because that's all their tooling could see. The unit that changes decisions is users trained: people observed performing the skill.
Why the ladder collapses at rung two
Completion measures exposure; quizzes measure recognition. Both are honest about what they are, and both are routinely presented as if they were rung three. The gap is cognitive, not cynical: recognizing the right answer among four options is a different act from producing a workflow under pressure in a live interface. Every operator has met the certified-but-frozen user. Programs kept reporting rungs one and two anyway, because witnessing rung three required a human assessor beside every learner, which scaled exactly as badly as it sounds.
What "users trained" means operationally
A user is trained on a workflow when they've performed it, in the product or its practice mode, on realistic data, unassisted, witnessed. That definition is measurable now because an in-product training agent can do the witnessing: it teaches on the learner's screen, watches the attempt, and records the outcome, performed cleanly, performed with prompts, or not yet, with the gaps named. The aggregate becomes an audited skill inventory: which workflows, which teams, which accounts, verified rather than attested. (Disclosure: this is the metric Skippr's Skippr AI training is built around, "measured on users trained" is the design, not the tagline.)
Connecting rung three to rung four
Performance data finally makes the business rung arguable. Compare trained against untrained cohorts on the metrics the training was supposed to move: support tickets on the trained workflows, adoption of the trained features, time-to-productivity for new hires, renewal outcomes for educated accounts. Two disciplines keep you honest. Baseline first: measure the cohort before the training changes anything. And mind selection effects: users who seek training differ from users who don't, so compare like cohorts, or, stronger still, roll training out to comparable groups in waves and read the difference.
What to keep from the old dashboard
Don't delete completion tracking; repurpose it. Compliance formats still require attendance records, and completions remain a useful funnel metric, you can't perform what you never started. The change is narrative: completion becomes an input measure, users trained becomes the output measure, and business impact becomes the argument. When leadership asks how training is going, the answer stops being a percentage of videos finished and becomes a list of teams that can demonstrably do the thing.
Questions buyers actually ask
What's wrong with completion rates?
Nothing, as long as they're presented as exposure, not competence. The trouble starts when finishing content is reported as ability to perform, which no completion number can support.
How do you measure training without an AI agent?
Sampled practical assessments: a human observes a subset performing real workflows. It's the same rung-three logic at manual scale; agents remove the sampling constraint.
Which business metrics move first?
Usually support volume on the trained workflows and feature adoption, within a quarter. Ramp time moves with the next hiring class; renewal effects read on the annual cycle.
Watch it train someone
Completion is not competence, and the difference shows up in the product rather than the dashboard. Fifteen minutes is enough to tell them apart.