Customer-facing AI has moved through four distinct stages in a decade, and knowing which stage a tool belongs to explains almost everything about what it can and can't do for you. The short version: chat answered questions; voice made it conversational; eyes and hands made it capable; live made it an employee.
Stage one: chat (2016-2023)
The first generation was the chat widget: rule-based bots that gave way to LLM-powered ones until, by 2023, a chatbot could answer a surprising share of documented questions. Chat deserves its credit, it proved customers would engage with automated help, and it deflected real ticket volume. But it carried two structural constraints no model upgrade could lift: it was blind (the user had to describe their screen in words, and the bot had to guess) and it was stateless (every exchange a quick hit, context dying with the window). Chat was the right first step, and it hit a ceiling. It answers questions; it doesn't finish jobs.
Stage two: voice (2023-2025)
Real-time speech models collapsed the latency that had made earlier voice assistants unbearable, and suddenly AI could hold a natural spoken conversation, interruptions, tone, mid-sentence language switching. Voice mattered twice over: bandwidth, because people explain problems much faster by talking, and presence, because a voice on a call is a participant, not a form field. This era produced AI phone support and avatar-fronted experiences of every kind. Most voice agents, though, inherited chat's blindness: they could hear the user and still couldn't see the screen or touch the product. A voice that cannot see is a phone call from someone who has never used your app.
Stage three: eyes and hands (2024-2026)
The third wave gave agents perception and action. Vision models learned to read live interfaces, not just parse structure, but understand which page the user is on, what state it's in, where they're stuck. In parallel, computer-use and browser-automation capabilities let agents click, navigate, fill forms, and complete flows. The proof points piled up fast across the industry, and every incumbent in every software category began bolting agents onto their existing form factor. The direction stopped being in dispute; the wave was general, and no one company invented it.
Stage four: live, where the synthesis lands
The fourth stage combines all of it in one continuous session: voice to talk, eyes on the user's actual screen, hands that act in the product, a brain that knows the business, plus the two things bolt-on agents lack. An agenda: live agents run long, structured sessions, a demo with a storyline, an onboarding plan spread over days, and track progress instead of reacting turn by turn. And membership in a team: the agent that ran the demo hands the account to the one that onboards it, memory intact. That synthesis is where the category stops being a better widget and becomes a colleague, which is why the stage is named for how it feels to work with: live. (Disclosure: stage four is where we build, Skippr's agents were born there rather than retrofitted toward it, but the staircase itself belongs to the whole industry.)
See what a live agent actually does
The category is easier to watch than to define. Fifteen minutes is enough to see where the mechanism differs from everything it gets confused with.