"Multimodal" gets marketed as a menu, your AI now does text, voice, and images, pick your channel. That reading misses the point entirely. What multimodal actually means for customer experience is simultaneity: an agent that talks while seeing while acting, the modes fused into one attention, the way a human colleague's are. The menu was never the product; the fusion is.
The fusion reading: modes as one attention
Watch a good human helper at a screen with a customer: they listen while watching the cursor, talk while pointing, notice hesitation while explaining, act while narrating the action. The modes aren't channels; they're one integrated attention doing one job. That's the bar fusion sets for AI: voice, vision, and action in the same live moment, each informing the others, the spoken answer shaped by what the screen shows right now, the action narrated as it happens, the pause in the user's voice cross-read against the stall on their screen. In the fused version, "can you describe what you're seeing?" is absurd, seeing is already part of the listening.
What fusion changes in practice
The interaction stops being turn-based. Diagnosis collapses (the confused description and the visible state are reconciled instantly), teaching becomes demonstration-plus-dialogue (the learner asks while doing, the correction lands at the wrong click), and trust compounds, because narrated action under shared gaze is legible in a way black-box automation never is. Fusion is also what makes long sessions viable: one attention can hold a forty-minute working engagement; a bundle of alternating channels cannot. The practical test for any "multimodal" claim: interrupt the voice mid-sentence to point at something on-screen, and see whether the answer already knows what you're pointing at.
The bar customer experience should hold
Multimodal-as-menu is a checkbox; multimodal-as-fusion is a colleague. The distinction will define the next few years of customer-facing AI, because customers don't experience modes, they experience attention, and attention is either whole or it isn't. (Disclosure: fusion is the design at Skippr, talk, see, and take action as one attention, "Skippr's multimodal AI stack" is the phrase, and the pointing test above is the honest way to check it, on us or anyone.)
See what a live agent actually does
The category is easier to watch than to define. Fifteen minutes is enough to see where the mechanism differs from everything it gets confused with.