An AI agent that talks with your customers, sees their screens, and acts in your product concentrates more trust than almost anything else you'll buy this year. The security evaluation deserves better than a certifications checkbox, here are the questions that actually sort vendors, organized the way a good review runs.
Start with the certifications, then keep going
SOC 2 Type II and ISO 27001 are the table stakes, and the Type II distinction matters: it attests to controls operating over time, not a point-in-time snapshot. But certification is the floor, not the review. The questions that differentiate vendors start after the badge: scope of the certification (does it cover the product you're buying or the corporate website?), recency, and whether the vendor volunteers reports and documentation without a meeting, security-mature vendors have this packaged; the improvisers schedule calls.
The data questions
What does the agent perceive, and what of it is retained? Screen awareness especially: is perception live-only or recorded, where does anything retained live, for how long, and under whose keys? What's the data path for conversations, voice included, and what are the residency options? How is customer data segregated between tenants? Who at the vendor can access session content, under what controls? And the question most vendors haven't been asked enough: what happens to your data at offboarding, deletion timelines, attested. Every answer should exist in writing; verbal assurance is a red flag wearing a smile.
The control questions
Consent and visibility: how do end users know the agent is present, seeing, or acting, and how do they grant or revoke it? Scope enforcement: what bounds the agent's perception to your product surface and its actions to permitted operations, and are those bounds enforced mechanically or by prompt-promise? Admin governance: SSO, role-based access, per-surface and per-environment policy, audit logs of both agent actions and human configuration changes. Guardrails: how are claim boundaries and forbidden actions defined, tested, and monitored? And escalation: what triggers a human, and what evidence travels with the handoff?
The honesty questions
Ask what the agent does when it doesn't know (the answer should be routing, not improvisation), how incidents are disclosed and how fast, what the vendor's own red-team practice looks like, and for references in your compliance neighborhood. Then run the tell that never lies: ask for all of the above in writing, and time the response. (Disclosure: we build Skippr, SOC 2 Type II and ISO 27001 certified, with consent, scoping, and human-in-the-loop as design principles, and this list is the interrogation we'd expect from a serious buyer.)
Questions buyers actually ask
Is SOC 2 Type II enough by itself?
No, it's the entry ticket. The differentiating questions are data retention, scope enforcement, consent design, and governance, none of which a badge answers.
What's the single strongest signal?
Documentation-on-request without a sales call. Vendors who have the answers packaged have usually done the work behind them.
Do these questions apply to self-service pilots too?
Yes, run the pilot on public-safe content while the review proceeds in parallel; a good platform supports exactly that sequencing.
See what a live agent actually does
The category is easier to watch than to define. Fifteen minutes is enough to see where the mechanism differs from everything it gets confused with.