1. Start with the job, not the vendor
Write one concrete job in plain language. Include the input, the desired outcome, the systems involved and the conditions that should trigger a human handoff. “We need AI for sales” is too broad. “Respond to inbound website leads, qualify fit, follow up twice and book a call without quoting prices” is testable.
2. Define pass/fail criteria before the demo
Decide what success means before you evaluate a product. Otherwise every vendor demo can look impressive. Use measurable criteria such as task completion, correct escalation, factual accuracy, number of human corrections and accepted-result cost.
| Weak criterion | Better criterion |
|---|---|
| “Sounds natural” | Answers correctly without inventing policy |
| “Can book meetings” | Books only valid slots and handles reschedules |
| “Integrates with CRM” | Reads/writes only fields required for this job |
3. Check permissions and failure consequences
An agent that can send email, edit CRM records or create appointments can also do the wrong thing at scale. List every permission the job requires and separate reversible actions from actions that can create financial, legal or customer impact.
4. Measure the cost of an accepted result
Subscription price alone is misleading. Include usage fees, setup, human review, corrections and any additional tools required. A cheaper agent that creates twice as much review work may be the more expensive hire.
5. Give candidates the same pilot
Compare candidates on the same representative scenarios and the same rules. Include normal work, edge cases and at least a few situations where the correct behavior is to stop and escalate.
Would you put AI candidates through a working interview?
HiredBot is in private beta. Tell us the job you would want tested and which AI tools you are considering.
Join Early Access