Build an AI SDR tool evaluation checklist before the demo
Generates a two-page, weighted evaluation kit for AI SDR platforms: a scorecard tuned to your situation, demo questions engineered to puncture vendor claims, instant disqualifiers, a 30-day pilot design with pre-agreed metrics, and a true total-cost view. It replaces demo-driven gut feel with an evidence process.
You are a GTM technology advisor who has evaluated over 40 AI SDR platforms for buyers and watched the same movie repeatedly: a slick demo, a 12-month contract, and three months later the tool is generating replies nobody wants and meetings that no-show. You evaluate on evidence, not demo theater. Build me a personalized evaluation checklist for AI SDR tools. Output: 1. SCORECARD — 12-18 weighted criteria grouped under: output quality, data and integrations, control and guardrails, deliverability posture, reporting honesty, and commercial terms. Weight them for MY situation from the interview, and state why each weight is what it is. 2. DEMO QUESTIONS — 10 questions to ask live, each designed to puncture a specific common claim (for example: 'show me the worst email your system sent this week', 'what exactly happens when a prospect replies angry'). 3. RED FLAGS — 6-8 instant disqualifiers (no reply-handling escalation, mandatory annual contracts with no pilot, refuses to show raw sending volumes, claims meetings booked without defining showed meetings). 4. PILOT DESIGN — a 30-day pilot structure with success metrics agreed upfront, sample size, and the comparison baseline (what my current motion produces). 5. TOTAL COST VIEW — the line items beyond license: data credits, inbox infrastructure, ramp time, and internal oversight hours. Rules: no criterion may be scoreable by demo impressions alone — each needs a verifiable artifact or number. The checklist must fit on two pages. Before you build it, interview me. Ask ONE AT A TIME, waiting for my answer each time: 1. What would the AI SDR replace or augment — describe the current motion and team. 2. What volume and what targets (roles, segments) would it work? 3. What is your real budget, all-in, monthly? 4. What is your risk tolerance for AI sending without human review? 5. Which tools must it integrate with? Then produce the checklist, weighted to my answers.
How to use it
- 1
Copy the prompt into Claude, ChatGPT, or any LLM.
- 2
Answer the interview before you book any demos, so the scorecard exists before the sales pitches start.
- 3
Score every vendor on the same sheet within a day of each demo, while artifacts are fresh.
- 4
Insist on the pilot design as written; vendors who refuse pre-agreed success metrics have answered your question.
Best practices
Ask every vendor the same 10 demo questions in the same order — comparability beats cleverness.
Demand raw examples: real emails sent, real reply threads, real escalation cases. Polished case studies are marketing.
Weight deliverability posture heavily if you send from your own domains; a burned domain outlasts any contract.
Score 'reporting honesty' by asking how they count a booked meeting — held, scheduled, or merely accepted.
Example: what this looks like in practice
A VP of Sales at a 90-person fintech is evaluating three AI SDR platforms to augment two human SDRs. The interview captures a $4,000 monthly all-in ceiling, low tolerance for unreviewed sending in a regulated space, and mandatory Salesforce plus Outreach integration. The generated scorecard weights guardrails and integrations at 40% combined, and the demo questions include 'show me an email your system decided not to send, and why'. One vendor fails two red flags immediately (annual-only contracts, wouldn't show raw examples). The pilot with the winner runs 30 days against the human-SDR baseline with 'meetings held with ICP accounts' as the metric — surfacing that the tool books 20% more meetings but with 15% worse show rates, which renegotiates the price before rollout.
Best fit
This prompt is one gear in a bigger machine. We orchestrate 20+ tools into outbound systems our clients own — and guarantee the results.
Apply for a Pilot Spot → →Frequently asked questions
For some motions, yes — high-volume, well-defined segments with strong data foundations see real leverage. But outcomes vary wildly by implementation, and the failure mode is expensive: burned domains, off-brand messaging, and no-show meetings. A weighted evaluation plus a 30-day pilot with pre-agreed metrics, which this prompt builds, is how you find out cheaply which side you're on.
More automation & workflow design prompts
Build a system prompt for an AI cold email reply agent
Interviews you about your offer, booking flow, and voice, then generates a complete production system prompt for an AI reply agent — classification taxonomy, response rules, escalation conditions, and a strict JSON output contract an n8n or Zapier workflow can parse. You get the artifact teams pay thousands to have written, tuned to how you actually talk.
Turn a manual sales process into an automation blueprint
Converts a fuzzy 'we do this by hand every week' description into a structured blueprint with a detectable trigger, labeled steps, edge cases, human checkpoints, and failure alerts. It forces the design decisions people usually discover mid-build, so what you hand to n8n, Zapier, or an ops hire is assembly work, not archaeology.
Architect a Clay table before you burn the credits
Designs the Clay table before you build it: sources, column order, run conditions, waterfalls, outputs, and a credit estimate. Column ordering and gating are where most teams silently overspend, so the prompt enforces filter-first architecture and makes every paid column defend its run condition — typically cutting credit burn 40-70% versus enrich-everything builds.