AI SDR: What It Actually Automates and Where It Breaks
The six jobs an AI SDR genuinely does, the deliverability ceiling nobody sells against, how to evaluate a vendor, and the honest cost comparison.
in this article
The demo is always the same. You paste in a domain, the product thinks for forty seconds, and out comes a list of two hundred contacts with verified addresses, a summary of what each company does, and a first email referencing the funding round they announced in March. It is work that used to fill a junior rep's week.
Then you buy it, and six weeks later you are in a meeting about domain reputation.
An AI SDR is a real category with a real capability and a marketing story running two years ahead of the software. Here is the separation, for whoever has to operate it.
The six jobs it genuinely does
List building from criteria. You describe an ICP, the product turns it into filters and queries a data provider. Apollo, Clay and ZoomInfo sit underneath most of these tools; the AI part is mostly translating a sentence into a filter set, over the same data your competitors buy.
Waterfall enrichment. A genuine engineering win. Rather than accept one provider's guess, the product queries several in order, stops at the first verified hit, and pays the expensive source only when the cheap ones miss. On a European mid-market list that is the difference between enriching half your records and enriching most of them.
Research summarisation. Given a domain, pull the site, the careers page and recent announcements into a short brief: what they sell in their own words, who they sell to, what changed. Tedious, rule-heavy, no judgement required.
First-draft sequencing. Four or five touches across email and LinkedIn, timed, with the opening grounded in whatever the research surfaced. Drafts. The most valuable output in the product, and the most mis-sold.
Reply classification. Sorting the inbox into interested, not now, wrong person, unsubscribe, out of office and hostile. Better than the keyword rules it replaced, and it removes half an hour of triage a day.
Meeting booking. Turning "Tuesday afternoon could work" into a calendar hold against real availability.
Those are the jobs that appear whenever agents are scoped honestly: recurring, high-volume, low-judgement work upstream of a person.
What it does not automate, whatever the site says
It does not decide who is worth contacting. The ICP is your judgement encoded; the tool executes filters and has no opinion about whether they describe a buyer. Feed it a bad ICP and you get a bad list faster than you ever have.
It does not know when a signal is stale, and nothing flags a fourteen-month-old funding round except a freshness rule someone writes. It does not judge whether a line reads as insight or as surveillance. And it does not handle the second reply, where a question about pricing tiers turns a confident wrong answer into a lost deal.
It does not fix the offer either. Applied to a proposition nobody wants, it produces a faster demonstration that nobody wants it.
Deliverability is the ceiling, and no vendor sells against it
The cost of drafting an email went to zero. The number of emails you can deliver did not move. Drafting time used to throttle volume; the tool removes the throttle, volume triples on infrastructure warmed for a third of it, and the failure arrives as a reputation problem rather than a copy problem.
The mechanics matter. Reputation accrues per sending domain and per mailbox, which is why outbound runs on secondary domains, never the one your invoices come from. Operators who stay out of trouble hold a warmed mailbox to twenty or forty sends a day, not hundreds. Warming means a ramp over several weeks with real replies; warmup networks where mailboxes email each other in a loop are increasingly detected and simulate reputation rather than build it. Since the major providers tightened bulk sender requirements in 2024, SPF, DKIM, DMARC and one-click unsubscribe are table stakes, with a complaint threshold around 0.3%.
Spam traps are the quiet killer. Recycled traps are dead addresses providers reactivate to catch people mailing stale lists; pristine traps are seeded on pages to catch scraping. Both punish importing a list someone exported eighteen months ago and blasting it because the tool made it cheap.
Run this check in ten minutes. From each sending inbox, send a plain message to a fresh Gmail account, a fresh Outlook account and a corporate mailbox behind a filtering gateway, and note where each lands. Then look up the DMARC record for every sending domain and confirm the policy is not still p=none. If anything is in Junk or Promotions, better copy is not the fix. The Deliverability Preflight is the longer version.
A burned domain is not repaired with a setting. You buy new ones, warm them for weeks and absorb the gap. That is why "more sends" is the characteristic failure mode: the one lever the tool makes effortless and the one infrastructure cannot absorb.
Reply handling and the handoff you have to design
Classification is safe. Generation inside a live thread is not. The rule that holds: the agent classifies and drafts, and a human sends anything containing a commitment, a number, a date, or a claim about what the product does.
The handoff needs instrumentation most tools do not ship. When a reply arrives, one named person owns the thread, every sequence for the account stops everywhere within minutes rather than at the next nightly sync, and the account drops out of paid audiences. That is ordinary suppression logic, and its absence is how two motions collide on one prospect, or how an out-of-office reply gets classified as engagement and restarts a cadence.
Four questions that stop the demo
Who owns the sending domains? If the vendor does, your reputation is pooled with their other customers and you cannot take it with you.
Show me a real deliverability view. Bounce, complaint and reply rates by domain and mailbox over ninety days, on a live account not the demo tenant.
Which messages does it send unattended? A precise answer, not "there is a human in the loop".
What is its refusal behaviour? Good products say "no usable signal, research or drop" and check claims against a proof file. Products that always produce a draft are producing invented relevance.
The cost comparison, done honestly
The slide says fifteen hundred a month against sixty-five thousand fully loaded for a junior SDR. Three things are missing from the left column. Data and enrichment credits are metered separately and frequently exceed the platform fee. Sending infrastructure is a line item: ten domains and thirty mailboxes, each with a monthly cost, plus a warming period you cannot send during. And someone operates it, reading the queue, maintaining the ICP and watching the complaint rate.
The honest framing is not tool versus person. It is tool plus one competent operator against two or three people doing the work by hand, and on that comparison the tool often wins, just not by the factor on the slide. The comparison also cannot price that a junior SDR becomes an account executive in two years and software does not.
Measure it the way you would measure a hire. Not emails sent, not meetings booked, but qualified meetings held that a rep would have accepted, per quarter, against total cost. And complaint rate, weekly. If the second number is rising, the first is borrowed against your future.
Frequently asked questions
What does an AI SDR actually automate?
Six things reliably: building lists from ICP criteria, waterfall enrichment across several data providers, summarising research about an account, drafting a first sequence grounded in that research, classifying inbound replies, and booking meetings against real availability. It does not decide which accounts are worth contacting, know when a signal has gone stale, handle a substantive second reply, or fix a weak offer.
Can an AI SDR replace a human SDR?
Not like for like. It replaces the execution layer: list building, enrichment, research and first drafts. What remains is judgement about which accounts matter, live conversation once a prospect replies, and operating the system. The realistic comparison is one operator plus the software against two or three people working manually.
Why do AI SDR tools get domains blocked?
Because they remove drafting time as the natural limit on volume without changing anything about deliverability. Sending rises on mailboxes warmed for far less, complaint rates cross provider thresholds, and stale lists hit spam traps. Reputation is per domain and per mailbox, takes weeks to build and months to recover. The defences are secondary domains, correct authentication, a real warming ramp, modest daily volume and list hygiene.
How should I evaluate an AI SDR vendor?
Ask who owns the sending domains, and insist on bounce, complaint and reply rates by domain and mailbox over ninety days from a live account. Establish which messages it sends without a human, what it writes back to the CRM, and whether it refuses to draft when there is no usable signal.
where this lives in the system
shorter reads on this, at aiporate.com
see where you stand
Twelve questions. Then your build order.
The diagnostic returns your operating stage, the three widest gaps in your motion and what to build first. Two minutes, no sales sequence, one human reply.