AI wrote our 200 cold outreach emails — here's what happened
We handed our launch outreach to AI. 200 emails, real signal stack, real numbers. The lesson isn't what you'd expect.
AI wrote our 200 cold outreach emails — here's what happened
We handed our launch outreach to AI. 200 emails, real signal stack, real numbers. The lesson isn't what you'd expect.
We had a launch window, no sales team, and about four hours to figure out our cold outreach strategy. So we did what any AI company should probably do: we handed the whole thing to AI and watched what happened.
Two hundred emails. Founders and ops leads at AI and B2B SaaS companies with 2–50 employees. No contact database. No agency. No hand-crafted copy. The AI built the list, sourced the signals, and wrote every email. We reviewed for accuracy and hit send.
Here are the results — and why the lesson from them is not what most people expect.
The bet we were making (and why it was a real bet)
Standard B2B cold email benchmarks in 2026 are commonly cited at around 21% open rate and 2–3% reply rate — though platform averages across large volumes run higher. Those numbers assume a competent human writing targeted copy. We were replacing that human with AI and asking whether the output would be closer to "founder outreach" or "enterprise spam."
The honest answer going in: we didn't know. There is a version of "AI wrote my cold emails" that means someone ran a first-name merge through GPT and called it personalization. That performs like spam because it is spam. We were betting on something different.
Our target segment — founders and ops leads at early AI companies — are precisely the people most likely to spot AI-generated boilerplate and delete it on sight. If you are selling an AI product to people who build AI products, you do not get to be lazy about this.
The signal stack: what AI actually used to personalize each email
This is the part that determines everything else. Here is what we fed the AI for each contact:
- Company LinkedIn page (not just the name — the "About" section, recent posts, headcount signals)
- Last 3 blog posts published by the company or its founders
- Active job listings (what they are hiring for tells you where they are hurting)
- Tech stack via BuiltWith (are they on Vercel? Supabase? AWS Lambda? Tells you a lot)
- Crunchbase funding date (how long since last raise, which tells you where they are in the cycle)
The AI's job was not to write a clever opening line. Its job was to surface one genuine, specific observation about each company and connect it to something we could actually help with. "You're hiring a Head of Sales" is a signal. "You've published four posts about your onboarding funnel in the last 60 days" is a signal. A funding date from 14 months ago is a signal about runway math that every founder in that position is thinking about.
The emails that used this signal stack averaged 34% open rate and 8.4% positive reply rate. For context, that is the range most founders achieve when they write personal outreach themselves — but these were written by AI, for 200 contacts, in about three hours of total human time.
The segment that performed best: companies that had raised seed funding 12–18 months prior and were actively posting about GTM, sales, or pipeline. The signal set was richest for them, and the AI had the most to work with.
The numbers — unfiltered
Metric Our campaign Standard B2B benchmark Open rate 34% ~21% Reply rate (any) 11.2% 2–3% Positive reply rate 8.4% ~1% Meeting booked rate 4.1% 0.5–1%We booked 8 meetings from 200 cold emails. Four of those turned into active conversations. One closed.
Those are founder-outreach-quality numbers from a process that ran mostly without us. The human time involved was: defining the target segment, reviewing the AI's emails for accuracy before sending (non-optional — more on this below), and responding to replies.
Where it went wrong: the hallucination problem
Eleven emails in the batch were flagged as problematic during review — and two of them got sent before we caught the pattern.
The failure mode: the AI confidently referenced a blog post, a podcast appearance, or a product launch that did not exist. "I saw your post about your enterprise migration" when no such post existed. "Congrats on the Product Hunt launch" when there was no launch.
One reply came back: "Not sure what post you're referring to — you might have me mixed up with someone else." Polite. But that was a relationship that started with a factual error.
How it happened: for some contacts, the signal sourcing came up thin — no recent blog posts, no job listings, no meaningful LinkedIn activity. Instead of producing a lower-personalization email, the AI hallucinated personalization context to fill the gap. It was pattern-matching to the format of the high-signal emails and fabricating the content.
The fix is simple but you have to actually do it: every email needs a human accuracy pass before it sends. Not a style pass — an accuracy pass. The question is not "does this sound good?" but "is every specific claim in this email verifiable?" If the AI says "I saw your post about X," does that post exist? If it says "you raised in February," does Crunchbase confirm that?
This step is not optional and it is not a sign that the AI failed. It is the system working correctly. The AI's job is to draft at scale; the human's job is to be the hallucination filter.
Signal quality beats copy quality, every time
The single most important thing we learned from this campaign is not about prompting strategy or email structure. It is about this: an AI writing about something that actually happened — a real Series A, a real product launch, a real job posting — outperforms a beautifully crafted template with no signal by 3–4x on reply rate.
The emails that underperformed were the ones where we ran out of signal. Not the ones where the copy was worse. The copy was fine across the batch. What made the difference was whether the AI had something real to connect to.
This changes how you should think about AI outreach. Most people try to improve the writing — better subject lines, better openers, better CTAs. That is the wrong lever. The lever is signal sourcing. How much real, specific, current information can you feed the AI about each target? That is what it runs on.
The traditional advice in founder-led GTM is "do things that don't scale." Send personal emails. Do manual research. Talk to customers. That advice was never really about writing — it was about the quality of attention and specificity. AI can now deliver that specificity at scale, as long as it has the signals to work from. The bottleneck has shifted. It is no longer writing; it is sourcing.
The workflow that actually worked
List-build → signal-source → first-draft → human accuracy pass → send.
Do not collapse steps. The list-build stage determines who gets an email. The signal-sourcing stage determines whether that email is worth opening. The first-draft stage is fast. The accuracy pass is slow-but-required. Skipping any step downstream of the signal-sourcing is low-risk; skipping the accuracy pass is not.
For teams running this as a repeatable GTM outreach automation — not just for a launch, but as a standing weekly workflow — the bottleneck becomes the signal-sourcing step. It requires fresh data: recent blog posts, recent job postings, funding events in the last 30–90 days. Signals go stale. A company that was hiring aggressively three months ago may have frozen headcount. An email built on that stale signal sends the wrong message.
This is exactly the loop SideKyk's Sales & Lead Gen specialist was built to run — sourcing fresh signals weekly, building the first-draft queue, flagging the accuracy review items. SideKyk is a team of AI agents that lives in your WhatsApp. No new app, no onboarding call, no subscription to cancel. If you want this running before your next outreach push, sign up at sidekyk.ai/ai-business — we walk new accounts through the signal stack setup in the first session.
The actual lesson
We went into this campaign thinking we were testing whether AI could write good cold emails. We came out knowing that was the wrong test.
The question is not "can AI write good cold email?" It can. The question is "can your pipeline reliably source the signals that make AI personalization real rather than fabricated?" If yes, you get founder-outreach quality at 10x the throughput and a fraction of the time. If no, you get sophisticated-looking spam with a hallucination problem.
The difference between those two outcomes is not the model. It is the data you give it.
Want this running in your WhatsApp every Monday morning?
Drop your number — we'll WhatsApp you the moment AI product teams goes live.
Join the waitlist on WhatsAppPowered by SideKyk · A team of AI agents in your WhatsApp