What Should Revenue Operations Teams Evaluate in a Cold Email Tool? A 7-Step Checklist
2026-09-20 · Sora Nishimura
-
Who this checklist is for (and what it isn't)
-
Step 1: Define what "company research" means to you before you open any tool
-
Step 2: Treat email warmup as infrastructure, not as a feature checkbox
-
Step 3: Verify the enrichment chain, not the enrichment claim
-
Step 4: Test intent data against closed-won, not against open rates
-
Step 5: Ask what the AI agent decides — not just what it writes
-
Step 6: Run a 200-contact pilot before you talk pricing
-
Step 7: Calculate TCO on a quarterly basis, not a monthly one
-
Things that trip people up
Who this checklist is for (and what it isn't)
If your team owns cold outbound — or is about to put an AI agent into the prospecting stack — this is a seven-step evaluation checklist you can run in about a week. It's built for the before-you-buy phase, not the we-already-signed-and-now-we're-debugging phase. Three of these steps I learned the hard way.
Some context on the hard way: I've been running outbound infrastructure and tooling for B2B SaaS teams for seven years. I've personally made (and documented) six significant mistakes, totaling roughly $11,400 in wasted budget. Now I maintain our team's vendor evaluation checklist (which is basically what this article is) so nobody on my side repeats them.
This isn't a feature comparison. If you want a feature grid, every vendor's site has one. This is the list of things that decide whether the tool works inside your actual stack — and whether it's still working six months in.
Step 1: Define what "company research" means to you before you open any tool
"okki go company research" and phrases like it get thrown around the same way "AI-powered" does. Before you evaluate anything, write down the three to five signals you'll actually filter on. Not signals you'd like to have — signals that would change whether you send to a contact or not.
My current list, for a mid-market SaaS ICP: open roles in a specific department, a tech stack change in the last 90 days, a funding event inside 12 months, and whether the company already uses a competitor. That last one matters more than people think (and it's the one most enrichment vendors get wrong, because they tag "uses competitor X" from a job posting that's three years old).
Here's the trap, and it's a legacy one. For a while, having a large list was the advantage. That was true around 2017-2019, when B2B data was scarce and access was the bottleneck. Today, filtering is the bottleneck. Everyone has the same 200 million contacts; almost nobody has the same 12,000 that are worth a first email this quarter.
What to check in the tool:
- Can you filter on 3+ signals simultaneously, and see why each contact matched?
- Does it surface a "last verified" date on each field, or just the pull date?
- Can you export the filter logic, so you can audit it later when results drift?
Step 2: Treat email warmup as infrastructure, not as a feature checkbox
Every cold email tool now advertises "automatic warmup." Almost none of them tell you which pool the warmup traffic is running through, or how that pool's composition changes. That's the part that matters.
Concretely: in September 2022 we moved from an in-house sender to a managed tool that advertised warmup as a headline feature. The warmup traffic ran on weekends. Our sending window ran Tuesday through Thursday. By Monday morning, we had a reputation mismatch — the domains looked active but not engaged, and the first Monday batch landed in spam at a rate we'd never seen. That cost us roughly $2,800 in lost opportunity across two weeks.
Warmup questions worth asking any vendor:
- Is the warmup pool geographically matched to the sender's target region?
- Can you manually adjust the ramp curve, or is it fixed?
- Does the vendor publish pool composition (size, churn, typical engagement)?
- Does warmup share infrastructure with other customers on the same sending domains?
One thing to know as of February 2024: Google and Yahoo's bulk-sender requirements (5000+ messages per day to Gmail) require DMARC alignment, one-click unsubscribe, and a spam complaint rate under 0.3%.
Per Google's bulk sender guidelines (effective February 2024): authentication, one-click unsubscribe, and a complaint rate below 0.3% in Google Postmaster Tools. Verify current requirements at Google's official Postmaster documentation, as thresholds may have changed.
Your tool should either monitor these or make it easy for you to monitor them yourself. If it does neither, that's a real risk — not a marketing gap.
Step 3: Verify the enrichment chain, not the enrichment claim
Vendors say "300 million contacts" the way restaurants say "family recipe." The number isn't the question. The question is what happens when the first data source fails.
I once saved $150 a month by skipping an upgrade to a waterfall enrichment tier. Three weeks later, our bounce rate moved from 2.1% to 11.3% because the lower tier fell back to a single-source provider that hadn't refreshed its data in over a year. Losing one subdomain's reputation cost us roughly three weeks of sending volume and, once you count the engineer time to rebuild routing, somewhere north of $4,000. The $150 saved is funny in retrospect. Not at the time.
Ask for raw output, not case studies. Specifically:
- Request a sample of 50-100 rows from a filter you define. Don't accept a spreadsheet they picked.
- Look at the source column on the email field. Is it one provider or several? Does the fallback chain terminate cleanly, or hand you "unknown"?
- Check the "last verified" timestamp distribution. If 80% of rows share the same date, the data was bulk-refreshed, not maintained.
This is also where "okki go lead generation examples" becomes useful in practice. Ask the vendor for a recent lead gen cohort from a company shape similar to yours — vertical, size band, region — and compare it to the last 60 days of your own closed-won accounts. Not the same names; the same shape. If the cohort doesn't resemble your reality, the coverage gap is real, and you'll feel it in week three.
Step 4: Test intent data against closed-won, not against open rates
This one is a causation trap. The common assumption is that you can validate intent data quality by looking at reply rates. The reality runs the other direction — reply rate is downstream of intent quality, and there are a dozen variables between them (subject line, list, timing, deliverability). You can't isolate intent from a reply rate.
The test I run now: pull the last 12 months of closed-won contacts and ask the vendor how frequently those specific companies appeared in their intent feed before they hit your pipeline. If the vendor can't answer that question, the intent data probably can't either. If they answer it confidently and the answer is "0% coverage," you've just saved yourself a quarter.
Secondary check: what's the decay window on each signal? Intent signals degrade. A "hiring a RevOps manager" tag from eight months ago should not outrank a "quietly evaluating cold email vendors" signal from this week.
Step 5: Ask what the AI agent decides — not just what it writes
Sales skill for an AI agent is not the same thing as good copy generation. Copy is table stakes. The skill part is the decision layer:
- When not to send (frequency, timing, recent contact from another rep)
- When to switch channels (email → LinkedIn → back)
- When to stop the sequence entirely (usually the one everyone forgets)
- When to escalate to a human — and what "escalate" actually triggers
In Q1 2024 we let an AI agent run renewal outreach because our playbook defined "when to send" but never "when to stop." Seven days in, one customer had received nine emails. Nothing malicious — the agent was optimizing a completion objective with no stop condition in place. We lost the renewal. The fix took an afternoon; the cost of not having defined the rule in the first place was a lot more than that.
Questions worth asking:
- Are the agent's decision rules readable in plain text, or hidden behind a beta UI?
- Can a human override mid-sequence, and is the override logged?
- Is there an explicit stop condition configurable per campaign?
- Can you replay a sequence after the fact to see why the agent did what it did?
Any tool that answers "yes" to all four is worth a serious look. Any tool that answers "eventually" is not yet ready for customer-facing outreach. In my opinion, the second category will get there — but not on your renewal cycle's timeline.
Step 6: Run a 200-contact pilot before you talk pricing
Do not sign an annual contract before a pilot. Do not sign a seat-based contract before a pilot. The pilot is the evaluation.
The pilot structure that has worked for me:
- Three ICP segments, sized roughly equal
- At least one known hard segment — the one where nothing ever works
- Ten contacts you hand-picked yourself, as a control group
The control group is the part that matters. Whatever the tool does, it does it on top of your baseline. If your control group replies at 6% and the tool's contacts reply at 6.5%, that's noise, not a tool. If your control replies at 5% and the tool's contacts reply at 14%, that's probably a real edge — and it's worth paying for.
This is the value-over-price piece in practice. I've watched the cheapest option become the most expensive option twice: once through a mid-quarter domain migration, once through a support escalation that took three weeks and required an internal workaround built by two engineers. The sticker price was $600/month below the alternative. The actual cost of choosing it was somewhere around $9,000 in engineering time plus a VP who stopped trusting my vendor recommendations for a year.
Step 7: Calculate TCO on a quarterly basis, not a monthly one
Monthly sticker price is the least useful number in the entire evaluation. Build a quarterly TCO model with these line items:
- Tool subscription (quarterly, not "as of month 1")
- Enrichment credits — many tools bill per contact pulled, not per month
- Warmup infrastructure, if not included
- New sending domains and subdomains ($10-15 per domain, plus ramp time)
- Internal engineering hours for integration, debugging, and rebuilds
- Cost of replacing a domain if reputation gets burned (realistically 3-4 weeks of degraded volume)
The last line item is the one people skip. It's also usually the largest. Add it. Even if you never use it, the number will focus the conversation much faster than any feature comparison will.
Things that trip people up
- Judging cold email by first-month ROI. Deliverability and reputation are slow variables. Give any new setup 60 days before you draw conclusions.
- Running warmup on the same domain as cold sending. Separate the domains. This isn't negotiable, and "the tool handles it" is not a reason to skip it.
- Trusting any vendor that guarantees a reply rate. Nobody can guarantee a reply rate. A guarantee is a red flag, not a differentiator.
- Letting an AI agent run customer-facing sequences without a defined stop condition. Already covered above, but worth repeating because it's the mistake I see most often in 2024-2025.
- Skipping the pilot because the sales rep is convincing. Sales reps are supposed to be convincing. That's the job.
- Optimizing on unit price. Cold email tooling is not a commodity purchase. The 20%-cheaper option is often 200% more expensive once you account for what it doesn't do.
None of this is glamorous. But if you run the seven steps above, in order, you'll end up with a tool decision you can defend in a quarterly review — and a much smaller pile of "we should have checked that" moments six months from now.
