Most prospect lists fail before a single message is sent, because they were assembled in the wrong order. Someone collects a few thousand addresses, then works out who those people are, then discovers most of them were never the right audience.

The order that works is the reverse: decide who you want, find the companies, find the people inside them, and only then chase addresses. It is slower at the start and produces a list perhaps a tenth the size — which then performs several times better and does not damage your domain. This guide walks through it end to end.

Define the target before you collect anything

Write down, in a sentence you could say aloud, who you are trying to reach. Not "SaaS companies" — something you could test a specific company against and get a clear yes or no.

Concretely, decide four things:

  • Firmographics. Industry, headcount range, geography, and any structural signal that matters (multi-site, regulated, publicly listed).
  • The trigger. What makes a company relevant now? Recent funding, a job posting for a role your product supports, a migration announcement, a new regulation in their sector. Outreach tied to a trigger outperforms untied outreach by a wide margin, and it is also what makes the legitimate-interests argument defensible.
  • The role. Which job title actually has the problem you solve, and which one signs. They are frequently different people, and writing to the signer first is a common way to get ignored.
  • The disqualifiers. Who looks like a fit but is not — too small to afford it, already using a competitor under contract, in a jurisdiction you cannot serve.

This takes an hour and saves weeks. Everything downstream is filtering against these four, and without them you have no basis to reject anything, which is how lists become large and useless.

Step 1: build the account list first

Collect companies, not people. A hundred well-chosen accounts is a better starting point than five thousand scraped contacts, because you can evaluate a company against your criteria and you cannot evaluate a bare address.

Sources that reward the effort, roughly in order of signal quality:

  • Your own closed-won accounts. Look for what they share and go find more of it. This is the highest-quality targeting input you will ever have and it is free.
  • Industry associations and trade bodies. Member directories are public, curated, and pre-filtered to your sector.
  • Conference exhibitor and attendee lists. A company that paid for a stand at a sector event has declared both its market and its budget.
  • Trigger feeds. Funding announcements, job boards, procurement notices, planning registers. These give you the company and the timing together.
  • Review sites and marketplace listings. Companies listed as users of an adjacent product are pre-qualified for a category they already buy.
  • Company registries. Formal but useful for size, sector code and jurisdiction filtering.

Record the company, its domain, why it qualified, and where you found it. That last field matters more than it looks: it is what you need for the disclosure requirement in your first message, and what lets you tell later which sources produced customers.

Step 2: find the right person at each account

Now go one level down. For each company, identify a named human with the job you decided on.

Public professional profiles are the obvious source, and the company's own site is underrated — team pages, press releases, case studies, conference speaker bios and support documentation all name people, with their exact job titles. Podcast and webinar guest lists are excellent for senior roles. For regulated industries, filings and annual reports name officers directly.

Two practical notes. Titles are not standardised: the person who owns your problem might be Head of Operations at one company and VP Business Systems at the next, so match on responsibility rather than on a title string. And check recency — a team page from 2023 will send you to someone who left, which is both a bounce and an embarrassing first impression if it does not bounce.

If you genuinely cannot identify a person at a company that otherwise qualifies, park it in a separate list rather than defaulting to info@. Role accounts convert poorly and complain often, and the account is worth more to you later when you can name someone.

Step 3: get the addresses

Only now does address collection start, and by this point you know exactly whose address you want, which changes the whole exercise.

Four routes, in descending order of reliability:

  1. Published directly. Team pages, press contacts, academic and professional directories, conference programmes, published papers. If it is there, take it — this is the most defensible source you can have.
  2. In documents you already hold. Signature blocks in existing threads, attendee PDFs, exports from a previous system. Copy the whole document, extract in one pass, and let deduplication sort it out. Long forwarded threads in particular carry far more contacts than the visible recipient list.
  3. Derived from the company's pattern. Covered in the next section.
  4. Commercial contact databases. Fast and expensive, with coverage that varies sharply by region and company size. Treat their data as a hypothesis to verify, not as fact — a meaningful share of any such database is stale.

Whichever route, pull the addresses out of the raw source rather than retyping them. Transcription is where typos enter a list, and a typo becomes a hard bounce that costs you reputation. Paste the source into an extractor, let it strip the surrounding markup, deduplicate, and export with the domain split out. Retyping forty addresses by hand will produce at least one mistake and you will not know which.

Working out a company's address pattern

Most organisations use one consistent scheme for everyone. Establish it from addresses you already have, and you can construct the rest.

The common patterns, in rough order of frequency:

PatternExample for Dana Wright
first.last@dana.wright@acme.io
firstlast@danawright@acme.io
first@dana@acme.io
flast@dwright@acme.io
first_last@dana_wright@acme.io
lastf@wrightd@acme.io

To find the pattern, gather every address you have at that domain — from a press page, a support thread, a PDF — and look at the shape. Two examples agreeing is usually enough. This is exactly the case where exporting with the domain as its own column pays off: sort by domain and each company's addresses sit together, making the pattern obvious at a glance.

Then verify what you construct. A derived address is a guess, and guesses bounce. Never send to a constructed address unverified, and never construct more than one variant per person and send to all of them — mailing dana@, d.wright@ and dana.wright@ simultaneously is a recognisable spam pattern and is treated as one.

Watch for the exceptions: companies that have rebranded often accept mail at both domains, large organisations use regional subdomains, and two people with the same name force one of them onto a variant.

Step 4: structure the list so it stays useful

A list that is only addresses cannot be segmented, cannot be audited and cannot be improved. The columns worth keeping from the start:

  • Email, domain, TLD as separate fields. Domain lets you group by company, spot one domain contributing two hundred addresses, and detect a scrape that went sideways.
  • First name, last name, job title, company — anything you personalise on has to exist as its own field.
  • Source and collection date. Both are required for the disclosure in your first message, both are what a regulator asks about, and both tell you which sources are actually working.
  • Qualification reason. One line on why this company is a fit. Future you, reviewing the list in three months, will not remember.
  • Verification status and date. So you know what is stale.

Keep it in a spreadsheet until it is genuinely working, then move it into your CRM. Building elaborate tooling around an unproven list is a way of avoiding the harder question of whether the targeting is right.

Step 5: clean, verify, segment

Before anything is sent, run the hygiene pass in the order that costs least: normalise and deduplicate, remove role accounts, apply your domain rules, subtract your suppression list, and only then pay to verify what remains. Doing verification first can easily triple the bill for the same outcome. The full reasoning is in how to clean an email list before your first campaign, and how to read the results in what email verification results actually mean.

Then segment before writing. At minimum split by the trigger that qualified them and by role, because those two determine what the message should say. A list segmented into four groups with four messages will outperform the same list with one message by a margin that makes the extra hour trivially worth it.

Check the jurisdiction split too. The rules differ materially between the EU, the US and Canada, and a domain TLD is a weak proxy for where someone actually is. What each regime requires is covered in using extracted email addresses lawfully.

Keeping it alive

A prospect list decays at roughly twenty to thirty percent a year through job changes alone, before counting acquisitions and closures. Maintenance is not optional, but it is small if it is routine:

  • Process bounces immediately. Hard bounces come off permanently. Repeated soft bounces come off too.
  • Honour unsubscribes globally and permanently, in a suppression file that outlives any single tool. The same people reappear in the next source you use.
  • Re-verify before major sends, and treat anything over six months old as unverified.
  • Retire the persistently silent. Someone who has ignored six sequences over a year is not going to respond to the seventh, and mailing them depresses the engagement signal for everyone else.
  • Record which sources produced revenue, and spend your time on those.

The mistakes that waste the most time

  • Collecting addresses before defining the target. Produces volume that cannot be filtered afterwards, because you did not keep the information you would filter on.
  • Treating list size as progress. Fifty thousand unqualified addresses is a liability with a hosting cost; two thousand qualified ones is a pipeline.
  • Sending to constructed addresses without verifying. The fastest route to a bounce rate that gets you throttled.
  • Skipping the source field. Cheap to record, impossible to reconstruct, and the first thing you need when someone asks where you got their address.
  • Retyping addresses by hand. Guaranteed typos, invisible until they bounce.
  • Mailing everything at once from a cold domain. Even a perfect list gets filtered if the volume pattern looks like a compromise. See why your emails land in spam.

The through-line: every step that feels like it is slowing you down at the start is the step that makes the list worth having. The work is in the qualifying, not in the collecting.