Select some of this text to see the custom selection colors.

Strategy

How to Score an ICP List Automatically Before Outreach

Score your list on fit before you spend a single enrichment credit. The signals to use, how to build automated ICP scoring in Clay or with an LLM, and the score-before-enrich order that saves budget.

Apr 16, 2025

4 minutes

Joep van Acht

How to Score an ICP List Automatically Before You Reach Out

Most teams do this backwards. They buy or build a big list, enrich the whole thing, and then start deciding who's actually worth contacting. That's paying full price for data on accounts you'll delete. Flip the order and scoring becomes the cheapest, highest-leverage step in the whole pipeline.

What does it mean to score an ICP list automatically?

Automated ICP scoring assigns every account or contact a fit score from firmographic and signal data, without a human reading rows one at a time. Instead of eyeballing a spreadsheet and trusting your gut on row 400, you define what a good-fit account looks like, encode it as rules or an LLM prompt, and let the system rank the entire list in one pass. The output is a prioritized list: a top tier that earns your time and budget, and a bottom tier filtered out before it costs you anything. The reason this matters isn't neatness. It's that manual triage doesn't scale and quietly gets worse the longer the list, while a scored list stays consistent whether it has 200 rows or 20,000. It's the same "let the system do the judgment work" principle behind picking the right targets with AI.

What signals should an automated ICP score use?

Two layers, and keeping them separate is what makes the score useful. Fit signals describe whether an account looks like your best customers: industry, company size, geography, business model, and tech stack. Intent or timing signals describe whether now is the moment: recent funding, hiring for a relevant role, a leadership change, or activity on your site. Score fit first, because it's cheap to compute from firmographic data you can get before spending on enrichment. Then let timing signals raise the priority of accounts that already clear the fit bar, rather than dragging in poor-fit accounts just because they did something noteworthy. Put simply: fit qualifies, timing sequences. A well-funded company that isn't your ICP is still not your ICP. Where those timing signals come from, and how to act on them, is the domain of signal-based selling.

How do you build ICP scoring in Clay or with an LLM?

Structure the table first. Stage the list Raw to Master to Campaign so scoring never runs on a polluted sheet: Raw is the untouched import, Master is deduplicated and standardized, Campaign is the scored, ready-to-work slice. Add your firmographic columns, then score with either a formula for hard rules or an LLM column for the fuzzy criteria a formula can't handle, like judging whether a company's description actually matches your ICP. Ask the LLM for a number and a one-line reason in the same call, so every score is auditable instead of a black box. The discipline that saves you: validate a sample by hand before you trust the column across the whole list. An LLM will confidently mis-score a batch if the prompt is loose, so tighten the prompt against real examples first. This staging and scoring layer is a core part of what a GTM engineering function sets up.

Why should you score before you enrich?

This is the part that saves real money. Enrichment costs credits and outreach costs sender reputation, but fit scoring from firmographic data is nearly free. If you enrich the whole list first, you pay to find emails and phone numbers for accounts you were never going to contact, then delete them anyway. Score on cheap fit data first, keep only the top tier, and spend enrichment credits exclusively on that tier. The order is the entire trick: score, filter, then enrich. Done consistently, score-before-enrich cuts enrichment spend sharply, because you stop paying a waterfall to resolve contact data for rows that fail the fit test in the first place. It also keeps your sending list smaller and higher quality, which protects deliverability. Cheap data qualifies, expensive data confirms. The list you feed in matters too, which is why sourcing contacts off LinkedIn pairs with this to widen the top of the funnel before you score it down.

What score threshold should trigger outreach?

Set the threshold by capacity, not by a magic number someone posted online. If your score runs 0 to 10, a common pattern is to fully enrich and sequence the top band, hold the middle for lighter touches or a later cycle, and drop the bottom entirely. The right cut is the one that fills your team's actual outreach capacity with the highest-fit accounts and stops there. Scoring more accounts than you can work well doesn't help; it just dilutes attention. Treat your first threshold as a hypothesis, not a fact. As replies come in, look at which score bands actually convert and move the line accordingly. If your top band underperforms, your fit criteria are wrong, not your threshold. That feedback loop, score, act, learn, re-score, is what turns a static list into a targeting system that gets sharper every cycle.