Strategy
Why Your 2019 Cold Email Playbook Now Lands in Spam
The tactics that filled pipeline in 2019 are the exact behaviours mailbox providers now score against you. Here is what changed at Gmail, Yahoo and Microsoft, which old moves became spam triggers, and the sending system that replaces them.
Apr 16, 2025
8 minutes
Joep van Acht
Nothing about your copy got worse. The scoring did. The cold email playbook most teams still run was designed for a filtering regime that no longer exists: one where the message was graded, and the sender was mostly assumed innocent. Since 2024 the assumption runs the other way. Mailbox providers grade the sender first, and the message barely gets a vote. That is why a sequence that booked meetings in 2019 can be technically identical in 2026 and land in Promotions, Junk, or nowhere at all. This post covers what actually changed at the providers, which specific 2019 habits are now liabilities, and the sending system we install instead.
Why does cold email that worked in 2019 now land in spam?
Because filtering moved from content scoring to sender reputation, and the 2019 playbook was optimised for the wrong one. In 2019 you could register a domain on Monday, connect a fresh mailbox, and push 2,000 sends that week with a tracking pixel, an image signature and a link in the first email. The filter looked at the words, found nothing egregious, and delivered. Today the same behaviour reads as a brand-new, unauthenticated, high-volume sender with no engagement history, which is the exact profile of a spam operation. Google's sender guidelines now spell out authentication, unsubscribe and complaint-rate requirements as conditions of delivery rather than best practice. Yahoo published a matching set of requirements. Microsoft applied comparable rules to consumer Outlook accounts. Your copy is being judged after your infrastructure already decided the outcome.
What do Gmail, Yahoo and Microsoft actually require now?
Four things, and they are non-negotiable for anyone sending at volume. First, authentication: SPF and DKIM on the sending domain, plus a DMARC record with alignment between the visible From domain and the authenticated domain. A DMARC policy of p=none clears the bar; it is the presence of the record and the alignment that is checked. Second, one-click unsubscribe via list headers on bulk mail, honoured within two days. Third, a spam-complaint rate held under 0.3% as measured in Google Postmaster Tools, with 0.1% as the level you actually want to sit at. Fourth, sane basics: valid forward and reverse DNS on the sending IP, and TLS in transit. The thresholds are formally scoped to bulk senders (roughly 5,000 messages a day to personal accounts at Gmail), but the same signals are read for everyone. Small senders do not get a different filter, only a smaller sample.
Which parts of the 2019 playbook became spam triggers?
The habits that used to signal professionalism now signal automation. Open-tracking pixels fired on every send, routed through a shared redirect domain used by thousands of other senders, put you on someone else's reputation. A link in the first email to a cold contact who has never engaged is one of the strongest negative signals available. Image-heavy HTML signatures and templated layouts push a one-to-one email into bulk-mail classification. Forty to sixty sends a day from a domain registered last week is a volume curve no real person produces. Spintax that rewrites the same pitch a thousand ways still produces a thousand near-identical fingerprints. Sending from your primary company domain means one bad campaign damages every invoice and contract email you send. And an unverified list drives bounces, which is the fastest way to convert a new domain into a burned one. Each of these was standard advice in 2019. Each of them is now a reason to be filtered.
Do spam words still matter, or is it all reputation now?
Spam words matter, but not the way the 2019 checklists implied. No filter is holding a banned-word list and rejecting the message the moment it sees "free" or "guaranteed." Content scoring is a tiebreaker applied to a sender whose reputation is already ambiguous. What genuinely moves the needle is behavioural: complaint rate, bounce rate, whether recipients reply, delete without reading, or hit report-spam. That is where the word choice re-enters, indirectly and decisively. Copy that reads like a mass template gets marked as spam by humans, human complaints are the metric providers weigh most heavily, and your placement drops for everyone on the list. So the practical rule is not "avoid these 40 words." It is: write something a specific person has a reason to read this week, because their reaction is the input the filter is actually measuring. Reputation is downstream of relevance, not of vocabulary.
How much can one mailbox actually send per day in 2026?
Far less than the 2019 answer, and the number is not the interesting part. Our operating parameters across TechTower client systems: two to three mailboxes per sending domain, a ramp of roughly three to four weeks before a new mailbox carries campaign volume, and a steady-state ceiling in the low tens of sends per mailbox per day rather than the fifty-plus the old playbook assumed. Sending domains are always secondary domains bought for outbound, never the client's primary domain, so a bad week never touches contracts, invoices or support mail. Replies route into a single shared inbox (we use Missive) so response time stays low, which matters because replies are a positive engagement signal. Volume is then a function of how many domains you are willing to operate, not how hard you push any one mailbox. Pushing a single mailbox harder does not scale; it just fails more expensively.
How do you rebuild volume without burning domains?
Most companies respond to a deliverability problem by buying more inboxes. We do the opposite: cut the number of sends and raise the relevance of each one. The arithmetic is straightforward. Placement is driven by complaint and engagement rates, both of which are ratios. Sending to people with no reason to care pushes the ratio the wrong way, so more volume on a weak list actively destroys the asset you need for the next campaign. Targeting is therefore a deliverability control, not just a pipeline one. That means building the list from a live trigger rather than personalising a static list, which is the core of signal-based selling, and widening the reachable pool so you are not forced to over-send to the same shallow segment. Our open-web sourcing recovers roughly 1.6x more reachable contacts than a LinkedIn-only approach, because LinkedIn-only sourcing leaves around 58% of a market untouched. Fewer sends, better aimed, compounding reputation.
What does a 2026-compliant sending setup look like?
2019 playbook | 2026 system | |
|---|---|---|
Sending domain | Primary company domain | Secondary domains bought for outbound |
Authentication | SPF, sometimes | SPF + DKIM + aligned DMARC record |
Unsubscribe | Link in the footer, sometimes | One-click list header, honoured in under two days |
Volume per mailbox | 40 to 60 a day, immediately | Low tens a day, after a three to four week ramp |
Tracking | Pixel on every email, shared redirect domain | Open tracking off, no link in email one |
List | Bought or scraped, unverified | Verified, deduped, suppression-checked before import |
Personalisation | Spintax over a static list | List built from a dated trigger |
Health metric | Reply rate | Complaint rate under 0.3%, then reply rate |
Read the table as a sequencing guide rather than a menu. The infrastructure rows are prerequisites: no amount of copy work rescues an unauthenticated domain sending from a cold start. The list and personalisation rows are where the compounding happens once the infrastructure is sound.
What should you fix first?
In this order, because each step protects the next. One: split outbound onto secondary domains so your primary domain is never at risk. Two: set SPF, DKIM and an aligned DMARC record on every sending domain, then verify in Postmaster Tools rather than assuming. Three: turn off open tracking and remove links from the first email in every sequence. Four: verify and dedupe the list, and run it against your suppression list, before a single import. Five: ramp new mailboxes over three to four weeks instead of opening at full volume. Six: rebuild targeting around live triggers so the people receiving the mail have a current reason to reply, with worked examples in our signal-based selling plays. Steps one through five stop the bleeding. Step six is the only one that makes the system better over time rather than merely legal.
The reframe
Deliverability is not a settings problem that a consultant fixes once. It is a standing constraint that your targeting either respects or violates, every single send. The 2019 playbook treated the mailbox as free distribution and the list as a volume dial. In 2026 the mailbox is a reputation account you draw down with every irrelevant send and build up with every reply. Teams that keep pushing volume through the old playbook will keep buying domains to replace the ones they burn. Teams that rebuild around triggers, verified lists and disciplined sending get the compounding version, where placement improves as the system runs.
That is the difference between running campaigns and operating infrastructure, which is what GTM engineering actually means in practice. We design the architecture and operate it, typically live in two to four weeks, with no internal engineering required on the client side.