Namefi

The Cat-and-Mouse War of Email Sender Reputation

A source-backed history of the thirty-year arms race over email sender reputation—the open relays, botnets, image spam, snowshoe campaigns, warmup networks, and other tricks bulk senders used to look trustworthy, the named operators who ran them, and the year and mechanism by which each trick was detected and shut down.

Aileen WrightAileen WrightAuthorVictor ZhouVictor ZhouEditorJul 14, 2026est. 23 min read
  • email
  • sender-reputation
  • deliverability
  • dmarc
  • domain-security
Share on X

Every few months someone asks a version of the same question: is there a shortcut to email reputation? A way to make a fresh sending domain look established, so cold outreach lands in the inbox instead of the spam folder? A whole industry sells answers—aged domains, warmup networks, rotation tools, "deliverability" services. Almost all of them work for a while and then stop working.

The reason they stop is the subject of this article, because it explains a thirty-year war. Email is one of the oldest adversarial systems on the internet, and its reputation machinery is the scar tissue from that war. Sender reputation is not really a score for how well you follow the rules. It is a proxy for one question that mailbox providers actually care about: do real humans want this mail? Every trick in the history of bulk email is an attempt to fake the signals that answer that question. Every countermeasure is an attempt to measure the real thing more directly. Because the ground truth—human wantedness—does not move, the measurement can only get closer to it over time. That is why the shortcuts decay.

What follows is the history of that arms race, told through the tricks themselves: what spammers did, why it worked, who ran the biggest operations, and the year and the mechanism by which each move was detected and shut down. As a registrar and DNS operator, Namefi lives in the part of this story where domains, DNS records, and authentication meet deliverability, so the naming layer gets particular attention. But the pattern is the point, and it repeats so cleanly that by the end you can predict where any new "reputation hack" will land.

What sender reputation actually measures

A mailbox provider like Gmail or Yahoo cannot read your mind, and it cannot survey a recipient before delivering. So it infers wantedness from behavior it can observe: do recipients open, reply, and keep your mail, or do they delete it unread, never engage, and click "report spam"? Google exposes a sanitized view of this in Postmaster Tools, which sorts a domain's and IP's reputation into High, Medium, Low, and Bad—where "High" mail is "rarely marked as spam by Gmail" and "Bad" mail is "almost always marked as spam or rejected by the receiving server."

The exact algorithms are proprietary and change constantly, so nobody outside these companies can describe Gmail's internal code. But the documented direction of travel is consistent and public, and it is enough to explain everything below. Keep the frame in mind through every era: each trick manufactures a fake version of "humans want this," and each countermeasure finds a way to notice the fake.

The open frontier: relays and the first blocklists (1994–2003)

In the early internet almost everything was trusting by default, and spammers exploited that trust directly. Mail servers were commonly shipped as open relays—configured to accept and forward mail from anyone. Spammers routed their bulk mail through other people's misconfigured servers, hiding the origin and dumping the bandwidth cost on the victim. Wikipedia's history notes that until the 1990s, "mail servers were commonly intentionally configured as open relays" and that closing them "reduced the percentage of mail senders that were open relays from over 90% down to well under 1% over several years." The turning point on the software side came in 1998, when Sendmail changed its default so that relaying was "denied by default. Note that this changed in sendmail 8.9; previous versions allowed relaying by default."

Closing relays one server at a time was too slow, so the defense that actually scaled was the blocklist. In 1997, Dave Rand and Paul Vixie turned their private list of spam-sending addresses into the Real-time Blackhole List, "created in 1997, at first as a Border Gateway Protocol (BGP) feed by Paul Vixie" and then as the first DNS-based blocklist (DNSBL). Others followed: SpamCop (1998), SORBS (2001), and the Spamhaus Project, founded in London by Steve Linford in 1998, which launched its Spamhaus Block List in 2001 and later added the exploited-host XBL (2004) and the Policy Block List (2007). Blocklists became infrastructure: a receiving server could reject a connection in milliseconds based on the sender's reputation.

Two smaller tricks from this era are worth naming because their countermeasures still matter. In a directory harvest attack, a spammer brute-forces likely addresses (jdoe@, johnd@) against a domain and keeps the ones the server does not reject—Wikipedia describes it as "a technique used by spammers in an attempt to find valid/existent e-mail addresses at a domain by using brute force." And in backscatter and the joe job, a spammer forges an innocent third party into the "From" line, either to dodge blame or to frame a rival; the term traces to a 1997 retaliation against the owner of joes.com. Both tricks pushed receiving servers toward verifying senders before accepting mail rather than bouncing afterward.

The escalation cut both ways. In 2003, spam operators fought back by attacking the blocklists themselves: the firm Osirusoft, "an operator of several DNSBLs... shut down its lists" in August 2003 after sustained denial-of-service attacks. The arms race was already, unmistakably, a war.

The law arrives, and the spam kings fall (2003–2012)

For a long time bulk email had no dedicated law in the United States. That changed with the CAN-SPAM Act, signed December 16, 2003 and effective January 1, 2004. Anti-spam advocates were scathing—the law was criticized because it "neglects to actually tell any marketers not to spam," legitimizing opt-out bulk mail rather than banning it—but it gave prosecutors a tool, and over the following decade the era's biggest operators fell one by one.

The mechanism here was law plus reputation-tracking. Spamhaus's Register of Known Spam Operations followed repeat offenders as they hopped IP ranges, so that a spammer terminated by one provider could not simply reappear clean somewhere else. But arresting kingpins did not end spam, because by the time the courts caught up, the underlying infrastructure had already moved somewhere the law could barely reach: other people's computers.

The botnet decade: renting a million strangers' PCs (2007–2013)

If reputation systems judge you by the server you send from, the winning move is to send from millions of servers you do not own. That is exactly what spam botnets did, harnessing consumer PCs infected with malware, each sitting on a residential IP address with no sending history. This was the golden age of spam volume, and the names ran the internet's junk mail for years: Storm, Srizbi, Rustock, Cutwail, Grum, Waledac, Kelihos, and later Necurs. At its September 2007 peak, "compromised proxies in the Storm worm botnet were throwing out 20 per cent of the world's junk emails." Cutwail at its 2009 peak was estimated to run "around 1.5 to 2 million individual computers, capable of sending 74 billion spam messages a day."

Two landmark academic studies from this period pulled back the curtain on why anyone bothered. In 2008, researchers behind the "Spamalytics" study parasitically infiltrated the Storm botnet and measured its actual conversion rate: "After 26 days, and almost 350 million email messages, only 28 sales resulted—a conversion rate of well under 0.00001%." Spam was profitable only because sending was essentially free once you were stealing the computers. In 2011, the "Click Trajectories" study followed the money end to end and found the real weak point was not sending but banking: 95% of spam-advertised pharmaceutical, replica, and software products were monetized through merchant accounts at just a handful of banks. The reporting of Brian Krebs, later collected in his book Spam Nation, put faces to that economy: rogue-pharmacy affiliate programs like GlavMed, which "generated revenues of at least $150 million" across 2007–2010, and its rival Rx-Promotion, run by ChronoPay's Pavel Vrublevsky under the alias "RedEye," whose feud with competitors—the "Pharma Wars"—helped bring the whole ecosystem down.

The countermeasures came from three directions, and the years are worth marking:

Botnets never fully disappeared, but by the mid-2010s the economics had shifted. Volume alone stopped working, because the receiving side had learned to judge the source network, not just the individual message.

Dressing up the message: the content-filter arms race (2002–2012)

Running parallel to the infrastructure war was a fight over the words themselves. Early filters were hand-written keyword rules, easily dodged by misspelling "Viagra." The reset came in August 2002, when Paul Graham's essay "A Plan for Spam" popularized Bayesian filtering—the insight that "you can filter present-day spam acceptably well using nothing more than a Bayesian combination of the spam probabilities of individual words." Crucially, it adapted on its own: "as spammers start using 'c0ck' instead of 'cock' to evade simple-minded spam filters based on individual words, Bayesian filters automatically notice." Within two years the technique was everywhere.

Spammers answered, and the filters answered back:

  • Bayesian poisoning and hash busters (2004). Spammers padded messages with random or innocent-sounding "word salad" to dilute the spam score, and appended random character strings to defeat checksum filters—a hash buster "randomly adds characters to data in order to change the data's hash sum." Filters adapted by 2004–2005 through remote-image blocking, down-weighting common words, and frequent retraining.
  • Image spam (peaked late 2006). To defeat text filters entirely, spammers rendered the pitch—often penny-stock pump-and-dump schemes—inside an image with visual noise to foil optical character recognition. Image spam "peaked at the end of 2006, when over 50% of spam was image spam." The countermeasures—OCR, image fingerprinting, and near-duplicate clustering that recognized the same image across recipients—worked fast: image spam "accounted for only eight percent of all spam during July, a drastic decrease from January, when it totaled 52 percent of junk email" in 2007.
  • Attachment spam (August 2007). As image spam collapsed, spammers moved the pitch into PDF and Excel attachments. The wave crested and broke within weeks: PDF spam "went from 30% of all spam sent on Aug. 7 to less than 1% on Aug. 29" once vendors added document parsing.
  • Template spinning (mid-2000s). Spin syntax—{hi|hello|hey}—and mail-merge fields made every copy byte-different to defeat exact-match filters. The counter was fuzzy hashing that tolerates small changes, in tools like the Distributed Checksum Clearinghouse, whose "fuzzy checksums are changed as spam evolves," and the locality-sensitive Nilsimsa hash. The same near-duplicate principle later caught spun image spam.
  • Link obfuscation (2001–2011). Spammers hid destinations behind redirectors, URL shorteners, and disposable domains, and used look-alike Unicode characters in homograph attacks (first published in 2001, demonstrated against PayPal in 2005). Browsers responded by displaying suspicious internationalized domains in raw Punycode, and filters began resolving and reputation-scoring links before delivery.

The deeper shift in this era was philosophical: filtering moved from reading keywords to reading structure and behavior. Once the filter stopped asking "what words are in this?" and started asking "have I seen this shape sent to a thousand people?", spinning the words no longer helped.

Outrunning reputation itself: the network-level games (2006–2020)

By the late 2000s, reputation systems were good enough that the remaining tricks were all about timing and identity—ways to send before the score could catch up, or to borrow a score you had not earned.

Proving who you are: the authentication era (2003–2021)

For decades the "From" address was pure decoration—you could put anything there, which is exactly what phishers did. Closing that hole took three DNS-based standards, layered over more than a decade because each fixed a gap the previous one left.

SPF (Sender Policy Framework) let a domain publish which servers may send for it. Its 2006 specification opened by stating the problem plainly: "E-mail on the Internet can be forged in a number of ways." But SPF only checks the hidden envelope sender, not the "From" a human sees, and it breaks when mail is forwarded. Microsoft's competing Sender ID tried to check the visible header but "fewer than 3% of mail domains" adopted it, partly because its patent licensing was incompatible with open-source software. DKIM, which grew out of Yahoo's 2004 DomainKeys, added a cryptographic signature so a receiver can verify a message "using public-key cryptography" and confirm a domain takes responsibility for it.

The piece that tied it together was DMARC, launched by an industry group on "January 30, 2012" and published as RFC 7489 in 2015. DMARC added alignment: the visible "From" domain must match the domain that actually passed SPF or DKIM, so a sender can no longer authenticate as one domain while displaying another. The transition was not painless. In April 2014, after a wave of address-book spoofing, Yahoo and then AOL published strict p=reject DMARC policies, which broke every mailing list that rewrites messages—one engineer reported "a blizzard of bounces... and the list got a whole bunch of rejections from Gmail, Hotmail, Comcast, and Yahoo itself." That breakage drove the next standard, ARC (2019), which preserves authentication results across forwarders, and BIMI (general in Gmail in 2021), which rewards authenticated senders with a verified logo.

Authentication does not prove you are wanted; a perfectly signed message can still be spam. What it does is weld every message permanently to a domain identity that accumulates reputation. Once you cannot forge the "From," you cannot escape your own track record by wearing someone else's name. That is why, when Google and Yahoo finally made SPF, DKIM, and DMARC mandatory for bulk senders in February 2024, it did not so much change the game as make explicit a foundation that had been building for twenty years.

Faking the ground truth: warmup and manufactured engagement (2018–2026)

The modern cold-email industry attacks the most direct target of all: not the IP, not the domain, not the "From" line, but the engagement signal itself. If reputation is driven by recipients opening, replying, and rescuing mail from spam, then warmup networks manufacture exactly those actions. A pool of mailboxes—thousands of them—automatically emails one another, opens the messages, replies in threads, marks them important, and drags anything that lands in spam back to the inbox. To the provider it is meant to look like a domain whose mail people love. Mechanically it is a fake engagement graph.

This is the hardest fake to sustain, and understanding why is worth a short detour into web search. In 2004, Stanford and Yahoo researchers published "Combating Web Spam with TrustRank," a variant of PageRank that fights link farms by "selecting a small set of seed pages to be evaluated by an expert" and letting trust flow outward from that known-good core, on the assumption that reputable pages rarely link to junk. The consequence is that a cluster of pages linking only to each other cannot bootstrap trust: there is no path from the trusted seed into the closed island.

A warmup network is that island. This is an analogy, not a claim about Gmail's proprietary code, but the structural problem is real: engagement that circulates only inside a pool of sender-controlled mailboxes has no edge connecting it to people who genuinely wanted the mail, and trust that transfers has to originate from real recipients. The documented reality now matches the theory. In early 2023, Google forced the largest such service to close, telling the operator of GMass—which ran "more than 80,000 email accounts at any given time"—to "shut down our warmup system or we'd lose our Gmail API access." In 2024, the sales platform Apollo.io retired its native warmup feature in favor of a tool that only paces send volume, with no engagement simulation. And independent testing has been unkind: one deliverability firm reported that after testing nearly all major warmup tools, "none of them demonstrated any meaningful improvements in deliverability or reputation." A curve of opens and replies that is too smooth and too perfect is itself a signal, because real human attention is noisy.

The same years quietly demolished the spammer's favorite measurement. Apple Mail Privacy Protection, introduced in 2021, broke open tracking by routing remote content through "two separate relays operated by different entities" and loading it in the background whether or not the recipient opened anything—so the tracking pixel that warmup tools rely on to prove "engagement" stopped meaning anything for a huge slice of recipients.

Two older defenses close the loop on list quality by measuring the real thing directly. Spam traps are addresses that exist only to catch senders who did not get permission: pristine traps that were never valid, so hitting one proves scraping or guessing, and recycled traps that were once real but decommissioned, "allowed to bounce for a minimum of 12 months and then reopened," which catch stale lists. Project Honey Pot has run a distributed trap network since 2004. And feedback loops report complaints straight back to the sender in a standard format so there is no ambiguity about whether recipients wanted the mail—they said so, by reporting it.

2026: how to do it right

Read the eras together and the shape is unmistakable. Every trick is a way to fake the answer to "do humans want this?"—a fake relay origin, a fake From address, a fake image of legitimacy, a fake individual message, a fake IP history, a fake engagement graph. Every countermeasure is a way to measure the real answer more directly, whether by welding identity to a domain, analyzing behavior instead of keywords, or reading complaint rates as ground truth. The measurement keeps converging on the thing it has always been a proxy for, which is why the shortcuts have a shelf life and the target does not move.

So here is what actually works in 2026, stated plainly.

Treat authentication and hygiene as the floor, not the achievement. SPF, DKIM, and DMARC with alignment are now required for bulk senders by both Google and Yahoo; Google defines a bulk sender as one sending "more than 5,000 messages per day to Gmail accounts" and requires you to "keep spam rates reported in Postmaster Tools below 0.3%." One-click unsubscribe, honored promptly, is mandatory. None of this earns you reputation. It only gets you to the starting line, and its absence disqualifies you.

Send mail people asked for and are glad to get. What actually moves reputation is unchanged from the beginning of this history: a real prior relationship, genuine two-way engagement, and content relevant enough that recipients open and reply because they want to. Low volume paired with honest personalization works not because it games the filter but because it tends to produce welcome mail. Permission is not a compliance checkbox; it is the thing every countermeasure in this article was built to measure.

Stop paying for shortcuts that are already decaying. Aged domains carry residual reputation you cannot see and did not earn. Warmup, at any real scale, is a gray tactic on a downward slope—detectable, increasingly penalized, and capable of dragging associated domains down with it. Money spent on either is better spent on a list of people who want to hear from you.

There is only one move in this entire war that never gets patched, because it is not a trick: send mail that people are glad to receive. Everything else is a temporary lead in a race the measurement always eventually wins.

For teams that run their own sending domains, this is also why the unglamorous infrastructure matters. The naming and authentication layer—your DNS records, your SPF and DKIM and DMARC alignment, the domain identity every message is welded to—is the foundation the whole reputation system is built on. It will not make unwanted mail wanted. But get it wrong and even wanted mail struggles to arrive, and let it be hijacked and your hard-earned reputation becomes someone else's phishing tool. As a registrar and DNS operator, that control plane is the layer Namefi works to get right, so that the reputation you actually earn is the reputation your recipients actually see.

Sources and further reading

Reputation and requirements (primary)

Blocklists, relays, and the early era

The law and the spam kings

Botnets and the spam economy

Content-filter arms race

Network-level games

Authentication

Warmup, engagement, and list quality

Contributors

Aileen Wright
Art & History Writer • Namefi

Aileen Wright is a student in her twenties living in New York City, where the distance between a museum wall and a library reading room is a short walk and a long afternoon. She came to name writing through art and history — the way a single portrait, coin, or manuscript margin can carry a name across centuries and change its meaning on the way.

Most weeks you can find her in Central Park with a paperback, or in the quiet of a public reading room chasing down where a name actually comes from rather than what a name-list says it means. She is also teaching herself to code, which has made her oddly precise about spelling, sorting, and the small details that decide whether a name ages well.

For Namefi she writes about the history and culture behind domain names, the stories brands carry as they rename, and the difference between a good story and a verified source.

Victor Zhou
Founder & Standards Editor • Namefi

Victor Zhou is a technology founder and standards editor focused on digital identity and trust. He founded Namefi, edits Ethereum Improvement Proposals, and previously led smart-contract architecture work at Google Labs.

His work sits at the intersection of naming, ownership, and the systems people use to establish identity online. That perspective makes him especially interested in the way names move between personal meaning, public recognition, and digital infrastructure.

For Namefi, Victor edits and writes about domains as durable digital identity: how names become ownable onchain assets, how tokenization changes custody and trust, and what naming can learn from the systems people use to establish identity online.

Related guides

Discuss this post

View the discussion on Namefi Discuss