URL Reputation and Content Fingerprinting
How modern content filters judge message bodies — URL/domain reputation lists (SURBL zones, query mechanics, redirector handling) and fuzzy hashing (rspamd shingles, near-duplicate campaign detection) — and what that means for senders' link domains and templates.
Reference9 min read
Who it is for Senders, ESP operators
Applies to senders on any platform
ContentsOn this page — 4 sections
A message can be fully authenticated and sent from a clean IP address and still be filtered because of what it contains. The reputation of IP addresses and sending domains (blocklists) judges who is sending. Content filters add two more checks that judge what is being sent: the reputation of every domain linked in the message body, and a fuzzy fingerprint of the message that is matched against spam campaigns seen before.
Both checks work independently of the sender's own authentication and IP reputation. A message that links to a blocklisted domain, or whose fingerprint matches a known campaign, is filtered anyway.
Part 1: URL and domain reputation (SURBL)
SURBL is the model example of a URI DNSBL. Instead of listing sending IP addresses, it lists domains (and some IP addresses) that appear in the bodies of unsolicited or malicious mail. URIBL.com and the Spamhaus DBL and HBL (see Blocklists & Spamhaus) work the same way. Spam filters extract every URL from a message, reduce each one to a domain, and query that domain against these zones. A match typically adds a heavy score or blocks the message outright.
List zones and return values
The public data is served as one combined zone with bitmasked values, multi.surbl.org, which brings together these datasets:
| List | Meaning | Bit value (last octet) |
|---|---|---|
| DM | Disposable email domains | 4 |
| PH | Phishing sites | 8 |
| MW | Malware sites | 16 |
| CT | Click-tracker domains | 32 |
| ABUSE | General abuse or spam sites | 64 |
| CR | Cracked (compromised legitimate) sites | 128 |
- Responses are A records in the form 127.0.0.X. A domain on several lists returns the sum of their bits. For example, 127.0.0.80 means MW plus ABUSE (16 plus 64).
- NXDOMAIN means the domain is not listed, and an A record means it is listed. SURBL states that the A record is "the strongly preferred response for automated use."
- The default TTL is 60 seconds on the live multi zone, and the data itself is updated roughly every 30–40 seconds.
- A response of 127.0.0.1 is not a listing. It signals that the querier's access is blocked because of excessive volume on the public mirrors, and that the querier must sign up for SURBL's Sponsored Data Service.
- Beyond its domain lists, SURBL also runs HASHBL, reputation queries based on hashes, in these categories: abuse, cracked, malware, phish, email, crypto, phone.
- For filter operators, data is delivered as DNS (Private Query Service), rsync (recommended for high-volume mail filtering), RPZ (web filtering), a REST API, CSV, and RTF (real-time feeds in JSON).
The very existence of the CT (click tracker) list matters: tracking and redirect domains are a listing category in their own right, not an edge case.
How filters extract and query URLs (SURBL implementation guidelines)
SURBL's published guidance to filter developers describes what production content filters actually do:
- Extract every URI from the message, fully resolving redirections to the final target domain. In other words, filters are told to follow redirectors and shorteners to the destination and check that domain, not only the visible link.
- Reduce URIs to domains or subdomains. With the wildcarded multi zone, subdomains of a listed domain match automatically, so filters do not need to normalize to the level of the registered domain.
- Do not resolve the extracted domains in DNS. The domain string itself is the lookup key.
- Query by prepending the domain to the zone (
domainundertest.com.multi.surbl.org) and doing an A-record lookup. - URLs with numeric IP addresses (
http://10.20.30.40/) are checked with the octets reversed, as DNSBLs do:40.30.20.10.multi.surbl.org(octets in base 10). - Keep a local allowlist of well-known, high-traffic domains (yahoo.com, w3.org, google.com) to avoid pointless queries.
- Check that answers fall within 127/8. An answer outside 127.0.0.0/8 indicates a DNS resolver that uses wildcards or redirects and corrupts the results; run a local caching nameserver.
SURBL explicitly prohibits two practices, because both cause false positives with shared hosting: do not use URI data from message bodies to check sender IP addresses, and do not resolve listed domains to IP addresses and blocklist those IP addresses.
What this mechanism implies
- Listing one domain affects every message that links to it, from every sender and on every IP address. For an ESP, a URI DNSBL listing of a shared domain is an incident that affects many customers.
- Because filters resolve redirects, hiding a bad destination behind a shortener or your own redirect domain does not evade the check. Instead, it exposes your redirect domain to listing when the destinations behind it are abused. This is exactly what SURBL's CT list captures.
- Because subdomains of a listed domain match under the wildcarded zone, a listing at the level of the registered domain takes down every customer subdomain beneath it.
Part 2: content fingerprinting with fuzzy hashing (rspamd)
Fuzzy hashing detects that a message is a near-duplicate of spam reported before, even after the spammer changes it. rspamd's fuzzy_check module and fuzzy_storage worker are a fully documented open implementation. Commercial fingerprinting systems, and shared hash networks such as Razor, Pyzor and DCC, follow the same principles. rspamd also provides a public shared fuzzy feed that many installations query by default, so a fingerprint learned anywhere can add to scores everywhere.
How the fingerprint is built: shingles
- The message text is split into words, then into overlapping sequences of three words (trigrams), called "shingles."
- Each shingle is hashed with multiple hash functions (32 hashes per shingle), and the fingerprint of the message is the set of these hashes. (rspamd cites Broder's research on resemblance and shingling as its basis.)
- Comparison is probabilistic. Similarity is computed from the number and positions of shingle hashes that match between the candidate message and the stored hashes, which gives a percentage match rather than exact equality.
- Hash algorithms for text:
mumhash(the current default, and recommended),xxhash,fasthash,siphash(legacy). Changing the algorithm invalidates all stored data. - Images and attachments are matched exactly, not fuzzily, using blake2b digests of their content. A single changed byte breaks the match. This is why image spam campaigns render their images again for each send, and why filters combine exact hashes with perceptual image hashing elsewhere.
Key defaults in fuzzy_check: min_bytes = 1k (the minimum size of an attachment or image considered), min_height/min_width = 32 px for images, min_length = 0 (all text parts are checked), text_multiplier = 4.0, timeout = 2s, retransmits = 1. mime_types selects which attachment types (for example application/*) are hashed, and short_text_direct_hash hashes texts too short for shingling exactly.
HTML-structure fuzzy hashing (rspamd v3.14+)
Since v3.14.0, rspamd can also fingerprint the structure of the DOM, independently of the text. It uses tokens of the form tagname[.class][@domain] (for example a.button@example.com). Only the first CSS class is used, known tracking classes are filtered out, and link domains are normalized to eTLD+1.
The combined hash is weighted as follows: structure shingles 50%, call-to-action (CTA) domains 30%, all link domains 15%, and counts of features (tags and links) 5%. To avoid matching generic templates, it requires at least min_html_tags (default 10; example configurations use 15), at least 2 links, and a DOM depth of at least 3.
The weight on CTA domains targets phishing. A phishing message that perfectly clones a brand's template gets a structure similarity of ~0.9, but with different CTA domains the combined similarity collapses (the documented example: 0.9 structure similarity with mismatched CTA domains gives a combined 0.45). Conversely, a spam campaign that keeps its CTA domain while shuffling its copy still matches.
Why mutations and token stuffing don't evade it
The fingerprint is a large set of overlapping word-trigram hashes, compared probabilistically. As a result:
- Small edits (swapped words, inserted names, reordered paragraphs) change only the shingles that overlap the edit. Most shingles still match, and similarity stays above the threshold.
- Token stuffing (appending random words, hidden text, or strings meant to break hashes) adds shingles but does not remove the ones that match. It only dilutes the similarity ratio slightly, while the absolute number of known-bad shingles still matches. Defeating shingle matching requires rewriting essentially the whole body, and even then HTML-structure hashing (v3.14+) still matches the unchanged template and CTA domain.
- Merge fields for each recipient, unsubscribe URLs and tracking tokens also leave most of the shared shingles intact. This is exactly why a campaign reported by its early recipients is recognized in the rest of the send.
Weights, thresholds, and gradual scoring
Stored hashes carry a weight that grows as sources (user reports, trap hits, honeypots) report the same fingerprint again. Scoring is deliberately gradual, not all or nothing:
- Each flag or rule has a
max_score(a threshold on hash weight). Below the threshold, the symbol scores 0. The score then rises from the threshold to 2× the threshold, along a hyperbolic tangent curve: ≈50% of the metric weight at the threshold, and the full score at 2×. For example, with a report weight of 1 and a threshold of 20, a fingerprint needs ≥ 20 independent complaints before it scores at all. - Hashes are organized by numeric flags that map categories to symbols. The conventional layout: flag 1 is confirmed spam (
FUZZY_DENIED, max_score 20), flag 2 is probable spam (FUZZY_PROB, max_score 10), and flag 3 is legitimate, allowlisted content (FUZZY_WHITE, max_score 2). Flag numbers must be unique across writable rules, andskip_hashesallowlists specific fingerprints. - Learning:
rspamc -f <flag> -w <weight> fuzzy_add <message>(or-S FUZZY_DENIED) adds a hash, andfuzzy_delremoves one.read_only = truemakes a rule query-only, and an optional learning condition written in Lua can block learning or change its flag.
On the operational side, the fuzzy_storage worker:
- has a single-writer architecture, with an update queue synced to disk every
sync = 1min; - stores data in SQLite (the default) or Redis;
- lets hashes expire through
expire(90d is the recommended initial value, so old fingerprints age out as campaigns die); - needs roughly 400k hashes per 100 MB of RAM (1.5M ≈ 500 MB);
- accepts learning only from
allow_updateIP addresses; - supports optional Curve25519 transport encryption (
encrypted_only = true), and master and slave replication (TCP 11335) with flag translation for each slave.
Sender implications
Together, the two mechanisms lead to concrete rules for anyone running an ESP or a sending program:
- Every domain in the body carries reputation: link domains, image-hosting domains, redirect and click-tracking domains, and even domains that appear in plain text. Their reputation is evaluated separately from the From domain and the sending IP address.
- Customers who share a tracking domain share its fate. A click-tracking or link-wrapping domain used across an ESP combines the reputation of every customer's destinations. One abusive customer can get it listed on a URI DNSBL (SURBL's CT category exists for exactly this), which filters every customer's mail at once. The mitigations are tracking subdomains for each customer on domains the customer owns (through a CNAME), proactive URI DNSBL monitoring of tracking domains, and screening destination URLs against SURBL, URIBL and the DBL at send time.
- Public shorteners do more harm than good. Filters resolve them to the destination anyway, and the shortener's domain carries the accumulated reputation of everyone who has abused it. See Content & Design for Deliverability.
- Keeping link domains clean is an ongoing monitoring task, not a one-time check. Query your tracking, image and landing domains against multi.surbl.org (and the Spamhaus DBL) as often as you check IP blocklists, and check the URLs of every outbound campaign before you send it.
- Fingerprinting makes complaints count against the rest of the send. The first few thousand recipients who report a campaign create its fuzzy hash and add weight to it, and the rest of the same send then matches it. Sending in segments with pauses (to the most engaged recipients first, watching early complaint signals, then continuing) takes direct advantage of the scoring window between the threshold and 2× the threshold.
- Changing a template does not reset its content reputation. Minor changes to the copy, rotating subject lines, or personalizing merge fields leave the shingle fingerprint (and the fingerprint of the HTML structure and CTA domains) intact. The only real reset is genuinely different content sent to an audience that wants it. Content reputation is a symptom of list quality, not a problem to solve by engineering the content.
- Consistent CTA domains cut both ways. Keeping your links on your own stable domains with good reputation helps structure hashing tell you apart from phishing clones of your template. Scattering links across throwaway domains looks like evasion.
Related articles
- Blocklists & Spamhaus, on how DNSBLs for IP addresses and domains work, including the Spamhaus DBL and HBL, the other major datasets of domain reputation
- Content & Design for Deliverability, with the practical content rules (shorteners, images, links) that these mechanisms explain
- Reputation Monitoring, on where URI DNSBL checks of tracking and link domains belong in a monitoring program
Check your own record
The free check reads what your domain publishes in DNS.
In this topic
- Deliverability Glossary
- Email Abuse Taxonomy
- DNS Blocklists and the Spamhaus Zones
- Spamhaus Listings Deep Dive — SBL, CSS, PBL, DBL Policy and Delisting