emailmarketing.net

Vetting Automation & Third-Party Risk Signals

The signal-vendor landscape for automating ESP customer vetting — email-risk APIs, IP/device fraud scoring, domain intelligence, payment-fraud signals, breach lookups — and the automation architecture that composes them into onboarding gates, progressive trust ladders, and continuous re-scoring.

Operational17 min read

Who it is for ESP operators

If you run an email service provider (ESP), you can automate much of customer vetting by buying signals from third-party vendors and combining them into risk scores, gates and progressive trust. Customer Vetting covers what to check: the questionnaires of the Messaging, Malware and Mobile Anti-Abuse Working Group (M³AAWG), the red flags, and the method for manual review. The guidance below covers which tools to use and how to automate the checks: the signal vendors an ESP can connect to its signup and onboarding flow, and the architecture that combines those signals into risk scores, gates and progressive trust.

The M³AAWG Hosting Abuse Best Common Practices (BCP) already requires the outcomes: "fraud-score prospective accounts and auto-reject below threshold," "put limits on new accounts," and "tiered rights allocation" (see the fraud-prevention practices). What follows is how to implement them.

How much of this is established. No published industry standard covers how to combine vetting signals. The patterns for combining them below are practitioner consensus, and inference from what major platforms publicly document about their own onboarding (attributed by name throughout). Claims about vendor capabilities come from vendor documentation. Where only the vendor's marketing asserts accuracy or effectiveness, the text says so. No third-party risk score has independently audited accuracy figures. Treat every vendor accuracy claim as marketing.

Two different questions, two different signal sets

Vetting automation answers two separate questions that are easy to confuse:

  1. Is this signup a fraudulent or abusive actor? (a spammer, a snowshoe operator, a fraudster using stolen cards, a phisher) IP and device intelligence, payment-fraud signals, domain history and identity checks on the customer answer this question.
  2. Is this customer's data dangerous? (purchased lists, harvested addresses, files full of traps) Email validation and risk APIs run against the customer's uploaded lists answer this one, together with the acquisition questions from the vetting questionnaire.

A customer can be legitimate and have terrible data, or look clean and intend fraud. Automation must score both. Gates that mix the two questions pass polished fraudsters and reject honest small businesses that have old lists.

The signal-vendor landscape

Email-address validation vs. risk scoring

Vendors bundle these two product classes together, but they overlap only partly and serve different purposes in vetting:

  • Deliverability validation answers "will this address accept mail?" It covers syntax, whether an MX record exists, mailbox probing at the SMTP level, catch-all detection, and flags for role accounts and disposable addresses. This is list-hygiene tooling (see List Hygiene) put to use for vetting. Run it on a sample of the list a customer uploads, and the overall profile (invalid %, role-account %, disposable %, catch-all %) is a strong indicator of how the addresses were acquired. High invalid rates point to old or purchased data, the same logic Mailchimp documents for Omnivore (below).
  • Risk scoring answers "is this address likely to hurt me?" It covers the likelihood of a trap, complaint history, and predicted engagement. It is probabilistic by nature and specific to each vendor, so use it to rank addresses relative to each other, not as a verdict.
Vendor (example of the class) Documented capabilities Notes
Kickbox Result classes deliverable, undeliverable, risky and unknown, with reason codes. A proprietary Sendex quality score (0–1) for each address The meaning of the score is documented. Specific cut-off values quoted in reviews (for example, "0.7 = safe") are practitioner heuristics, not guarantees the vendor publishes
ZeroBounce Statuses valid, invalid, catch-all, spam_trap, abuse, do_not_mail and unknown. "AI scoring", a 0–10 prediction of engagement. Activity data (engagement over the past year) The abuse status (people who habitually click the spam button) and the spam_trap flag come from the vendor's own datasets. Their coverage and accuracy are marketing claims that no one has verified independently. Real trap operators do not sell trap lists (see Spam Traps)
BriteVerify (Validity) Verification focused on deliverability: valid, invalid, accept-all, unknown Typical of the pure validation class (with NeverBounce, Emailable and others), with no fraud-risk layer

How to use it for vetting (practitioner consensus): validate a random sample of any list a prospect brings, which follows the sampling logic of the M³AAWG test-send method. Decide on the overall profile, not on individual addresses. A list that is 15% undeliverable was not built by confirmed opt-in, whatever the questionnaire says.

IP, device, and identity fraud scoring

These vendors score the signup itself: the IP address, device, email address and identity the prospect used to register.

Vendor Documented capabilities Score model
IPQualityScore (IPQS) Detection of proxies, VPNs and Tor. IP reputation. "Abuse velocity" (recent abusive activity from the IP). Device fingerprinting. An email risk API (disposable detection, "history of fraudulent behavior online"). Phone reputation A fraud score from 0–100. IPQS's own documentation treats a score of 90 or more as high risk (a threshold the vendor recommends; tune it to your own outcomes)
MaxMind minFraud Transaction risk scoring over the IP, email, address and payment fields. Email and domain intelligence, including the domain's first_seen date on the minFraud network (recorded since 2019) and a domain risk score. IP geolocation data (GeoIP) A risk score from 0.01–99, documented as a calibrated probability of fraud (a score of 20 means about a 20% chance)
Sift Machine-learning scores for each type of abuse: payment fraud, account abuse and content abuse. Trained on the customer's own event stream and on Sift's network across customers. An event-driven API (you send signup, login and content events, then poll for scores or receive them) A score from 0–100 for each type of abuse. A user can score high for one type and low for another

Vendors that work across a network (minFraud, Sift, IPQS) get their value from seeing the same actor across many customers. That is also their weakness. Their coverage claims ("data from hundreds of millions of users") are marketing that cannot be verified, and you cannot know how often they flag legitimate users in your own population until you measure it. The facts documented independently are the API contracts and the meaning of the scores above.

Signals specific to deliverability add to the same data. A signup from an IP address on Spamhaus AuthBL (sources of credential stuffing; see Blocklists & Spamhaus), or a signup through an anonymizing proxy, is the automated form of the registration-abuse red flags in the Hosting Abuse BCP, described in Compromised Accounts.

Disposable / temporary-domain detection

Signups that use throwaway mailboxes correlate strongly with abuse of free trials and with snowshoe registration (practitioner consensus; no base rates are published).

  • Open-source lists: the community-maintained disposable-email-domains blocklist (maintained since 2014, with a submission process that requires evidence) is the de facto free baseline. Several automatically generated forks track providers that change domains quickly.
  • Commercial detection (IPQS, ZeroBounce, Kickbox, and services like UserCheck) claims fresher coverage of disposable providers that keep changing domains. The claim is plausible but not audited.
  • Design caveat: disposable-domain providers rotate domains precisely to evade static lists. Treat a domain missing from the list as "unknown," not "clean." A newly registered domain signal (below) catches much of the rotation the lists miss.

Domain intelligence: the customer's sending domain

The domain a prospect wants to send from is the richest single thing to vet. All of these checks can be automated:

Check Mechanism Signal
Registration age Creation date from RDAP or WHOIS A domain registered days ago that claims years of sending history contradicts the questionnaire. New domains carry almost no reputation (Sending Infrastructure Practices)
Zero-reputation window Spamhaus ZRD lists newly registered domains, and domains that were dormant, for 24 h, with return codes that encode the domain's age in hours A documented Spamhaus dataset, available only through DQS. A prospect whose domain is on ZRD at signup has, by definition, no history to vet
Blocklist history Spamhaus DBL (spam domains, with return codes that separate spam, phish, malware and abused-legitimate), equivalents such as SURBL, through DQS or data feeds Spamhaus explicitly markets its domain-reputation data for vetting ESP prospects and for placing customers in pools by risk. The abused-legitimate distinction matters, because it flags a compromised customer rather than a bad one
DNS posture SPF, DKIM and DMARC records that exist and make sense. An MX record. The quality of the name servers A prospect with no control of its DNS cannot authorize the ESP (Customer Domain Authentication). DMARC at enforcement on a domain the prospect claims not to control is a sign of impersonation
Web presence An HTTP fetch of the domain: a real site, a parked page or a template. Whether the content category matches the stated business A parked or empty domain that claims to be an established business is a classic snowshoe pattern (practitioner consensus)
Prior-IP reputation Blocklist lookups, and SNDS or other reputation lookups, on the IP addresses the prospect says it sent from Automates the vetting-questionnaire question "which IPs did you previously mail from"
Corporate registry / WHOIS transparency Registrant data from RDAP (heavily redacted since the General Data Protection Regulation, GDPR), business-registry APIs Automates the verification of the business. Expect WHOIS to return little and fall back to registry or manual checks. See the 2011 caveat in Customer Vetting

Payment-fraud signals as spam predictors

Signups with stolen cards and spam operations overlap heavily. Spammers do not pay with their own cards, and serial rebillers keep changing their payment identities (the Hosting Abuse BCP's "frequent changes of payment information" trigger). If the ESP takes payment at signup, the payment processor's fraud tools provide a free vetting signal:

  • Stripe Radar (an example of the class) documents a machine-learning risk score from 0–99 for each payment, with the risk levels normal, elevated and highest (defaults in Radar for Fraud Teams: elevated at 65 or above, highest at 75 or above). It also provides "risk insights" that explain the contributing factors, and it is trained across Stripe's network. Adyen, Braintree and Kount publish equivalent scores.
  • Signals worth feeding into vetting: the payment risk score or level, a mismatch between the card's bank identification number (BIN) country and the claimed business country, BINs of prepaid or virtual cards, the velocity of card testing, and later chargebacks (a chargeback on an active sending account is a strong trigger to terminate and review).
  • That payment fraud predicts spam is practitioner consensus, not the finding of a published study. It is, however, built into the registration-abuse guidance of the M³AAWG Hosting Abuse BCP, which treats stolen or disposable cards as the signature of abuse through paid accounts.

Breach and abuse-history lookups

  • Have I Been Pwned (HIBP) documents a v3 API: breaches, pastes and stealer logs by email address, and domain search for domains you own, in subscription tiers with API keys. Its relevance to vetting is indirect. A signup email address that appears in recent stealer logs raises the risk of account takeover (require stronger authentication), and it is a standard input to fraud tools. Paradoxically, an email address that has never appeared in any breach can indicate a newly created identity. Several fraud vendors (IPQS, minFraud first_seen) use this "digital footprint age". The interpretation is industry practice, not documented fact.
  • Internal deny history: the most valuable abuse-history lookup is your own. The Hosting Abuse BCP says to "keep records of previously terminated fraud accounts and match new signups against them." Match on email address, domain, payment fingerprint, device fingerprint and, with care, IP range. No vendor sells this, so build it. Abuse-desk practitioners report that terminated customers signing up again under a fresh identity is the single most common way vetting fails.

Automation architecture

Composing an onboarding risk score

This is the pattern used across the industry. It is inferred from how platforms behave and from the conventions of fraud tools, because no ESP has published a scoring formula:

  1. Collect signals at signup (synchronously, within a budget of about 1–2 s): the IP and proxy score, the disposable check, domain age with ZRD and DBL status, payment risk if applicable, and any match against the internal deny history.
  2. Enrich asynchronously (within minutes): the web-presence fetch, prior-IP reputation, breach lookups, and validation of a list sample once a list is uploaded.
  3. Combine the signals into a single score plus hard-fail overrides. Two rules for combining them matter more than the weighting math:
    • Hard fails bypass the score: a match against the internal deny list, a sending domain listed on DBL (with codes other than abused-legitimate), or a confirmed stolen card. No weighted average should wash these out.
    • Weight by evidence that a signal predicts abuse, then recalibrate: start with the thresholds vendors recommend, log every signal for every signup, and once you have enough outcomes (terminations, breaches of complaint-rate limits, chargebacks), re-weight the signals against your own labels of abuse. A score you cannot test against past outcomes is only for show. Do not adopt anyone else's numeric thresholds, including the vendor examples quoted above, without validating them locally.
  4. Record the full snapshot of signals with the decision. You need it for appeals, for recalibration, and for the abuse desk when the account later misbehaves.

Gate design: four bands

This band structure is standard practice. The bands are consensus, and the boundaries are yours to calibrate:

Band Action Design notes
Auto-approve Full provisioning at the entry rung of the trust ladder (never unlimited) Auto-approval does not mean the account goes unmonitored. Every approved account still enters outbound monitoring and the ladder below
Sandboxed / limited send The account is active but capped: low daily volume, sending only to verified recipients or to the customer's own domain, a shared low-tier IP pool, and restricted API scopes The SES sandbox is the standard documented example (below). This band routes risk into a quarantine pool, as in Spamhaus's guidance on pooling by risk and in Multi-Tenant Architecture
Manual review The account waits in a queue for a person, with the full snapshot of signals. It stays inactive or sandboxed while the review is pending The vetting questionnaire from Customer Vetting is the review script. Review time matters: review queues that take days lose legitimate customers to competitors
Reject Decline service For rejections based on a fraud score, follow the abuse-desk rule: no data retrieval, and no explanation of which signal fired (Abuse Desk Operations). For rejections based on policy (a prohibited content category), state the policy plainly

Two rules shape the structure. First, the reject band should be narrow and made up mostly of hard fails. Weighted scores near the boundary belong in manual review, because rejecting a good customer costs more than sandboxing a bad one. Second, band thresholds must be adjustable at runtime without a deploy, because waves of attacks require tightening within hours (compare the platform-wide responses in Subscription Bombing).

Progressive trust ladders

This is the M³AAWG principle of "tiered rights allocation", automated: initial caps are lifted after the account shows a clean history, not on request alone. Major platforms document these examples:

Platform Documented mechanism
AWS SES Every new account starts in the sandbox: a maximum of 200 messages in 24 h and 1 message per second, with recipients restricted to verified addresses or domains. Production access requires a request, reviewed by a person, that describes the use case and mail practices (typically 24–48 h). A typical initial production quota is 50,000 a day, later raised automatically or on request based on sending metrics. Quotas apply per region, and trust does not transfer between regions.
Mailgun New or flagged accounts are put on probation: an hourly sending limit (commonly cited as 100 messages an hour per domain) until the customer completes business verification, a review of identity and use case by a person.
Postmark Every new account is reviewed manually. Until it is approved, it can send only to its own verified domains. Review typically takes under 24 h on weekdays. After approval there is no fixed daily cap, and enforcement is based on metrics (complaint rate below 0.1%, bounce rate below 10%).
Mailchimp Omnivore Gating on data rather than on volume: on every import, or first send to new addresses, the Omnivore system predicts the likelihood of traps, bounces and complaints for that group of addresses. It blocks sending to audiences whose predicted rates are too high. In other words, the list itself must earn trust.

These examples teach three lessons. (a) The entry rung restricts recipients and rate, not only volume. Allowing sending only to the customer's own domain (Postmark) is a cheap and strong fraud filter, because spammers need to reach strangers. (b) Promotion combines automated metric checks with a person for the big step, from sandbox to production. (c) Ladders also go down: Mailgun puts established accounts back on probation, and SES pauses tenants based on findings (Multi-Tenant Architecture). Each rung above entry (larger caps, dedicated IP addresses, wider API scopes, DKIM for any domain) should have its own metric prerequisites. The Hosting Abuse BCP's bar of about 12 months of clean history for elevated privileges is the only published tenure figure on this topic.

Continuous re-scoring

Vetting at onboarding is a snapshot, so the risk score must keep changing. Feed it from outbound monitoring:

  • Changes in behavior trigger a new score: the M³AAWG triggers for ongoing monitoring (Customer Vetting), such as sudden list growth, changes to content or privacy policy after metrics shift, repeated changes of payment information and sending that stops and starts, together with bounce, complaint and trap telemetry and findings on each tenant's reputation.
  • Query vendors again when a trigger fires, not on a schedule: re-running paid fraud lookups on every customer every night is expensive and finds little. Query again when a behavioral trigger fires. A new sending domain triggers the domain-intelligence checks, a new card triggers the payment signals, and an unusual login triggers the IP and device signals.
  • A changed score goes back through the same gate bands: a worse score moves an account down the trust ladder (lower caps, the quarantine pool, a strict reputation policy) or into manual review. Reusing the onboarding machinery for enforcement over the account's life keeps one set of policies instead of two.
  • Route by past behavior: before enforcement, run the comparison of compromised and malicious accounts from Compromised Accounts. A history of good behavior calls for the account-recovery process, not the fraud process.

Integration patterns

  • Synchronous gate at signup: use only signals that answer within the page-load budget, and run everything else asynchronously. Fail open into the sandbox band, never into auto-approve. A vendor outage should mean that more accounts start limited, not that vetting is switched off. (Practitioner consensus.)
  • Asynchronous enrichment through a queue and webhooks: vendors with event models (Sift's event API, Stripe Radar's webhook events, HIBP polling) push changes in score, and your consumer maps them onto the customer's risk record. Idempotency and handling events that arrive out of order are the usual webhook disciplines.
  • Batch re-validation: list-validation vendors run batch APIs for files (submit a list, poll, fetch the results). Route list uploads through this path, with sending blocked or limited to a sample until the results return, as Omnivore does.
  • Cache with lifetimes that match how fast each signal changes: domain age never changes, blocklist status changes hourly (Spamhaus DQS answers propagate in minutes), and IP fraud scores go stale within days. Do not cache hard-fail signals for longer than they stay valid.
  • One internal schema for the risk record: normalize every vendor's score into your own record (signal, raw value, normalized contribution, timestamp, vendor), so that you can swap vendors and keep a uniform trail for appeals and audits. Lock-in through the meaning of scores is real: a "90" means something different in every API above.
  • Contract and terms-of-service check: most fraud and validation vendors restrict resale and retention of their data. Keep vendors' raw data out of anything the customer sees, including appeal messages, which also avoids teaching fraudsters which signals you use.

False positives and appeal paths

Every automated gate wrongly rejects some legitimate customers: users on shared IP addresses or VPNs, new businesses with new domains, nonprofits with lists that are old but consented. Design for it:

  • An appeal is an escalation to manual review, with the snapshot of signals attached, handled by staff who can override any automated decision. This is by far the most common practitioner pattern, and there is no published standard.
  • Tell the customer what to provide, not which signal fired: ask for the evidence from the vetting questionnaire (business registration, consent records, past metrics) rather than disclosing which signal tripped. Disclosure teaches evasion, the same logic as the vague rate-limit pages in Compromised Accounts.
  • Measure the false-positive rate: track the number of appeals, the share of appeals that overturn the decision, and how accounts behave after an overturn (overturned accounts that stay clean are your ground truth for false positives). Feed errors in both directions into recalibration.
  • Put a time limit on the sandbox: a legitimate customer who stays on the entry rung with clean metrics and is never promoted is a false positive that no one notices. Alert on accounts that stall on the ladder despite a clean history.

Where human review must remain

Automation shrinks the queue for human review. It does not empty it. Keep people on:

  • Edge cases and scores near a boundary: anything the score cannot place with confidence. The score's job is to make the human queue small and well briefed, not to eliminate it.
  • High-value and high-volume accounts: a prospect that wants to send millions of messages a month deserves the full M³AAWG questionnaire and a conversation about compliance, however clean its automated score. The damage a wrong auto-approval can do grows with volume (Multi-Tenant Architecture: aggregate liability is never delegated).
  • Regulated and sensitive content categories: financial services, health, political, adult, gambling and CBD. These need judgment on category policy and on the rules of each jurisdiction (Compliance), which no fraud score encodes. Automated classification of content categories can route an account to the right reviewer, but it should not decide.
  • Promotion from sandbox to production at meaningful volume: every documented platform above puts a person on this step.
  • Appeals and terminations: the highest-stakes decisions, in both directions.
  • Recalibration itself: someone must own the score. That person reviews the weights against outcomes, watches for drift and for demographic skew in false positives, and decides when a vendor signal is no longer worth its cost.

Check your own record

The free check reads what your domain publishes in DNS.

In this topic

All 16 in ESP Operations →