emailmarketing.net

ESP Outbound Monitoring Systems

The continuous-monitoring architecture of a sending platform — pre-send scanning, send-time anomaly detection against per-customer baselines, post-send feedback watchers, alert design (leading vs lagging, threshold vs trend, per-pool vs per-customer), and the graduated automated-response ladder.

Operational16 min read

Who it is for ESP operators

If you run an email service provider (ESP), you need to watch your own outbound traffic continuously, so that a customer's bad day is caught before it damages everyone on your platform. A monitoring system does this in layers. It compares each customer against their own baseline, runs watchers across the whole platform, and turns signals into throttles, pauses and escalations to people.

What happens after a specific incident is detected is covered separately: Abuse Desk Operations for handling reports and the remediation loop, Compromised Accounts for account takeover, Spam-Trap Incident Response for trap hits, and Reputation Incident Recovery for the rebuild. The monitoring described here is the detection that runs all the time and feeds those procedures.

An ESP's deliverability is set by its worst customers on their worst days, and "deliverability failures are unpredictable — they happen to different customers at different times for different reasons." Word to the Wise describes how well-run ESPs organize around this. An early-intervention desk catches customers who show warning signs before the damage lands. An enforcement desk handles customers who ignore guidance. The compliance department tracks metrics internally to stack-rank clients, rather than guessing who the good senders are. Everything below is the instrumentation those desks rely on.

Authoritative published documentation of monitoring inside ESPs is thin, because vendors treat their detection logic as sensitive. The mechanisms below are documented, from the Messaging, Malware and Mobile Anti-Abuse Working Group (M3AAWG) Best Common Practices (BCPs) and from vendor documentation, unless they are marked as practitioner consensus.

The three layers of the monitoring stack

Layer When it acts Latency Signals Can it prevent harm?
Pre-send scanning Before or at injection, before any mail leaves Milliseconds to seconds Content, URLs, list composition Yes. It is the only layer that stops the first message
Send-time anomaly detection During the send, against behavioral baselines Minutes Volume, velocity, mix, early bounces Partially. It stops a send while it is under way
Post-send feedback watchers After delivery, from external signals Hours to days Complaints, blocklists, trap hits, engagement, DMARC reports No. It limits recurrence, not the send that triggered it

Each layer trades coverage for latency. Pre-send scanning sees everything but knows the least, because no recipient has reacted yet. Post-send feedback knows the most but arrives after the damage. A platform needs all three because each catches what the others cannot. A send with clean content to a purchased list passes layer 1 and is caught only by layers 2–3. A phishing template caught by layer 1 never produces the layer-3 evidence.

Layer 1: pre-send scanning

The platform inspects customer mail before it leaves the network:

  • Content scanning against spam and phishing signatures. Run outgoing customer mail through the same kind of engine that receivers use. Practitioner consensus describes "pattern matching against known spam templates, URL reputation checks on every link", plus comparison against known phishing and credential-harvesting templates, which pauses sending before the campaign completes. Commercial and open engines (SpamAssassin, rspamd, Cloudmark, Vade) can score mail at injection; Halon documents these as anti-spam engines that plug into the mail transfer agent (MTA). Milter hooks are the classic integration point. The M3AAWG Web Messaging BCP sets out the options: heuristics, a spam-filtering engine in the submission path, reputation models, and batch clustering. They are summarized in Compromised Accounts, content filtering.
  • URL and link domain checks. Check link domains against URL reputation lists before sending. A customer whose template links to a domain listed on SURBL or the Spamhaus Domain Blocklist (DBL) will be filtered at receivers whatever the sending reputation, so catching it before the send protects the shared pool. URL Reputation and Content Fingerprinting explains how receivers extract and judge URLs. Spamhaus data such as the Hash Blocklist (HBL) and the DBL can be used on outbound mail as well as inbound.
  • Template fingerprinting across tenants. Apply fuzzy hashing to customer sends and cluster near-duplicates. The same template injected by many unrelated new accounts is the mark of scripted signup abuse. The Web Messaging BCP gives an example from batch analysis: "accounts all registered June 7 from one /16 sending 103-word messages". This is the platform's counterpart to the fingerprinting receivers do, described in URL and Content Filtering.
  • Screening list imports. The pre-send check with the most effect is on the recipient data, not the message. Scan imports for signs of purchased lists (column headers such as "jigsaw" or "append"), role accounts, addresses known to be bad or associated with traps, and invalid syntax or MX records. This is the same review Customer Vetting prescribes at onboarding, applied to every import. The shape of an import also feeds layer 2 (below). Its size relative to the customer's existing base, its domain distribution and its validation failure rate are the earliest predictors of the send that follows.
  • A missing unsubscribe link. Check every marketing or list message for an unsubscribe link before it leaves. The M3AAWG Sender Best Common Practices (Version 4.0, August 2026, section 2.7) lists a missing unsubscribe link among the account activity an ESP should look for. It gives two reasons: compliance with laws such as CAN-SPAM, the Canadian anti-spam legislation (CASL) and the ePrivacy Directive, and avoiding deliverability and legal problems. For the header form of the link, see List-Unsubscribe.

Pre-send scanning must be fast enough not to hold up mail (practitioner consensus). Blocking every send on a slow verdict destroys transactional latency. The usual pattern runs cheap checks inline (URL lists, signature matches), while expensive analysis (clustering, machine learning scores) runs seconds to minutes behind and can stop the rest of a campaign.

Layer 2: per-customer behavioral baselines and deviation signatures

The core send-time mechanism is a rolling behavioral baseline for each customer, and for each tenant or subaccount (see Multi-Tenant Architecture for where the metrics attach). It records typical daily and hourly volume, sending cadence, the mix of recipient domains, and normal bounce, complaint and engagement levels. The signal is a deviation from that customer's own history.

Absolute thresholds applied to the whole platform miss both the 500/day sender who suddenly sends 200,000 in an hour and the 10M/day sender whose 2% shift is a real incident. Research on subscription bombing measured why: anomaly detection on aggregates drowned in noise, while detection for each entity (volume ≥10 standard deviations above the entity's 7-day mean) worked. See Subscription Bombing for the study.

M3AAWG publishes one volume figure of its own. Its Sender Best Common Practices (Version 4.0, August 2026, section 2.7) says a customer who sends more than 40% above their usual volume can be flagged, even when no irregular list import explains it. The cause may be a malicious actor or a mistake by the customer, and a customer who made a mistake should be educated so it does not happen again. The same section encourages ESPs to combine automated heuristics with review by a person, so treat the 40% figure as a trigger for review, not for an automatic pause. It also lists an unprecedented jump in list growth as a signal, to be checked by validating the added addresses.

What each deviation most likely means:

Deviation Main hypotheses (roughly in order of likelihood) First check
Volume spike (for example "normally 500/day, suddenly 200,000 in an hour", the classic trigger for an automatic pause) A compromised account or API key; a new list import; a legitimate seasonal event or campaign The authentication trail (new IP addresses or locations, a new API key), the import history, contact with the customer
Shift in recipient domain mix (for example a B2B sender suddenly at 80% consumer domains, or one unusual domain over-represented) A new or purchased list; harvested data; a concentration of spam-trap domains; a compromise that reuses stolen recipient data Compare the mix with the baseline; correlate with imports; check the engagement of the over-represented domain. A domain with near-zero engagement concentrated in one segment is a sign of a trap network (see Spam-Trap Incident Response)
Hard bounce velocity (an unknown-user rate far above the customer's norm, visible within minutes of the send starting) An old or stale list sent again; a purchased or appended list; a list built by dictionary attack The mix of bounce categories (Metrics and Benchmarks); whether bounces cluster in one import or segment
Complaint velocity (feedback loop (FBL) reports arriving faster than the customer's norm) A permission failure on a new segment; a change in content or frequency; a compromise Complaint attribution for each campaign; which providers are reporting
Engagement collapse (opens and clicks fall at one provider while volume holds) Mail going to the spam folder at that provider, which is a lagging placement signal. Measurement caveats apply (Non-Human Interactions, Tracking Distortion) The engagement trend for each provider; postmaster dashboards
New list import signature (a large import followed immediately by a send to the whole list) The single most common trigger of trap hits, bounce spikes and blocklistings The import size compared with the existing base; the validation failure rate of the import; hold the first send to the imported segment for review (practitioner consensus)
Sending that stops and starts; frequent changes of contact or payment details; changes to content or privacy policy after metrics shift Spreading reputation damage across several ESPs; resale of the account; evasive behavior The post-send triggers documented in the M3AAWG Vetting BCP (Customer Vetting, ongoing monitoring)
Login and session anomalies (unfamiliar locations, impossible travel, new user agent strings, mass deletion of sent mail) Account compromise Compromised Accounts, detection covers the response. The baseline system only raises the alarm

How to keep baselines (practitioner consensus, consistent with the per-user statistics in the Anti-Phishing Working Group (APWG) study): keep at least a 7-day rolling mean and variance for each metric of each customer, and a longer (30-day) window for seasonal senders. New customers have no baseline, which is exactly why tiered rights for new accounts, with low caps raised as history builds, stand in for one.

The telemetry baseline in the M3AAWG Hosting Abuse BCP is the minimum set of events to collect: registrations, logins, password changes, actions completed, and identities that reach any limit. Keep this data for a week or longer to catch abuse that is slow and low in volume (summarized in Compromised Accounts, platform telemetry).

Layer 3: platform-level watchers

These monitors watch external signals continuously across all of the platform's IP addresses and domains, independently of any one customer:

Watcher What it watches Feeds and mechanics In depth
Blocklist monitor Every sending IP address on the platform, and every sending, tracking and link domain, on a fixed polling cycle plus parsing of SMTP rejections Automated DNSBL lookups. Parsing rejection text in delivery logs names the list before polling does Reputation Monitoring, Blocklists and Spamhaus
Trap hit alerts Trap hit metrics from blocklist and reputation vendors and from SNDS, attributed to a customer or segment Vendors report hits without revealing which addresses are traps. Microsoft's Smart Network Data Services (SNDS) shows trap hits for each IP address Spam-Trap Incident Response, Microsoft SNDS
FBL complaint streams Abuse Reporting Format (ARF) reports from every available provider FBL, parsed and attributed to a customer and campaign. They drive both suppression and complaint velocity metrics Intake shared with the abuse desk; routing by source Complaint Feedback Loops, Abuse Desk
Queue health and deferral dashboards The distribution of queue age for each destination, the deferral rate, and connection failures. Throttling at a provider is the step before blocking, and it shows here first The qshape queue-age model; automated traffic shaping driven by bounces reacts to the same signals MTA Delivery Tuning, Per-Provider Tuning Baselines
DMARC report monitoring Aggregate (RUA) reports for the domains the platform operates: unexpected sources using platform domains, authentication failure rates, and spoofing of the ESP's own brand Daily XML aggregate feeds; alert on new failing sources DMARC Aggregate Reports, DMARC Deployment
Postmaster dashboards The receivers' own view: domain and IP reputation and spam rate in Google Postmaster Tools, and SNDS filter verdicts Set them up before problems occur, because history cannot be rebuilt later Google Postmaster Tools, Microsoft SNDS and JMRP
Customer authentication watcher Whether customers' SPF, DKIM and DMARC records are still valid (customers break their own DNS routinely) Periodic verification of delegated records Customer Domain Authentication

Two points on scope. First, watch shared assets hardest. A listing of a shared pool IP address punishes every tenant on it, so blocklist and trap watchers should page at platform severity for shared assets and at customer severity for dedicated ones.

Second, attribute everything. A platform-level signal (a listing, a trap hit, a spike in deferrals at one provider) becomes actionable only once it is mapped to the customers and campaigns that caused it. That requires tagging every send (tenant IDs, campaign IDs, a DKIM d= for each customer, or headers in the style of Feedback-ID), designed in from the start.

Composite health scoring: a documented example

SparkPost's Signals Health Score is the best-documented public example of combining these feeds into one leading indicator for each sender. The SparkPost support document now redirects to Bird marketing pages, so the details below come from the document as it was previously published.

The score is calculated daily on a scale of 0–100 (≥80 is considered good) and predicts engagement relative to all senders on the platform. It combines twelve components, including:

  • subscriber quality: the share of injections that match address patterns associated with problematic ways of acquiring lists;
  • hard bounce %, block bounce %, complaint % and transient failure %;
  • spam-trap hit %;
  • suppression list hit %: how often the customer tries to mail addresses that are already suppressed, which is a sign of poor list hygiene;
  • 3-day historical engagement;
  • the share of unsubscribes.

The minimum is about 1,000 messages/day with open tracking; below that, no accurate score is possible.

The design lessons are these. Score rates against injections, so that they are normalized for volume. Count attempts to send to suppressed addresses as a negative signal. Refuse to score below a statistical minimum rather than produce noise. The per-tenant reputation findings in Amazon SES (rolling windows of 24h–7d, a minimum representative volume, low or high severity) apply the same idea and connect it directly to enforcement (see Multi-Tenant Architecture).

Alert design

Leading and lagging indicators

Leading (they predict trouble) Lagging (they confirm trouble)
Examples Import shape and validation failure rate; deviations in volume or velocity; shifts in recipient mix; suppression hit rate; subscriber quality patterns; a rising deferral rate at one provider Blocklistings; FBL complaint totals; trap hits reported by vendors; a reputation downgrade in Google Postmaster Tools (GPT); engagement collapse
Latency Seconds to minutes Hours to days (GPT and FBL data arrive after sends; listings follow the behavior that earned them)
Right response An automated throttle or hold, which is cheap to apply and cheap to get wrong Incident procedures and root-cause work, which are expensive and need people

The operational goal is to catch customers on leading indicators so that the lagging ones never fire. Word to the Wise's early-intervention desk is the team that acts on leading indicators; the enforcement desk acts on lagging ones.

Thresholds and trend detection

Use both, for different jobs:

  • Absolute thresholds where the email ecosystem imposes them: ceilings published by providers (Gmail's spam rate thresholds and bounce rate norms, in Metrics and Benchmarks and Gmail Sender Requirements) and platform policy caps (limits on new accounts). These are contractual lines, not statistics.
  • Baseline and trend detection for all behavior. That means deviation for each entity (the rule of ≥10 SD for each entity, validated in the list-bombing research), trends from week to week rather than noise from day to day (Reputation Monitoring), and velocity: complaints per hour into a send, not per calendar month. The SES finding example, "bounce rate exceeded 15.0% based on a representative volume of 664 emails" over about 2 hours, shows the pattern of a rolling sample.
  • A minimum volume for every rate metric. A "20% bounce rate" made of 2 bounces out of 10 sends must not page anyone. SparkPost (a 1,000/day minimum) and SES ("minimum representative volume") both document this guard.

Scoping: per-customer, per-pool, per-destination

Every alert needs an explicit scope, because the same metric means different things at different scopes:

  • For each customer or tenant: behavioral deviations and the velocity of complaints and bounces. These drive a response at customer level: throttle, pause or vet again.
  • For each pool: the aggregate complaint rate, blocklistings and deferral rates on shared pools. These drive isolation decisions, such as moving the offender out, as described in Spam-Trap Incident Response (for how pools work, see Multi-Tenant Architecture). An alert for a pool without attribution to a customer is only half an alert.
  • For each destination: deferral and rejection rates for each receiving provider. A spike in deferrals at Microsoft only, across the whole platform, points to MTA tuning or a problem with shared reputation (MTA Delivery Tuning). The same spike limited to one customer points to that customer's reputation at Microsoft.
  • Across all customers on the platform: signals that no single scope can see. Examples are the same target address signed up through many customers' forms (Subscription Bombing), the same template across unrelated new accounts, and the same range of submitting IP addresses registering across tenants.

Alert hygiene (practitioner consensus):

  • Deduplicate by incident, not by event. One listing is one incident, not one page for each bounced message.
  • Route by severity. Signals about shared assets, traps and blocklists page someone; behavioral deviations of a single customer go to the queue of the intervention desk.
  • Record every alert in the customer's history. Stack-ranking depends on accumulated incident records, and the unblocking practice in Abuse Desk shows why a written history for each sender matters.

Automated responses and the escalation ladder

The documented and consensus response ladder follows the principle that "the response scales with the severity: throttling first, then pausing, then account suspension". In increasing severity:

  1. Throttle: cap the customer's sending rate, or hold them to their baseline volume. This is the cheapest response. It turns a runaway send into a slow one while analysis catches up. The Web Messaging BCP says rate limiting "should be employed on almost all web services accepting or relaying user-generated content". Dynamic blocking is rate limiting with an allowance of zero.
  2. Quarantine or hold: set aside the rest of a campaign (or the first send to a suspicious import) for automated deep analysis or review by a person. This applies the Web Messaging BCP's three responses, challenge, quarantine and reject, at the scale of a campaign. Suppress or reroute the mail rather than silently dropping it, because silent deletion creates support work and hides the problem.
  3. Pause sending, for a customer or a tenant: the automated stop. The tenant reputation policies of SES are the public reference design. Standard pauses on high-severity findings, Strict pauses on any finding, and None only monitors, for supervised onboarding. Sending resumes only when a person decides, followed by a grace period after reinstatement to prevent repeated pauses (Multi-Tenant Architecture).
  4. Isolate: move a problem customer off shared infrastructure onto dedicated IP addresses or pools, so that continued analysis does not cost other tenants (M3AAWG guidance on responding to spam traps).
  5. Suspend or terminate: the abuse desk's remediation loop. Validate, notify the customer with a citation of the terms of service (ToS), allow a remediation window, suspend, then terminate. Fraudulent accounts skip these courtesies (Abuse Desk).

When to escalate to a person. Automation should throttle and pause. People should decide anything that is expensive to get wrong (practitioner consensus, consistent with the repeated warnings in the M3AAWG documents to verify before acting):

  • Unclear verdicts between a compromised and a malicious account. Previous good behavior suggests compromise, and the two responses are completely different (Compromised Accounts). The Compromised User ID BCP requires a person to intervene in all cases of customized exploits.
  • Spikes that may be legitimate: seasonal peaks, breach notifications and recalls (Mandated and Regulatory Email), product launches. The baseline system cannot tell these apart from abuse; the customer's history and a phone call can.
  • Resuming after any automated pause (deliberately left to a person in the SES design).
  • Anything that affects the reputation of shared assets outside the platform: blocklist delisting requests and provider escalations (Escalation and Mitigation Channels). Repeated requests that are badly handled can make listings permanent.
  • Termination decisions, always (Abuse Desk).

The cost of false positives completes the picture. Every automatic pause that wrongly stops a legitimate sender costs support work and customers. Every abuse the automation misses costs reputation on the shared pool. The tiered trust model in Customer Vetting is how you tune the balance. New customers and customers flagged before run under strict automation (low caps, holds on deviation). Long-standing clean customers run under loose automation (alert and review). Sensitivity then follows demonstrated risk instead of being the same for everyone.

Check your own record

The free check reads what your domain publishes in DNS.

In this topic

All 16 in ESP Operations →