emailmarketing.net

Shared IP Pool Recovery — Disaster Runbook for a Poisoned Pool

The disaster-recovery runbook for a shared IP pool torched by a bad tenant — distinguishing a poisoned pool from one struggling sender, attributing and isolating the offender, containing them, delisting the pool's IPs across blocklist zones, evacuating clean tenants to a healthy pool, and deciding whether to rehabilitate or retire the burned IPs.

Operational14 min read

Who it is for ESP operators

If one tenant's behavior has pushed the IP addresses of a shared pool onto blocklists, into spam folders or under provider throttling, every clean tenant on that pool pays for it. This runbook covers that situation: one customer, or a few, have damaged a shared pool, and the innocent tenants who share it are now losing inbox placement.

The runbook puts existing procedures in order and adds only the coordination a shared-pool disaster needs. Each procedure it relies on is described in full elsewhere, and is linked at the end of this page.

How this differs from recovering a single account: Reputation Incident Recovery assumes that a single account owns the damaged asset and can pause everything while it rebuilds. A shared pool cannot do that. You must remove the offender without pausing the clean tenants, whose only fault was sharing IP addresses. That constraint shapes the whole procedure: isolate the guilty, protect the innocent, and decide the fate of the IP addresses last.

The recovery sequence at a glance

The order matters, and each step assumes that the previous one is done. If you delist before containment, the IP addresses are simply listed again. If you move clean tenants before delisting, they leave the damaged IP addresses, but the pool is still not fixed.

# Phase Goal Main reference
1 Detect & classify Confirm that the pool is poisoned, not that one sender is struggling Outbound Monitoring
2 Attribute Identify which tenant or tenants caused it Spam-Trap Incident Response
3 Contain Stop the damage: throttle, pause or evict the offender, and purge their queue Account Enforcement, Abuse Desk
4 Protect clean tenants Move the innocent tenants to a healthy pool (at the same time as 5) Multi-Tenant Architecture
5 Delist Clear the pool's IP addresses in every zone, after fixing the root cause Spamhaus Listings Deep Dive
6 Rehabilitate or retire Warm the cleaned IP addresses up again, or replace them Reputation Incident Recovery

Once attribution is firm, steps 3 and 4–5 run in parallel: containment stops new damage while the move and delisting go ahead. Do not delist while the offender is still sending bad mail into the pool.

Phase 1: detect and classify a poisoned pool or one struggling sender

The first decision is a diagnosis, and a wrong answer is expensive. If you treat a poisoned pool as one sender's problem, the damage to the other tenants continues. If you treat one sender's ordinary bad week as a pool disaster, you move tenants for no reason. Outbound Monitoring describes the monitoring that produces the signals. The question here is how to read them to tell a poisoned pool from a single sender.

What tells them apart is scope. Does the damage follow the pool's IP addresses, whichever tenant sent the mail? Or does it follow one tenant, whichever IP address it used?

Signal Points to a poisoned pool Points to one struggling sender
Blocklisting A shared-pool IP address is listed, and the evidence for the listing spans several tenants' mail Only a dedicated IP address is listed, or the listing traces back to one tenant's domains or content
Collapse in placement or engagement Every tenant on the pool sees the same provider get worse at the same time One tenant gets worse while the other tenants on the same IP addresses hold steady
Spike in deferrals or rejections Rejections at one provider rise for all senders on the pool Rejections are concentrated in one tenant's traffic
Complaint or trap rate The pool's overall rate breaches the threshold, but no single tenant dominates it One tenant's rate is an outlier, and the pool's overall rate is fine once that tenant is removed
DMARC or authentication failures The pool's shared sending domain is being flagged A customer's own delegated domain is failing

The essential tool is scoping by pool, with every alert attributed to a customer. An alert on a shared asset is only half an alert until it is mapped to the tenant or tenants behind it (see Outbound Monitoring § Scoping).

A practical test: remove the traffic of the tenant you suspect most from the pool's totals, and calculate them again. If the pool's metrics fall back within the threshold, you have one offender, which is a containment problem. If they stay red, the pool is broadly contaminated. That means moving every tenant, possibly several offenders, and a harder root cause. Watch shared assets most closely. A listing of a shared-pool IP address punishes every tenant on it, which is why blocklist and trap monitors alert at platform severity for shared assets.

Is this really a disaster? Use this runbook, rather than the routine enforcement loop, when any of the following is true: a shared-pool IP address gets a blocking listing of the SBL or CSS class (not a merely informational one); a major provider (Gmail, Microsoft or Yahoo) starts placing the whole pool in spam or rejecting it outright; or two or more clean tenants report a loss of inbox placement they did not cause. Below that level, this is ordinary enforcement against one tenant, not pool recovery.

Phase 2: attribution, finding the offender or offenders

Both containment and delisting require knowing who caused the problem. You cannot cite an acceptable use policy (AUP) clause against an unknown tenant, and Spamhaus will ask how the problem was solved before it delists (see below). Spam-Trap Incident Response § Detection and attribution covers attribution in full for poisoning driven by spam traps. The tagging on each send that makes any shared-asset signal usable (tenant IDs, campaign IDs, a DKIM d= for each customer, headers in the style of Feedback-ID) is also what maps the poisoning to a tenant.

Questions specific to shared pools to answer before you act:

  • One offender or several? The test from Phase 1, removing a tenant and calculating again, answers this. A pool contaminated by many marginal tenants shows a failure in how the pools are tiered, not a single bad actor. The fix is to reorganize the pool's tiers (see Avoiding Blocklistings § Risk-tiered IP segmentation), not to evict one customer.
  • Compromise or malice? A tenant with a clean history whose metrics suddenly spike is probably a compromised account, not a spammer. The symptoms in the pool are the same, but the remediation is completely different, and so is the tone of your contact with the customer. A good history is the sign. Settle this before you send an eviction notice.
  • Which IP addresses carried the bad mail? In a pool, the offender's traffic may have used only some of the pool's IP addresses. Map the bad sends to specific IP addresses. That limits the delisting work in Phase 5 and informs the decision to retire or rehabilitate in Phase 6.

Keep spam traps confidential throughout. Blocklist and reputation vendors report trap-hit metrics without revealing which addresses are traps, and your notice to the customer must keep that confidentiality (see Spam-Trap Incident Response).

Phase 3: containment, stopping the damage

Once attribution is firm, cut the offender off from the pool immediately. This applies the response ladder in Outbound Monitoring and the state machine in Account Enforcement under the pressure of a disaster. The difference from routine enforcement is that every hour the offender keeps sending gets the IP addresses listed again just as you try to delist them. Containment is therefore not negotiable, and it comes before any delisting request.

  1. Throttle the offending tenant to zero, or pause it, at the level of the customer or tenant, with a person approving any resumption. For a confirmed poisoning, go straight to a pause and skip the gentle throttle used when the case is unclear. Fraudulent or clearly malicious accounts get none of the courtesy steps (Abuse Desk § Remediation workflow). A suspected compromise is paused and also secured (credentials rotated, API keys revoked) rather than terminated.
  2. Purge the offender's queued and deferred mail. Deferred bad mail keeps retrying for up to 72 hours after you "stopped" it, and it keeps generating complaints, trap hits and new listings on the pool's IP addresses. Draining the queue is part of stopping the damage, not a later cleanup (see the queue purge in Reputation Incident Recovery § Stop the bleeding).
  3. Notify the customer and cite the specific clause of the AUP or terms of service. This keeps the customer agreement intact and protects the platform against a dispute from either side (Account Enforcement, Abuse Desk). The eviction decision itself (suspension until the problem is fixed, or termination) follows the enforcement ladder. A termination holds up in a dispute only with documented communication, several attempts to intervene, and agreement among stakeholders (Avoiding Blocklistings § Enforcement of last resort).
  4. Rebuild suppressions from the period of the incident. The pool must honor every bounce, complaint and unsubscribe the offender generated from now on. Merge them in before any tenant resumes sending at volume (Reputation Incident Recovery § Rebuild the data layer, Suppression-List Architecture).

Containment is complete when no bad mail is entering or leaving the pool. Only then will delisting last.

Phase 4: protecting the clean tenants by moving them to a healthy pool

The innocent tenants cannot wait on damaged IP addresses for delisting to finish. Every day on a listed shared IP address costs them inbox placement they did nothing to lose. Move them to a healthy pool while the delisting work goes on. Multi-Tenant Architecture describes the building blocks for routing and isolation. The move changes a routing binding. It does not rebuild the tenants.

How to move tenants:

  • Change the routing binding, not the tenant. A tenant sends through whichever pool its routing binding names (a configuration set, or ip_pool_name). Move clean tenants by binding them to a healthy pool. Their identities, suppressions and metrics are logical, so they stay where they are, and only the physical routing to IP addresses changes (see the split between logical tenants and physical pools in Multi-Tenant Architecture § The three isolation layers).
  • The destination pool needs spare capacity and the right warm-up. Moving tenants onto a pool that is already near capacity, or that is cold for the providers they send to, only moves the problem. Warm-up is tracked for each IP address at each receiving mailbox provider, and it fades. A healthy pool that has not carried Gmail volume is not warm for Gmail (see IP Warm-Up and Multi-Tenant Architecture § Managed pools). If no warm pool has spare capacity, use the shared pool's overflow mechanism to absorb the shock: spread the moved tenants' volume across warmed capacity and ramp up the rest, rather than sending large volumes from cold IP addresses.
  • Move tenants in order of exposure. Move the clean tenants with the highest volume and the most sensitivity to placement first, because they lose the most for each hour on listed IP addresses. Tenants with low volume can wait a short time.
  • Do not move tenants onto the pool of another risky tenant. Keep the risk tiers: clean, established senders go to a clean, established pool, not to whatever pool has space. Contamination then stays within a tier (Avoiding Blocklistings § Risk-tiered IP segmentation).
  • Update authentication if the sending domain changes. If the moved tenants now use IP addresses under a different shared sending domain, verify SPF, DKIM and DMARC alignment again before they resume. Customers break their own DNS all the time, and the monitor of customer authentication catches it (see Customer Domain Authentication).

On a platform with well-organized tiers, this is cheap: the healthy pools already exist, and moving tenants is a routing change. A platform with a single undivided shared pool has nowhere to move tenants to. That is the argument for setting up tiers before an incident, not during one.

Phase 5: delisting the pool's IP addresses across the blocklist zones

With the offender contained and the queue purged, work on delisting. Spamhaus Listings Deep Dive and Blocklists & Spamhaus give the criteria for each zone, the return codes, and when removal is self-service or needs an investigation. What changes for a shared pool:

  • Fix the root cause first, in a way you can prove. Delisting a shared-pool IP address while the cause persists uses up your limited self-removals, and for the SBL it invites a wider listing. The abuse desk must be able to describe how the problem was solved. For a tenant that was spamming, Spamhaus expects the account to be truly removed, not just paused (Spamhaus Listings Deep Dive § SBL delisting). This is why Phase 3 comes first.
  • Check every listed IP address in every relevant zone. List each shared-pool IP address that carried bad mail (from Phase 2), and check each one against the zones that apply. Match the workflow to the list:
Listing class Removal path Note for shared pools
CSS (127.0.0.3, automated, email with low reputation) Fix the cause, then remove the listing yourself through the checker, or wait ~72 h for it to expire The list most often hit when a careless tenant poisons a pool. Self-removals are limited, so do not use them up before containment is real
SBL (127.0.0.2, manual) The abuse desk of the responsible network writes to the SBL Removals Team, explaining the fix This is where escalation happens: tolerating the offender can widen the listing from one IP address to your whole allocation, which is the strongest case for eviction
DBL (127.0.1.x, domains) Expires on its own once the criteria stop matching, or through the checker form Relevant when the poisoning involved the shared sending or click-tracking domain, not just IP addresses
Internal to a provider (reputation at Gmail, Microsoft or Yahoo) Escalate to postmaster or support with a description of the incident and the remediation Not a DNSBL. It clears on its own timeline as clean volume rebuilds reputation. See Escalation & Mitigation Channels
  • Listings caused by a compromise clear quickly, and listings caused by a spammer do not. If the poisoning came from an account compromise that is documented and fixed, operators and providers routinely reverse listings quickly once you show that the pool is secure. Poisoning by a deliberate spammer is harder and slower, and the delisting request must show that the customer was actually removed.
  • Delisting does not unblock you instantly. A removal from a zone reaches each receiver on its refresh schedule (minutes for DQS and ZEN subscribers, up to 24 h for slower ones), and reputation damage at the receiver fades separately on top of that. Do not tell moved or returning tenants that the pool is "clean" the moment the zone clears.
  • Always have a person handle delisting requests. Automated delisting requests that are mishandled or repeated can make a listing permanent (Outbound Monitoring § Escalate-to-human criteria).

Phase 6: rehabilitating or retiring the poisoned IP addresses

Once the offender is gone, the clean tenants have moved and the zones are clearing, decide what to do with the damaged IP addresses. Reputation Incident Recovery § Replace or rehabilitate? sets out the general analysis: new assets start with a negative reputation rather than a neutral one, reputation follows the mail and not only the asset, and listings caused by a compromise clear quickly. The default is still to rehabilitate. These factors, specific to shared pools, push the decision one way or the other:

Factor Favors rehabilitating the pool's IP addresses Favors retiring or replacing them
Cause A single offender, contained and cleanly evicted, or a compromise that is documented and fixed Contamination spread across many tenants (a failure of tiering), or repeated poisoning of the same pool
Listing status Can be delisted, and operators respond to the evidence of remediation A permanent or repeated SBL listing after the offender was tolerated too long, or IP addresses on lists with no practical way to be removed
Length of history A short incident on otherwise clean IP addresses A long history of abuse, or IP addresses the offender had been quietly damaging for months
Cost Warming up again costs less than acquiring and warming new IP addresses from zero Hard blocks at major providers continue despite a clean warm-up over several weeks

If you rehabilitate the pool: treat the cleaned IP addresses as a second warm-up, not a cold start. Receivers hold concrete bad history on them, so ramp up only traffic to recently engaged recipients, hold or step back at any provider that is still deferring, and check blocklists again before each increase in volume. Reputation Incident Recovery § Re-warming vs cold warm-up sets out the differences between the two in a table. Note the rule on dormancy: an IP address left idle for over 30 days during the incident and recovery needs a new warm-up, whatever its previous status. A damaged pool kept out of rotation for weeks creates its own need for a new warm-up.

If you retire the pool: if the pool cannot be rehabilitated, release the damaged IP addresses cleanly and set up a new pool warmed from zero. Once it is warm, move the tenants (who have already left the damaged pool) onto it. Retire the damaged IP addresses rather than quietly giving them to new tenants: new senders who inherit an IP address with a recent history of abuse lose reputation faster than the pool recovered. When you acquire replacement IP space, check its reputation history first (IP Acquisition Diligence). A new pool built on space with an inherited bad reputation is not a fresh start.

Either way, the offender does not return to shared infrastructure. A tenant able to poison a shared pool belongs on a dedicated IP address (where the reputation is theirs alone), on a stricter risk tier, or off the platform. It never goes back into the pool it damaged.

Check your own record

The free check reads what your domain publishes in DNS.

In this topic

All 16 in ESP Operations →