Reputation Incident Recovery
The recovery runbook after a compromise or reputation collapse — securing the account, purging queues, rebuilding suppressions, re-warming vs cold warm-up, reading deferrals during recovery, automated warmup schedules as vendors implement them, and when to replace vs rehabilitate an IP or domain.
Operational12 min read
Who it is for ESP operators, Senders
Applies to senders on any platform
ContentsOn this page — 6 sections
When a compromised account or API key has pumped spam through your infrastructure, or a list accident or another event has collapsed the reputation of your IP addresses or domains, recovery takes days or weeks of deliberate work. The runbook below covers the recovery sequence, how re-warming differs from a cold warm-up, and how to read push-back from receivers while you rebuild.
For detection, diagnosis and blocklist delisting, see Reputation Monitoring and Remediation. For building reputation from scratch, see IP Warm-Up.
Post-compromise restoration sequence
The sequence below combines SendGrid's documented guidance on restoring reputation after a compromise with standard incident practice at email service providers (ESPs). Order matters, because every step assumes that the previous one is complete. Re-warming on top of an unsecured account or a dirty queue only burns again the reputation you are trying to rebuild.
1. Stop the bleeding
- Pause all sending for several days. SendGrid's guidance is explicit: stop sending, to prevent further reputation damage, before you attempt anything else. Mail sent while receivers are actively penalizing you makes the damage deeper.
- Secure the account. Rotate every credential and API key, revoke unknown API keys and teammate or subuser access, enable and enforce two-factor authentication (2FA), and check for domains, senders, webhooks or templates the attacker added. Close the route of the compromise before any recovery mail goes out. Otherwise the attacker resumes the moment you do.
- Purge the queues. Drain or delete everything that is still queued or deferred from the period of the incident. Deferred spam that keeps retrying for hours or days after "recovery" goes on generating complaints, trap hits and blocks under your name.
2. Rebuild the data layer
- Rebuild suppression lists. Attackers commonly delete or export suppression lists (bounces, complaints, unsubscribes), and the mail sent during the compromise generated new bounces and complaints that you must honor. Restore suppressions from backup, merge in every bounce and feedback loop (FBL) complaint from the period of the incident, and permanently suppress any address the attacker mailed that was not on your opt-in list.
- Cut the send list down to recently engaged recipients ("remove unengaged recipients from your contact list", in SendGrid's words). Recovery sending uses only your best data, the recipients who recently opened or clicked, exactly as at the start of a warm-up, but with even less margin for error.
3. Re-verify infrastructure
- Confirm authentication end to end: SPF, DKIM and an enforced DMARC record on all sending domains. SendGrid lists setting up DMARC on all sending domains as a core restoration step. It blocks continued spoofing of your domain, and it signals to receivers that you are in control.
- Check again that you meet the providers' requirements for bulk senders (Gmail, Yahoo, Microsoft). An incident is often the moment when enforcement that was previously dormant becomes active.
- Enroll in every monitoring channel before you resume: Google Postmaster Tools, Microsoft SNDS and JMRP, and Yahoo's complaint feedback loop. They give you each provider's verdict, which decides whether you can take each step up in volume.
4. Delist and notify
- Request blocklist delisting by following the delisting workflow, after you have fixed the root cause. A compromise is one of the few causes of listing where operators routinely delist quickly, once you can show that the account is secured.
- If blocks persist at a specific provider after recovery, contact that provider's postmaster or support channel directly, with a description of the incident and of the remediation. SendGrid explicitly recommends asking the support teams of the major providers for help tailored to your case.
5. Gradual re-ramp
- Resume with a manual warm-up process: small volumes to engaged recipients, increased gradually (see the schedules below). On platforms with automated warmup, you may need to re-arm the automation or put the IP back into warmup. An IP that kept its "graduated" warmup status inside the platform has still lost its reputation with receivers outside it.
- Set expectations using SendGrid's stated timeline. Internal reputation measures (the platform's own engagement scores) recover relatively quickly. External reputation with recipient servers can take up to a month or longer of patient, consistent sending. Full inbox recovery after severe damage, as documented by the community, takes 8–12 weeks (see Recovering inbox placement).
- Track the quality of engagement continuously during the ramp. SendGrid points to engagement recency and unique open rate as the scores to watch. On any platform, the equivalents are the trends in opens, clicks and complaint rate at each provider.
Re-warming after an incident vs cold warm-up
Re-warming uses the same mechanics as a cold warm-up, but the situation differs in ways that change the plan:
| Dimension | Cold warm-up (new IP or domain) | Re-warm (after an incident) |
|---|---|---|
| Starting reputation | Slightly negative: unknown, and treated with suspicion | Actively negative: receivers have a concrete bad history tied to the asset |
| Prerequisite | Nothing beyond setting up DNS and authentication | Root cause fixed, queues purged, suppressions rebuilt, delistings requested. Ramping before this triggers the penalties again |
| Pacing | A standard schedule (for example, +50% per week, or a vendor's stage plan) | The same shape, but expect to hold or step back at individual providers. Receivers that blocked you throttle again, harder and for longer |
| Audience | Recently engaged recipients | Even stricter: only recently engaged recipients, with high-risk flows (win-back, re-engagement, third-party) suspended for the whole recovery |
| Signals that gate each increase | Deferral and failure rates, inbox placement | The same, plus a blocklist check before each increase. A new listing during a re-warm is treated more harshly than the first listing |
| Timeline | Weeks to full volume | External trust: from up to a month (SendGrid) to 8–12 weeks for full inbox recovery |
Two further rules documented by vendors bear on re-warming:
- Dormancy resets warm-up. SendGrid says that if an IP has not sent mail in over 30 days, warm-up should be resumed before returning to volume. A recovery pause of several weeks therefore creates a re-warm requirement by itself, independently of the incident.
- Prevention is cheaper. In SendGrid's own words, "establishing a positive reputation as a sender takes less effort than repairing an existing reputation." Budget recovery time accordingly.
Deferral handling during recovery
Deferrals (4xx temporary failures) are the main real-time feedback during a re-ramp. Receivers rarely explain a reputation penalty, but they always throttle. The concepts below come from SendGrid's documentation on deferrals, and the mechanics apply to any mail transfer agent (MTA).
What a deferral is: the recipient server cannot accept the message at the moment. It is a temporary condition, not a rejection, and retries may succeed. All senders receive some deferrals. The signal is in their rate and their pattern, not in their existence.
Retry behavior (SendGrid's implementation, typical of ESPs): deferred messages are retried with exponential backoff for up to 72 hours. A message that is still undelivered after that is finalized as blocked or expired. During automated warmup, a warming IP that reaches its hourly cap stops sending, and the account's other IP addresses carry the traffic. With no backup IP addresses, messages are retried roughly every 15 minutes for 72 hours before they expire.
Types of deferral. Work out who is doing the throttling:
| Type | Example reasons | Meaning during recovery |
|---|---|---|
| External (initiated by the receiver) | "IPs were throttled by recipient server" | The provider is limiting you because of your reputation. This is the core recovery signal |
| External (based on limits) | "IPs reached ISP-suggested hourly limits" | You exceeded the warmup or ramp threshold for that provider. This is a pacing problem, not necessarily a reputation problem |
| Internal (initiated by the platform) | "reached ISP-suggested max connection limits", "max port limit", "max connection limit" | Your own platform is deliberately slowing delivery to protect reputation. Do not fight it |
Operational rules during a re-ramp:
- When the deferral rate rises at one provider, hold or reduce volume at that provider. Deferrals are throttling, the step before blocking (see warning signs). Do not advance the ramp while a provider is deferring at elevated rates.
- Do not force your way through deferrals. Opening more connections, or injecting deferred mail again as new messages, turns throttling into blocks.
- For deferrals caused by limits, reduce the delivery rate, or spread delivery across more warmed IP addresses. For deferrals where the receiver is throttling you, let the retry system pace delivery. SendGrid's documented remediation table says "no action required; delivery auto-slowed."
- Measure the impact through end-to-end delivery time, the gap between your platform accepting the message and its final delivery. Deferrals with short end-to-end times mean the throttle is absorbing your ramp without harming campaigns. Growing end-to-end times mean you are ramping faster than the receiver will accept.
- Pause entirely when deferrals at a provider turn into hard blocks, or into 72-hour expirations, at meaningful rates. That provider is telling you the reputation is not ready. Drop volume there to near zero, and rebuild with only your most engaged recipients.
Automated warmup schedules as implemented
Vendor automation is useful reference data when you build your own ramp logic. The mechanics are described conceptually, and each number is attributed to its vendor.
SendGrid automated warmup (41-day hourly-cap schedule)
SendGrid throttles a warming dedicated IP with an hourly sending cap that grows by about 40% per day over 42 days (day 0–41). When the cap is reached, the warming IP stops for the rest of the hour, and the account's other IP addresses carry the overflow, including other warming IP addresses that still have room. With no backup IP addresses, mail is retried about every 15 min for up to 72 h. After day 41 the IP leaves warmup. The schedule, as published by the vendor:
| Day | Hourly cap | Day | Hourly cap | Day | Hourly cap |
|---|---|---|---|---|---|
| 0 | 20 | 14 | 2,222 | 28 | 246,953 |
| 1 | 28 | 15 | 3,111 | 29 | 345,735 |
| 2 | 39 | 16 | 4,356 | 30 | 484,029 |
| 3 | 55 | 17 | 6,098 | 31 | 677,640 |
| 4 | 77 | 18 | 8,583 | 32 | 948,696 |
| 5 | 108 | 19 | 11,953 | 33 | 1,328,175 |
| 6 | 151 | 20 | 16,734 | 34 | 1,859,444 |
| 7 | 211 | 21 | 23,427 | 35 | 2,603,222 |
| 8 | 295 | 22 | 32,798 | 36 | 3,644,511 |
| 9 | 413 | 23 | 45,917 | 37 | 5,102,316 |
| 10 | 579 | 24 | 64,284 | 38 | 7,143,242 |
| 11 | 810 | 25 | 89,998 | 39 | 10,000,539 |
| 12 | 1,000 | 26 | 125,997 | 40 | 14,000,754 |
| 13 | 1,587 | 27 | 176,395 | 41 | 19,601,056 |
SendGrid adds two caveats. Transactional streams should not be forced onto a strict schedule, because you cannot control how often they are triggered. And no schedule replaces good sending practices: a gradual ramp alone does not guarantee reputation.
Mailgun warmup (volume-stage model and API)
Mailgun's automatic warmup uses stages based on volume rather than calendar days. Each stage has a daily cap, and the IP moves to the next stage once it has sent that stage's volume. Stage progression and the 24-hour window are independent: sending the full volume of a stage moves you on however many hours have passed, and the 24-hour window starts with the first message.
The stage caps Mailgun publishes are 1,000 per day for Stage 1 and 2,500 per day for Stage 2. Later caps are not published, and according to the schedule model in the API, a plan can run to 15 stages. Reaching a daily cap early stops all sending from that IP until the 24-hour window resets. Overflow traffic is rerouted, not dropped: volume beyond the warming IP's cap moves to the account's shared IP addresses or to other dedicated IP addresses, so the mail is still sent, just not from the warming IP. A full warmup typically takes 4–8 weeks.
You can drive the warmup through the API (Mailgun API reference, specific to this vendor):
| Operation | Endpoint | Purpose |
|---|---|---|
| GET | /v3/ip-warmups |
List the status of IP warmups in progress |
| GET | /v3/ip-warmups/{addr} |
Status of one warmup in progress (current stage, stage and hourly limits, total number of stages) |
| POST | /v3/ip-warmups/{addr} |
Create a warmup plan for an IP |
| DELETE | /v3/ip-warmups/{addr} |
Cancel the warmup plan |
Mailgun's guidance for manual warmup (an archived help-center article, 2023 snapshot) makes three recommendations:
- Use dedicated IP addresses from 100,000 emails per month. This is Mailgun's own rule. To compare it with other vendors' minimums and with the general statistical floor for each IP, see the canonical table of volume floors for dedicated IP addresses.
- Start at 100 emails on day one and increase by about 20% per day, sending every day. Sending three times a week or less works, but builds trust more slowly.
- Do not use domains less than 30 days old at all. Receivers associate newly purchased domains with spammers who cycle through burned domains.
Design implications for your own ramp automation. The two implementations agree on four points. Cap the warming asset instead of queueing and forcing mail, and send the overflow down a path that is already warm. Grow caps geometrically (about 20–50% per day). Make progression depend on volume actually sent, not only on elapsed time. And treat push-back from receivers (deferrals) as a stop condition that the automation must respect.
Replace or rehabilitate? Decision points
After a severe incident, the tempting shortcut is a fresh IP address or domain. The default answer is to rehabilitate, because replacement rarely escapes the problem:
- New assets start negative, not neutral. Receivers presume that a new IP belongs to a blocked spammer who has moved (see IP Warm-Up), and they treat domains younger than about 30 days as a sign of spammers cycling domains (Mailgun). Swapping assets is exactly the pattern filters are built to catch.
- Reputation follows the mail, not only the asset. Domain reputation, content fingerprints and list quality move with you to the new IP. If the root cause is not fixed, the new asset burns out faster than the old one recovered.
- Listings caused by a compromise clear quickly. Blocklist operators and providers routinely reverse listings caused by a compromise that is documented and remediated, so rehabilitation is a real option in this scenario (see the delisting workflow).
Replacement, with a full warm-up from zero, is the right choice only when rehabilitation is not possible or costs more than starting over:
| Factor | Favors rehabilitation | Favors replacement |
|---|---|---|
| Listing status | Can be delisted; the operator responds to evidence of remediation | Permanent or repeated listing after badly handled requests; the asset is on lists with no practical way to be removed |
| History depth | A short incident on an otherwise good asset | A long history of abuse that predates you (for example, an IP or domain you inherited that was abused before) |
| Provider response after remediation | Deferrals easing, and reputation dashboards recovering within weeks | Sustained hard blocks at major providers despite a clean re-warm over several weeks |
| Asset role | The primary brand domain (organic, inherent reputation, Brand Indicators for Message Identification (BIMI), customer recognition: effectively irreplaceable) | A dedicated sending IP or a sending subdomain, cheap to rotate and warm up again |
| Time economics | Recovery expected to take roughly no longer than a warm-up (a re-warm on a rehabilitated asset can be faster than the 4–8+ weeks a cold warm-up takes) | Recovery has stalled beyond the timeline of a cold warm-up |
For domains, there is a practical middle path. Never replace the organizational domain. Rehabilitate it, and move mail streams to a fresh, properly delegated subdomain warmed up from zero. This isolates the damage without the penalty for a new domain. For the mechanics of isolating streams, see M3AAWG Sending Domains BCP and Advanced IP Segmentation. Whatever you replace, retire the burned asset cleanly, keeping it authenticated and its protective DNS records in place (see Brand Protection: Domain Management).
Related articles
- Reputation Monitoring and Remediation, including the recovery playbook based on engaged segments
- IP Warm-Up
- Sending Infrastructure Practices, on subdomain strategy and domain warm-up
- Blocklists and Spamhaus
- Mandated and Regulatory Email, for sends that cannot wait for a recovery window
Check your own record
The free check reads what your domain publishes in DNS.
In this topic
- Deliverability Metrics and Benchmarks
- List Hygiene and Sunset Policies
- Sending Infrastructure Practices
- MTA Delivery Tuning