emailmarketing.net

Multi-Tenant ESP Architecture — IP Pools and Tenant Isolation

How mature sending platforms structure multi-tenancy — dedicated vs managed IP pools, shared-pool fallback, per-tenant reputation tracking, and automated tenant enforcement — extracted as design evidence from AWS SES and SendGrid.

Operational10 min read

Who it is for ESP operators

If your platform sends mail on behalf of many customers, you have to decide how to allocate IP addresses, how to keep one customer's reputation from damaging another's, and how to enforce limits on each customer. AWS SES and Twilio SendGrid publicly document how their own platforms answer these questions. Their documentation is used here as design evidence for anyone running a multi-tenant sending platform, not as a product tutorial.

For the same questions seen from the sender's side (how many IP addresses, which streams to separate), see Basic IP Allocation and Advanced IP Segmentation.

The three isolation layers

Mature platforms separate three concerns that are often confused:

Layer Isolates Mechanism (vendor examples)
IP pool IP reputation between mail streams (transactional and marketing) or between tiers of customers SES dedicated IP pools; SendGrid IP pools
Tenant or subuser Reputation metrics, enforcement and resources at the level of the customer SES tenants; SendGrid subusers
Suppression scope Do-not-send data between customers SES suppression at the level of the account, tenant or configuration set; see Suppression-List Architecture

A tenant is a logical container for metrics, policy and permissions on resources. A pool is a physical routing decision about which IP addresses carry the mail. A tenant can send through shared IP addresses, through a dedicated pool, or through both.

IP pools: the routing primitive

Standard (self-managed) pools

Both SES and SendGrid model a pool as a named group of dedicated IP addresses, selected at send time:

  • The typical use case, at both vendors, is one pool for marketing and one for transactional mail, so a campaign that draws many complaints cannot degrade transactional delivery. The pool is the boundary of reputation.
  • How a pool is selected: SES binds a pool to a configuration set, and each send names the configuration set (directly, or through a default configuration set attached to the sending identity). SendGrid selects the pool for each message with an ip_pool_name parameter on the send call.
  • Exclusive membership: in SES, an IP address belongs to exactly one pool, and assigning it to a new pool removes it from the previous one. In the same way, a configuration set points to exactly one pool.
  • Fallback behavior: in SendGrid, a send that names no pool "will use any IP available, including pooled addresses". In other words, sends without a pool do not respect pool boundaries. The design implication is that pool isolation is only as strong as the discipline of always naming the pool.
  • A deliberate route to the shared pool: SES lets a configuration set route mail to the shared pool (IP addresses shared by all SES customers). This is documented for "email that doesn't align with your usual sending behaviors", meaning unusual sends that you do not want to affect the reputation of your dedicated IP addresses. The reasoning is the same as routing risky one-off sends away from your best IP addresses (see Mandated & Regulatory Email).

Documented limits:

  • SES: at most 50 dedicated IP pools per account per region (managed and standard pools combined), and pool names must be unique across both types of pool.
  • SendGrid: up to 100 IP pools per user, and pool names of at most 64 characters. Dedicated IP addresses must have reverse DNS configured and be activated before they are assigned to a pool. IP pools require higher-tier plans (Pro or Premier Email API, or Advanced Marketing Campaigns).

Managed (auto-warmed) pools: the SES "dedicated IPs (managed)" pattern

SES's managed pools show what an automated layer for managing pools looks like. AWS states each of these architecture points directly:

  • The number of IP addresses scales automatically. The platform decides how many dedicated IP addresses the pool needs from the sending patterns it observes, and scales that number up and down. Scaling takes each ISP into account: if a receiving ISP enforces a low daily quota for each IP address, the pool scales out to spread that ISP's traffic across more IP addresses.
  • Warm-up state for each ISP. The level of warm-up is tracked for each IP address and each receiving ISP, not globally. An IP address that has sent only to Gmail is warm for Gmail and cold for Hotmail, and increasing traffic to Hotmail starts a new gradual ramp for Hotmail only. (Compare the manual schedules for each provider in IP Warm-Up.)
  • Adaptive warm-up, including decay. The warm-up percentage drops when volume to an ISP drops. Warmth is treated as something that fades, not as a one-time achievement.
  • Overflow to the shared pool during warm-up. Early in warm-up, volume above the capacity already warmed spills over to the platform's shared IP pool rather than being sent cold, which protects the reputation of the new IP addresses. In later stages of warm-up, the excess is instead queued, slowed down and retried later through the dedicated IP addresses. Even fully warmed pools are not guaranteed 100% dedicated routing: a sudden spike in volume triggers the allocation of another IP address, whose warm-up again uses the shared pool.
  • Demotion when volume is low. If a sender with one dedicated IP address falls below the minimum volume needed to maintain its reputation, the platform removes the dedicated IP address and routes everything through the shared pool. AWS's stated threshold for allocating the first dedicated IP address is sending volume that reaches "hundreds of emails over a period of a few days". Below "a few hundred per day", AWS steers senders to shared IP addresses outright. This is the lowest of the published minimums for dedicated IP addresses. To compare it with other vendors' house rules and with the general statistical floor for each IP address, see the dedicated-IP volume floors table.
  • Promotion works in one direction only. A standard pool can be converted to a managed pool, but a managed pool cannot be converted back. On conversion, IP addresses that the volume does not justify are removed. AWS explicitly warns against converting IP addresses that appear on external allowlists, because they may be given up. Configuration sets and tags carry over.
  • The billing model changes: standard dedicated IP addresses are billed per IP address, while managed pools are billed on the volume sent through the pool. Deleting the last managed pool gives up all its IP addresses and stops the charges immediately.
  • A constraint: with managed dedicated IP addresses, an SES account is limited to 10,000 sending identities per region.
  • Shared responsibility, in AWS's own words: "managed" covers only the mechanics of scaling and warm-up. The customer remains responsible for the reputation results: bounce rates, complaint rates, and most requests for delisting from blocklists (RBLs).

Summary of a pool's lifecycle, the ladder of promotion and demotion these designs imply: shared pool, then a first dedicated IP address (once the volume threshold is reached), then the pool scales out as ISP demand requires, then scales in and gives up IP addresses when volume drops, and finally returns to the shared pool. At every transition, overflow to shared IP addresses absorbs the shock.

Tenant isolation and per-tenant reputation (SES tenant model)

SES tenants are the most explicit public documentation of isolating customers from each other inside a single sending platform. AWS's stated motivation is exactly the ESP problem: before tenants, "one customer's poor email practices could pause an entire SES account, affecting all other customers."

Structure

  • A tenant is a logical container that groups verified identities (domains and addresses), configuration sets and templates. The intended users are independent software vendors (ISVs) that send for many customers, enterprises with several business units, service providers that isolate each client or application, and organizations with different regulatory regimes for each tenant.
  • Resources are assigned either as dedicated (to one tenant) or shared (between several tenants). Every send in the context of a tenant is validated: the identity, the configuration set and the template must all be associated with the named tenant, or the send fails. A resource associated with a tenant cannot be deleted until it is disassociated.
  • Attributing a send: each send names its tenant, either with an API parameter (TenantName) or, over SMTP, with a message header (X-SES-TENANT: <name>). Every tenant send requires a configuration set associated with the tenant (named directly, or as the identity's default).
  • Flat and regional: tenants cannot be nested, cannot span AWS accounts, and exist in one region each. Senders in several regions configure and monitor tenants separately in each region.
  • Scale limits: 10,000 tenants per account by default, with increases approved automatically up to 300,000 for qualifying accounts. Pricing is per tenant per month, based on email volume.

Per-tenant reputation tracking

For each tenant, the platform continuously tracks the bounce rate, the complaint rate (including complaints from mailbox providers' feedback loops, FBLs), feedback signals from mailbox providers as third parties, and appearances of sending IP addresses on reputation blocklists. When a threshold is breached, the platform creates reputation findings at two levels of severity:

  • Low severity: minor issues that could affect deliverability if left unaddressed.
  • High severity: serious issues that are probably already affecting deliverability, and may trigger enforcement.

Each finding carries a type (BOUNCE, complaint, third-party feedback, blocklist), its impact, a description with the rate and sample that triggered it, and links for remediation. AWS's own documentation gives an example of how detailed findings are: "The bounce rate exceeded 15.0% based on a representative volume of 664 emails" over a window of ~2 hours. In other words, findings fire on rolling representative samples, not on totals for a calendar month. Metrics are computed over rolling windows of roughly 24 hours to 7 days depending on the metric, and some findings require a minimum representative volume before they can trigger.

Metrics for each tenant (Sends, Bounces, Complaints) are published to the monitoring system with the tenant as a dimension. Changes in a tenant's status, and findings, are emitted as events (EventBridge detail types: Sending Status Enabled or Disabled, and Advisor Recommendation Status Open or Closed), so the platform operator can automate alerts and responses.

Automated enforcement: reputation policies

Each tenant gets a reputation policy that decides when the platform pauses it automatically:

Policy Behavior AWS guidance
Standard (default) Pauses the tenant's sending on high-severity findings The recommended balance for most tenants
Strict Pauses on any finding, including low-severity ones For high-risk tenants or repeat offenders
None Never pauses automatically; findings are still recorded Only for onboarding under monitoring; carries a risk for Trust & Safety

A tenant's sending status can be Enabled, Paused (manually or by policy), Enforced (AWS Trust & Safety paused it for serious reputation issues) or Reinstated (reactivated after a pause). What these mean in practice:

  • A paused tenant's sends fail until an operator reviews the tenant and re-enables it manually. Resuming deliberately requires a person.
  • Grace period after reinstatement: once a tenant is re-enabled, its active findings are ignored for a while so it can recover. The tenant stays Reinstated until all findings are resolved. This prevents a loop of immediate new pauses.
  • Enforcement from upstream is targeted: when AWS Trust & Safety detects abuse, it can pause only the offending tenant instead of the whole account, and opens a support case for remediation. The tenant structure turns enforcement on a whole account into precise enforcement.
  • Liability is still combined: AWS states explicitly that the combined activity of the tenants still affects the reputation of the account as a whole. Isolation limits the damage but does not make bad traffic acceptable. The account owner is responsible for monitoring all tenants.

AWS's stated operating practices for tenant fleets

  • Start tenants on the Standard policy, and apply Strict to tenants that are high-risk or have offended before.
  • Onboard new tenants under None, with event monitoring, to observe their patterns before you turn on automatic enforcement.
  • Alert on findings (through events) so operators can act before automatic pausing.
  • Review tenant metrics regularly, even without findings, to catch patterns as they emerge.
  • Teach tenants good sending practices, and match the granularity of tenants to the business (for each customer, business unit, type of application or regulatory regime).

Design lessons for an ESP

The vendors' architecture can be read as a blueprint:

  1. Separate logical tenancy from physical IP routing. Reputation metrics, suppression scope and enforcement attach to the tenant, and IP addresses attach to pools. A routing object (a configuration set, or ip_pool_name) binds the two for each send.
  2. Metrics for each tenant, with automated pause policies, are what keep one bad customer from damaging shared infrastructure. They need severity levels, a person to approve resuming, and a grace period after reinstatement. They are the platform's counterpart to the thresholds for senders in Metrics & Benchmarks.
  3. The shared pool carries real load, and is not just an entry-level product. It absorbs overflow during warm-up, sudden spikes, demotions for low volume, and unusual sends routed there on purpose. A multi-tenant platform without a healthy shared pool has nothing to absorb shocks, which is why policing the shared pool (vetting, enforcement) matters so much.
  4. Warm-up state belongs to each pair of IP address and receiving ISP, and it decays. Automating pool scaling requires modeling the acceptance capacity of each ISP separately (see IP Warm-Up).
  5. Isolation has limits. The platform's own account (the ESP itself, as its upstream providers and peers see it) is still judged on its combined traffic. Isolating tenants limits the damage, but it does not excuse bad traffic.

For suppression scope, the third isolation layer, see Suppression-List Architecture.

Check your own record

The free check reads what your domain publishes in DNS.

In this topic

All 16 in ESP Operations →