emailmarketing.net

MTA Delivery Tuning

Outbound queue theory and delivery tuning as a discipline — queue-age diagnosis, per-destination concurrency and rate shaping, connection reuse, adaptive backoff and dead-site detection, scheduler fairness, and bounce-driven traffic-shaping automation.

Operational10 min read

Who it is for ESP operators

If your outbound queues are growing or providers are deferring your mail, the cause is often how hard your mail server pushes each destination. Delivery tuning covers how a mail transfer agent (MTA) decides that pressure, and how operators diagnose and correct delivery congestion.

The concepts below apply to any MTA. The models for queue diagnosis, concurrency feedback and scheduling come from Postfix's documentation, and the models for shaping rules and automation come from KumoMTA's. Every serious MTA, commercial or open source, implements some version of each. For the concrete numbers for each provider that these mechanisms enforce, see Per-Provider Tuning Baselines.

Why this is a deliverability discipline

Mailbox providers enforce caps on connections per sender, on message rate and on messages per connection, and they answer violations with 4xx deferrals (often 421) or throttling penalties. An MTA that pushes past those limits turns good mail into deferred mail and ages its queues. At providers that score "unusual traffic patterns", it also damages IP reputation. Delivery tuning is the practice of keeping outbound pressure just under each destination's tolerance, backing off automatically when the destination signals distress, and diagnosing queue problems quickly when they appear.

Queue-age distribution as diagnosis (the "qshape" model)

The single most informative view of an outbound MTA is a two-dimensional table, with destination domain on one axis and message age on the other, and counts in each cell. Postfix ships this view as qshape. Age buckets are fine-grained for recent mail, typically 5, 10, 20, 40, 80, 160, 320, 640, 1280 and 1280+ minutes, with a total column for each domain. Run it separately against the incoming and active queues (mail being worked on now) and against the deferred queue (mail that has already failed temporarily).

Reading rules:

  1. Problems live in the top left. Large counts for a single domain in the young buckets mean an active, ongoing problem. A concentration in old ages means a past problem that is now draining.
  2. Concentration on a domain identifies the bottleneck. Sort by row total. The domains that dominate are where delivery is failing or throttled.
  3. A backlog in the active queue matters more than a backlog in the deferred queue. Mail stuck in the active queue blocks new work, while deferred mail is only waiting for a retry.
  4. A full active queue with an incoming queue that is not full means that one or more destinations drain more slowly than mail arrives for them.
  5. A view by sender detects backscatter. Pivoting the same table by sender instead of recipient shows when bounce messages (MAILER-DAEMON) dominate the queue, which means the load is self-inflicted bounce processing, not fresh mail.

Characteristic patterns:

Pattern Signature in the age table Diagnosis
Healthy Incoming and active queues nearly empty; small deferred totals spread across old buckets Normal operation
Deferred backlog, old ages Large deferred totals, with counts growing toward the oldest buckets; short incoming and active queues A past problem already resolved, with retries draining. "high volume of deferred mail is not a direct cause for alarm" by itself
Bounce backscatter As above, but the sender view shows mostly MAILER-DAEMON An earlier bad send (for example, to a list hit by a dictionary attack) still generating bounce work
Active-queue saturation Active count near its cap (Postfix example: 9,996 of a limit of 10,000; 20,000 in later versions), almost all in the youngest bucket, with one domain dominating A destination is draining far more slowly than mail arrives. Delivery agents are tied up, and new mail is delayed for all destinations
High-volume destination backlog One domain with large counts in every bucket younger than the start of the problem, and still growing The destination has been down or throttling since roughly the oldest affected bucket; congestion is ongoing

Scale expectations (Postfix figures, indicative for any store-and-forward MTA): a deferred queue handles roughly 100,000–1,000,000 messages, and good performance is unlikely beyond that.

Per-destination concurrency shaping

MTAs open a limited number of parallel connections to each destination, and adapt that number according to delivery outcomes. This is deliberately similar to TCP slow start.

  • Initial concurrency is small (Postfix default: 5 simultaneous deliveries), so that a new or recovering destination is probed gently.
  • Maximum concurrency caps parallelism for each destination (Postfix default: 20). Overrides for individual destinations are how provider caps are honored, for example for a provider that tolerates only 2 connections.
  • Positive feedback raises concurrency after successful deliveries, and negative feedback lowers it after connection or handshake failures. The feedback increment is a function of the current concurrency N:
    • +1 per success gives an exponential ramp (5, then 10, then 20): fast, but prone to oscillation;
    • +1/N per success gives a linear ramp, one extra slot after N successes: the gentlest;
    • +1/√N sits in between. Fractional feedback combined with integer truncation gives natural hysteresis: concurrency only steps up after 1/g(N) consecutive successes.
  • The asymmetry matters. Negative feedback is applied at the start of a sequence of failures ("reverse hysteresis"), so that overload is corrected immediately, not after a full window of failures.
  • Measured effect (Postfix scheduler tests against a receiver that enforces concurrency limits with 421 responses): fixed feedback of ±1 deferred about 50% of mail; 1/N feedback deferred about 16.5%; 1/√N about 24.5%. Gentler feedback sharply reduces deferrals at receivers that enforce limits.
  • Caveat: feedback below 1 has little effect when volume to a destination is low, because deliveries are too infrequent for the counters to move.

Dead-site detection and penalty boxes

This is separate from concurrency feedback. When a whole cohort of delivery attempts (N attempts, where N is the current concurrency) fails with connection or handshake errors, the destination is declared dead after a configurable number of failed cohorts (Postfix default: 1 cohort). Dead destinations go into a penalty box, where they are skipped entirely for a period, instead of being hammered at a concurrency of 1. Separating "the site is down" from "the site is slow" lets concurrency feedback stay gentle without wasting delivery agents on unreachable hosts.

Rate shaping, connection reuse, and the shaping-parameter vocabulary

Beyond concurrency, modern MTAs shape traffic to each destination along several independent axes (the names come from KumoMTA, and every MTA has equivalents):

Parameter (concept) Meaning Typical global default (KumoMTA)
Connection limit Maximum simultaneous connections to the destination 10
Connection rate Maximum new connections per unit of time 100 per minute
Message rate Maximum messages per unit of time, regardless of connections 100 per second
Deliveries per connection Messages sent before an SMTP session is closed and reopened 100
Idle timeout How long a cached connection may stay idle before it is closed 60 s
Data timeouts Timeouts for the DATA phase and the final dot 30 s and 60 s
TLS mode Opportunistic or Required, per destination Opportunistic
Consecutive connection failures before delay The number of failures that triggers backoff 100

Connection reuse (sending many messages over one connection) is a big lever. Setting up a connection, negotiating TLS and the receiver's accounting for each connection all cost more than an extra MAIL FROM. But providers cap messages per connection and answer the excess with errors such as "Max message per connection reached", so the right value depends on the provider (50 for Gmail, 20 for Yahoo, 5 for Apple in community baselines). A related setting is the number of recipients per copy of a message (Postfix default: 50 recipients per copy; larger recipient lists are split into parallel copies).

Destination grouping (MX rollup): shaping limits should apply to each receiving infrastructure, not to each recipient domain. Thousands of hosted domains share the MX fleets of Google or Microsoft. Rolling up all the domains whose MX records match a suffix (for example, .google.com or .googlemail.com) into one shaping bucket prevents you from accidentally multiplying your effective connection count by the number of recipient domains. Conversely, a provider whose MX names are regional silos (for example, Mimecast site names) should be limited per MX site, not per provider.

Retry scheduling and adaptive backoff on 4xx

Mail that failed temporarily (4xx) is given a future timestamp with exponential backoff between a floor and a ceiling (Postfix defaults: minimum 300 s, maximum 4,000 s), and the deferred queue is scanned again periodically (by default every 300 s). Messages that cannot be delivered within the queue lifetime (default 5 days; bounces often get their own lifetime, also 5 days) are returned as permanent failures.

Operational rules:

  • Do not shorten retry intervals to "clear the backlog." Retrying undeliverable mail more often saturates the active queue and ties up delivery agents on dead sites, which starves deliverable mail. Fixing the cause always works better than retrying harder.
  • Timeouts limit worst-case throughput. For example, with 100 delivery agents, a destination with 2 MX hosts of which one is down, and a connect timeout of 30 seconds, throughput tops out at about 6 messages per second, because half the connection attempts use up a full timeout. Shorter connect timeouts (and trying the next MX quickly) restore throughput during partial outages.
  • Isolate problem destinations. Give destinations that are chronically slow or unreliable a dedicated delivery pool with shorter timeouts and adjusted concurrency, or move them to a fallback "graveyard" relay with its own retry schedule, so that they cannot degrade mainstream delivery.

Scheduler fairness and preemption

Inside the MTA, a scheduler decides which message gets the next free delivery slot:

  • Round-robin across transports and destinations makes sure that no single destination monopolizes delivery agents.
  • First in, first out (FIFO) within a destination, by the time mail entered the queue, except for preemption.
  • Preemption lets small messages get past bulk ones without starving the bulk mail. As a large message's recipients are delivered, it accumulates "delivery slots" (one for every K delivered recipients; Postfix default cost K=5). A candidate to preempt it is picked by maximizing enqueue_time / recipient_count (the oldest relative to its size first), and needs a minimum number of slots (default 3). Slot loans (by default up to 3 slots, at a 50% discount) let a small message jump ahead immediately and "pay back" the advance from later deliveries. This expedites mail the size of transactional messages without giving up fairness to bulk campaigns. It is the equivalent, inside the MTA, of the practice among email service providers (ESPs) of separating transactional and bulk streams.

Bounce-driven traffic-shaping automation

Static shaping tables set the ceiling, and automation adjusts below it in real time by parsing what providers say in their SMTP responses. The model is KumoMTA's Traffic Shaping Automation, which runs as a separate daemon that watches delivery logs and pushes configuration back to the MTA. It can be clustered, so that all nodes react together:

  • Match: regex rules over the text of bounce and deferral responses, scoped to a provider or site, a tenant, or a campaign.
  • Trigger: fire immediately, or only past a threshold (for example, 2 matches an hour, which tells a one-off from a pattern).
  • Act: either Suspend (stop sending to that destination from that queue entirely) or SetConfig (tighten a shaping parameter: drop connection_limit to 1, cap max_message_rate at 1 per minute, cut max_deliveries_per_connection). Suspension scoped to a tenant exists for responses that indicate a sender problem (for example, missing authentication) rather than a traffic problem.
  • Expire: every action has a duration (commonly 15 minutes–4 hours). When it lapses, normal shaping resumes automatically. This makes automated throttling self-healing: no operator has to remember to remove the throttle.

Representative community rules (see Per-Provider Tuning Baselines for the full table):

  • Gmail "unusual rate" responses: cap at 10 messages per minute for 30 minutes.
  • Yahoo [TSS04] complaint deferrals: suspend for 2 hours.
  • Outlook "exceeded the maximum number of connections": a connection limit of 1 for 1 hour.
  • A generic default rule matches the family of texts seen across providers, such as "temporarily deferred", "rate limited due to IP reputation" and "Server busy", and drops both the message rate to 1 per minute and connections to 1, for 90 minutes.

Design principles for building your own rules:

  1. Parse the provider's words, not only the code. 421 alone is ambiguous, while [TS02], [TS03] and "Max message per connection reached" each call for a different response (see the error code guides for Gmail, Yahoo and Microsoft).
  2. Match the action to the complaint. For complaints about connections, cut connections. For complaints about volume, cut the message rate. For deferrals about reputation or complaints, and for blocklistings, suspend and let the reputation cool. For content rejections, suspend and alert a person, because throttling will not fix content.
  3. Throttle harder for reputation signals than for mechanical limits. Hitting a connection cap is a tuning miss; "deferred due to user complaints" is a reputation event.
  4. Keep durations short, and let rules fire again. A penalty of 30–120 minutes that renews while the provider keeps complaining converges on the right rate without permanently crippling the queue.

Layering shaping configuration

Keep shaping data in layers, so that updates do not destroy local knowledge: a vendor or default layer, a community or shared layer, and a local layer with your own rules. Later layers override earlier ones, with an explicit option to "replace everything for this domain". Never edit the distributed layers in place, because upgrades overwrite them. Most deployments need local rules, because the shared baselines are deliberately not all-encompassing.

Check your own record

The free check reads what your domain publishes in DNS.

In this topic

All 18 in Operations →