# MIME & Transfer Encodings (RFC 2045/2046)

> Deliverability-focused MIME digest: structure and headers, multipart/alternative ordering, quoted-printable and base64 rules and pitfalls, size inflation, and the malformed-MIME defects that trigger filtering.

Source: emailmarketing.net — https://emailmarketing.net/learn/rfc/mime-and-encoding

Every HTML campaign you send is a MIME message, and defects in its encoding or structure are a recurring cause of broken rendering and filter penalties.

MIME extends the [RFC 5322 message format](https://emailmarketing.net/learn/rfc/rfc5322-message-format), which allows only ASCII text in lines, so that a message can carry several parts, non-ASCII text and binary data over transports that only guarantee 7-bit lines. RFC 2045 defines the framework and the encodings, and RFC 2046 defines the media types.

## The five MIME header fields (RFC 2045)

| Field | Requirement | Notes |
|---|---|---|
| `MIME-Version: 1.0` | MUST be present, exactly as written, on every MIME message | Without it, receivers may treat the message as plain text from before MIME |
| `Content-Type: type/subtype; parameter=value` | The subtype is mandatory. The names of the type, the subtype and the parameters are case-insensitive, and parameter values are case-sensitive unless defined otherwise | Default when absent: `text/plain; charset=us-ascii` |
| `Content-Transfer-Encoding: mechanism` | One of `7bit` (the default), `8bit`, `binary`, `quoted-printable`, `base64`, or an extension | See below |
| `Content-ID: <msg-id>` | Optional; MUST be unique worldwide | Used to reference inline images (`cid:`) from HTML parts |
| `Content-Description` | Optional, in US-ASCII (RFC 2047 for non-ASCII) | Rarely matters in practice |

## Content-Transfer-Encoding mechanisms

| Mechanism | Constraint on the data | Actually transforms? |
|---|---|---|
| `7bit` | Lines of at most 998 octets between CRLFs; no octets above 127; no NULs | No. It declares that the data already conforms |
| `8bit` | Lines of at most 998 octets; octets above 127 allowed; no NULs; CR and LF only as CRLF pairs | No. It requires the 8BITMIME SMTP extension from end to end |
| `binary` | Any sequence of octets, with no limit on line length | No. It requires BINARYMIME and [CHUNKING](https://emailmarketing.net/learn/rfc/esmtp-extensions), and is essentially unused in Internet mail |
| `quoted-printable` | The output is 7-bit, with lines of at most 76 characters | Yes |
| `base64` | The output is 7-bit, with lines of at most 76 characters | Yes |

**Restriction on composite types:** it is "EXPRESSLY FORBIDDEN" to use any encoding other than `7bit`, `8bit` or `binary` on `multipart/*` or `message/*` entities. Encoding applies only to leaf parts, never to a container, which prevents nested passes of encoding and decoding. RFC 6532 later relaxed this for `message/global` only (see [SMTPUTF8 and EAI](https://emailmarketing.net/learn/rfc/smtputf8-eai)).

### Quoted-printable rules (RFC 2045 §6.7)

Quoted-printable (QP) is designed for data that is mostly ASCII text, so the encoded output stays largely readable by people.

1. Any octet may be represented as `=XX` with **uppercase** hexadecimal digits. Lowercase is not permitted when generating, although robust decoders accept it.
2. Octets 33–60 and 62–126 (printable ASCII except `=`) may appear literally.
3. TAB (9) and SPACE (32) may appear literally, but **MUST NOT be literal at the end of an encoded line**. Trailing whitespace must be written as `=09` or `=20`, because the transport may strip trailing spaces silently.
4. A hard line break in text is represented as a real CRLF.
5. Encoded lines MUST be at most **76 characters** long (not counting the trailing CRLF). Longer lines of content use a **soft line break**: an `=` as the last character, which the decoder deletes along with the CRLF that follows it.
6. When binary data is encoded in QP, CR and LF must be encoded as `=0D` and `=0A` (a binary CRLF becomes `=0D=0A`). QP for binary data is legal, but base64 is the right tool.

In practice, QP inflates typical HTML in English or Western European languages only slightly (3 bytes for each encoded octet). For scripts where nearly every character is multi-byte UTF-8, such as Chinese, Japanese and Korean (CJK) or Cyrillic, QP approaches 3× inflation, and base64 (about 1.37×) is smaller.

### Base64 rules (RFC 2045 §6.8)

- The alphabet has 65 characters: `A–Z` (0–25), `a–z` (26–51), `0–9` (52–61), `+` (62), `/` (63), and `=` for padding.
- 3 input octets become 4 output characters. A final partial group is padded: 1 octet becomes 2 characters plus `==`, and 2 octets become 3 characters plus `=`. Any `=` marks the end of the data.
- Encoded output MUST be in lines of at most **76 characters**.
- Line breaks in text must be converted to CRLF **before** encoding (the canonical form), so a decoded text part uses CRLF line endings whatever platform composed it.
- **Size inflation: 4/3 (about 33%), plus a CRLF every 76 characters, comes to about 37% in total.** Provider limits on message size (for example, Gmail's 25 MB) apply to the encoded message, so a "20 MB attachment" is about 27 MB on the wire. The size declared with the SIZE extension also includes this inflation (see [ESMTP Extensions](https://emailmarketing.net/learn/rfc/esmtp-extensions)).

### Choosing an encoding (sender guidance)

- For HTML and plain-text parts in UTF-8, `quoted-printable` is the safe universal default. It survives 7-bit hops, keeps diffs and DKIM canonicalization predictable, and avoids depending on 8BITMIME negotiation. `8bit` is fine when your MTA verifies 8BITMIME support on every hop. Support is essentially universal today, but a hop that accepts only 7-bit data must then re-encode the message, which **breaks DKIM signatures**. Sign after the final encoding, or use QP so that no hop ever re-encodes.
- For attachments and images, always use `base64`.
- Never emit lines longer than 998 octets in any part. Long unwrapped HTML lines are a classic generator bug, and the soft wrapping of QP at 76 characters fixes this at no cost.

## Media types (RFC 2046)

The discrete types are `text`, `image`, `audio`, `video` and `application`. The composite types are `multipart` and `message`. Unrecognized subtypes fall back as follows: text with a known charset is treated as `text/plain`, other discrete types as `application/octet-stream`, and an unknown `multipart/*` as `multipart/mixed`.

### Multipart syntax

- The `boundary` parameter is **mandatory**. It is 1–70 characters long and must not end in a space. The allowed characters are digits, letters, `'()+_,-./:=?` and the space (but not as the last character). Because `/`, `=`, `?`, `:` and space can appear, the boundary value usually needs quotes in the Content-Type header.
- The delimiter is a line that starts with `--boundary`, and the closing delimiter is `--boundary--`. The boundary **must not appear in any encapsulated part**. Generators should use boundaries with high entropy rather than checking the content.
- Each body part consists of optional headers, a blank line and the body. Parts with no Content-Type default to `text/plain; charset=US-ASCII`, except inside `multipart/digest`, where the default is `message/rfc822`.
- MIME readers ignore the preamble (before the first boundary) and the epilogue (after the closing delimiter). The preamble is occasionally used for a "this is a MIME message" note to legacy clients.

### Multipart subtypes that matter for senders

| Subtype | Semantics | Deliverability notes |
|---|---|---|
| `multipart/mixed` | Independent parts, in order | A message with attachments |
| `multipart/alternative` | The same information in alternative renditions. **Parts MUST be in increasing order of preference and faithfulness, with the best and richest version LAST.** Clients display "the last part of a type supported by the recipient system's local environment." | The classic structure is `text/plain` first and `text/html` last. In the reverse order, many clients show the plain-text part. A well-formed pair, where the plain part actually matches the HTML content, has long been a positive signal for spam filters, while an empty or boilerplate plain part is a mild negative |
| `multipart/related` (RFC 2387, included for completeness) | HTML with inline `cid:` images | The standard structure for embedded images |
| `multipart/digest` | Parts default to `message/rfc822` | Rare in the traffic of email service providers (ESPs) |
| `multipart/parallel` | Order is not significant | Effectively unused |
| `multipart/report` (RFC 6522) | Machine-readable reports | The container for [delivery status notifications (DSNs)](https://emailmarketing.net/learn/bounce-handling/delivery-status-notifications) and for feedback loop (FBL) reports in the Abuse Reporting Format (ARF) |

The canonical nesting for a campaign is `multipart/mixed( multipart/related( multipart/alternative( text/plain, text/html ), inline images ), attachments )`. Collapse the levels you do not need.

### `message/*` subtypes

- `message/rfc822`: a full encapsulated message, used for forwards, FBL reports and the original messages returned in bounces. Its encoding is restricted to 7bit, 8bit or binary.
- `message/partial`: fragmented delivery with reassembly (the parameters `id`, `number` starting at 1, and `total` on the last piece; 7bit only). It is obsolete in practice, and filters treat fragmentation as an evasion technique. Never use it.
- `message/external-body`: a body by reference (access types FTP, ANON-FTP, TFTP, LOCAL-FILE and MAIL-SERVER). It is also effectively dead, and filters treat it as suspicious.

### text/plain and charsets

The default charset is `US-ASCII`, which means ANSI X3.4-1986 specifically. Always declare `charset=` explicitly, and label content with the lowest common denominator that actually covers it. In modern practice, use `charset=utf-8` everywhere and the question is settled.

## Malformed-MIME defects that trigger filtering

Spam filters score MIME conformance because sloppy generators correlate with spamware. SpamAssassin's `MIME_*` family of rules, and the equivalent rules in commercial filters, test exactly these defects. They are worth linting for in an ESP's message builder:

| Defect | RFC rule violated |
|---|---|
| Missing `MIME-Version: 1.0` on a message with MIME structure | 2045 §4 MUST |
| Missing or unmatched final `--boundary--` closing delimiter | 2046 §5.1.1 |
| Boundary string that occurs inside the body of a part | 2046 §5.1.1 "must not appear inside any of the encapsulated parts" |
| Raw 8-bit bytes in a part declared `7bit` or `quoted-printable`, or in headers | 2045 §2.7 and §6.7 |
| Lines over 998 octets (unwrapped HTML) | 2045 §2.7; [5321 limit](https://emailmarketing.net/learn/rfc/rfc5321-smtp#size-limits-section-4531--minimums-implementations-must-handle) |
| QP lines over 76 characters, lowercase `=xx` hex, literal trailing whitespace | 2045 §6.7 rules 3 and 5 |
| Base64 with invalid characters or broken padding | 2045 §6.8 |
| Content-Transfer-Encoding other than 7bit, 8bit or binary on a multipart | 2045 §6.4 |
| Charset declared but content not valid in it (for example, a `us-ascii` label on an 8-bit body) | 2045 §5.2; 2046 §4.1 |
| `multipart/alternative` with HTML before plain text, or HTML alone inside an alternative | 2046 §5.1.4 ordering |
| Bare CR or LF (not in CRLF pairs) in transmitted data | 2045 §2.8; 5321 |
| `message/partial` fragmentation | Filter evasion heuristic |

None of these defects alone typically blocks delivery at the major providers, but their effect adds to reputation. Several of them (boundary bugs, encoding mismatches) also visibly corrupt rendering, which drives "report spam" clicks.

## Related articles

- [RFC 5322 (Message Format)](https://emailmarketing.net/learn/rfc/rfc5322-message-format), the header framework MIME extends
- [SMTPUTF8 and EAI](https://emailmarketing.net/learn/rfc/smtputf8-eai), on raw UTF-8 in headers and `message/global`
- [ESMTP Extensions](https://emailmarketing.net/learn/rfc/esmtp-extensions), including CHUNKING and BINARYMIME
- [Content and Design for Deliverability](https://emailmarketing.net/learn/operations/content-and-design-for-deliverability), on content signals beyond structure
