# Gmail Filtering Internals

> How Gmail actually decides placement — per-domain authenticated-sender reputation, ML/LLM spam and phishing classification, tabbed-inbox categorization, and the per-user personalization signals that make one-size diagnosis impossible.

Source: emailmarketing.net — https://emailmarketing.net/learn/providers/gmail-filtering-internals

If your mail reaches the inbox for some Gmail users and the spam folder for others, the reason lies in how Gmail decides placement. The [Gmail sender requirements](https://emailmarketing.net/learn/providers/gmail-sender-requirements) tell you what Gmail checks. What follows is what Gmail does with the result. Two facts dominate, and they explain why Gmail deliverability resists one-size-fits-all diagnosis:

1. **Reputation is computed on the authenticated sending domain, not on the IP address.** This is the reverse of most receivers, which center reputation on IP addresses. Authentication is the prerequisite that makes a message possible to score at all.
2. **Placement is learned for each user.** The same message from the same sender can reach the inbox for one recipient and the spam folder for another, because Gmail lets a user's own history of reporting and marking override the sender's general reputation. There is no single answer to "am I reaching the inbox at Gmail?", only a distribution across recipients.

This is documented rather than inferred, mainly from Google's own research (Taylor, CEAS 2006; Taylor, Fingal and Aberdeen, NIPS 2007) and from its help and blog pages. The 2006 paper describes a system from the beta period, and the mechanisms have since been rebuilt on machine learning (see [ML classification](#ml-spam-and-phishing-classification)). But the paper is the only place where Google has set out the logic of reputation concretely, and its principles still describe how the system behaves: scoring on the authenticated domain, self-correction through user votes, and a middle band handed to a statistical filter. Treat the formulas as illustrations of the model, not as the code that runs today.

## The authenticated-domain reputation model

Source: **"Sender Reputation in a Large Webmail Service"**, by Bradley Taylor of Google, CEAS 2006 (Third Conference on Email and Anti-Spam). Gmail launched as a beta in April 2004, and the paper describes the reputation system built during that beta.

### Why the domain, not the IP

Google rejected IP reputation as the primary key because the IP "is a crude form of authentication". The same sender may not always use the same IP. Because of forwarding, the connecting IP is not always the true sender. IP addresses in `Received` headers cannot be authenticated reliably. And when several domains share a set of IP addresses, one spamming domain "can ruin the reputation for all the other domains." Gmail therefore classifies mail by **who is sending (the authenticated domain)** rather than by **what the content is**. The paper noted Gmail as "the only one" of the major reputation systems that worked on the authenticated domain rather than the IP.

These are the authentication methods used to establish the domain, and how they were distributed in the mail arriving at Gmail:

| SPF authentication method | Share of SPF-authenticated mail | Mechanism |
|---|---|---|
| Plain SPF | 52% | A published SPF record |
| Best-guess SPF | 38% | No SPF record, or a failing one: assume the mail is authenticated if the sending IP is in the same range as the domain's A or MX records, **or** if the reverse DNS name of the sending IP matches the domain claimed in the email |
| PTR zone | 10% | If plain and best-guess SPF both fail, and the sender is a subdomain of the DNS zone of the PTR, treat the mail as if it came from the zone itself (for example, `subdomain.example.com` from an IP with the PTR `host.example.com`) |

DomainKeys, the predecessor of DKIM, was the other signal. Gmail adopted it early, partly because SPF breaks under forwarding while a signature survives. Google's recommendation, then and now, is to **implement both**, so that if forwarding or changes to the message break one, the other still authenticates: "if both are present, it is that much stronger of an authentication signal."

Authentication rates were very different for wanted and unwanted mail, so authentication is itself a reputation signal:

| Method | Nonspam (wanted) | Spam (unwanted) |
|---|---|---|
| Both SPF and DomainKeys | 20% | 1.5% |
| SPF only | 53% | 39% |
| DomainKeys only | 1% | 0.5% |
| Not authenticated | 26% | 59% |
| **Authenticated (any)** | **about 75%** | **about 40%** |

### How the reputation number is computed

Every delivered message logs an event with its classification and authentication. Four counters accumulate for each authenticated domain:

| Variable | Meaning |
|---|---|
| `autononspam` | times mail from this sender went to the inbox automatically |
| `autospam` | times mail from this sender went to the spam folder automatically |
| `manualnonspam` | times a user hit **Not Spam** on this sender |
| `manualspam` | times a user hit **Report Spam** on this sender |

Reputation is a number in the range **0–100**: "you can think of the reputation as the probability that a given sender's mail is not spam." 0 is the most spammy, and 100 the least.

```
good       = autononspam + manualnonspam − manualspam
total      = autospam + autononspam
reputation = 100 × good / total
```

The refined version limits the manual terms, so that a small number of votes cannot dominate, and guards against division by zero (if `total = 0`, there is no reputation):

```
manualnonspam2 = min(autospam,    manualnonspam)
manualspam2    = min(autononspam, manualspam)
good2          = autononspam + manualnonspam2 − manualspam2
reputation     = 100 × good2 / total
```

Worked examples from the paper:

- `weliketospam.com` sends 100 spam messages. 60 land in the spam folder automatically and 40 get into inboxes as false negatives, which gives a reputation of **40**. When 30 users then mark the spam that was missed, the reputation drops to **10**.
- `weneverspam.com` sends 100 messages. 5 wrongly go to the spam folder as false positives, which gives a reputation of **95**. When 3 users unmark them, the reputation rises to **98**.

The reputation also works as an expected false-positive rate. For a sender with a reputation of 2, with the spam-folder threshold set below a reputation of 5, "about 2% of the time you'll have false positives on that sender's mail."

### Self-correction, vote-weighting, and recovery dynamics

- Reputations are computed **over many days**, so a solid domain "can build up a very solid reputation" and "a little blip of spam ... is absorbed as noise." The other side is that a domain that spams for a long time and later cleans up **takes a few days to recover**.
- **Both the Report Spam and the Not Spam buttons are critical.** With only spam reports as input, the reputation has no way to move back up after a filter mistake, so it cannot correct itself. Google calls having both buttons "critical for the success" of the system.
- **Not every user votes, and votes are rate-limited.** Only a subset of users, "the ones that will provide the best information", are counted, and users who never report spam or nonspam are excluded. To stop users with heavy mail from dominating, each user is limited to **one spam report per domain per hour**, that is, up to **24 votes per domain per day**.

### How the reputation is used (the three-way split)

After authenticating the sender and computing the domain reputation, Gmail acts as follows:

1. A reputation **above a threshold** sends all of the sender's mail to the **inbox**.
2. A reputation **below a threshold** sends all of it to the **spam folder**.
3. **In between (or unknown)**, the message goes, with the reputation value, to a **statistical spam filter that makes the final judgement** on content.

Crucially, **a user can override the general policy for individual sender addresses**. This is the documented reason why placement differs between users.

### What the reputation distribution looked like

- Spammy domains cluster tightly around **0**, and domains with wanted mail cluster in the **90–100** range. For legitimate bulk senders, a score **below 90** is "a reflection of less-than-ideal sending practices". Hygiene that is not quite good shows up as a middling score, not as a zero.
- Selected domain reputations (2006, illustrating the model): `aol.com` 98.5; `yahoo.com` 95.6 with SPF and 95.0 with DomainKeys (DK); `hotmail.com` 95.4; `ebay.com` 95.2 with SPF against 98.2 with DomainKeys; `earthlink.net` 93.3 with SPF against 98.0 with DK. Domains flagged as spam scored 1.5–3.3.
- The gap between SPF and DomainKeys for `ebay.com` is instructive. eBay signed only its transactional mail with DomainKeys, and transactional mail "would likely be more wanted by users than other mail," so the signed stream scored higher. This is an early illustration of **stream separation** paying off (see [advanced IP segmentation](https://emailmarketing.net/learn/ip-management/advanced-ip-segmentation) and [sending infrastructure practices](https://emailmarketing.net/learn/operations/sending-infrastructure-practices)).

### Documented failure modes and the sender rules that follow

The paper calls out these problems, and all of them are still live concerns:

- **Forwarding:** users forward mail (including spam) with tools such as `procmail` that rewrite the envelope sender. The forwarded spam then authenticates as the forwarder, and "hurts the reputation of the forwarding domain unfairly." The rule that follows is not to authenticate mail you are only forwarding unless you have filtered out the spam first, because signing it means taking responsibility for the resulting stream. [ARC](https://emailmarketing.net/learn/authentication/arc) later addressed the same problem.
- **Mailing lists:** a list domain (the paper's example is Yahoo Groups, signed as `yahoogroups.com`) can have a good overall reputation while individual groups within it send spam, so blanket trust "will allow some spam through." Users also often find **reporting a list as spam easier than unsubscribing**, which hurts the list sender's reputation "worse than it needs to be". This is an argument for a prominent unsubscribe option with little friction (see [List-Unsubscribe](https://emailmarketing.net/learn/list-management/list-unsubscribe)).

Sender policies the paper recommends:

1. Authenticate with **both** SPF and DomainKeys or DKIM (you may already have best-guess SPF implicitly).
2. **Do not authenticate mail you are only forwarding**, unless you filtered out spam first.
3. Keep spammers and zombie machines off your network, and follow good practices for bulk sending: "if a sender is doing something bad such as using single opt-in instead of double opt-in, it simply translates into a poor reputation and a greater likelihood of ending up in the spam folder." **Users enforce Gmail's bulk guidelines** through the reputation, not a rulebook.

## ML spam and phishing classification

The reputation model above answers the question "is this sender trusted?". For everything in the ambiguous middle band, the verdict on content comes from machine learning (ML) classifiers that have been rebuilt and scaled up repeatedly.

- **"The War Against Spam: A Report from the Front Line"** (Taylor, Fingal and Aberdeen, NIPS 2007 Workshop on Machine Learning in Adversarial Environments) presents Gmail's spam filtering as a real-world success story for adversarial machine learning: the overwhelming majority of spam is positively identified, with high precision (few false positives). The adversarial framing matters. Spammers adapt, so filters must retrain continuously instead of running a fixed set of rules.
- **2014: a rebuild at TensorFlow scale.** Google's rule-based filters had reached about 99% accuracy. From 2014, Google added **machine learning algorithms built on TensorFlow** that "continuously regenerate" the spam filters, finding new patterns and adapting "far quicker than previous manual systems." The reported outcome: **more than one billion Gmail users avoid spam**, at roughly **99.9%** effectiveness.
- **Claims about the security model in 2017** (Google, May 31 2017): **over 99.9% accuracy** in blocking spam and phishing, and **50–70% of all mail Gmail receives is spam**. Classifiers combine **thousands of signals for spam, malware and ransomware** with **heuristics on attachments** and **sender signatures**. Phishing defenses combine **Google Safe Browsing** with **analysis of the reputation and similarity of URLs**. A model for early phishing detection can **selectively delay** risky messages for deeper analysis, which affects **less than 0.05% of messages on average**. Warnings at click time, warnings on replies to external recipients, and defenses against ransomware and polymorphic malware complete the picture.
- **2024: large language model classifiers.** Google deployed **a large language model (LLM) trained on phishing, malware and spam** that blocks **20% more spam** and reviews **1,000× more spam reported by users each day**, together with a second "supervisor" model that evaluates **hundreds of threat signals** when a risky message is flagged. See [holiday tightening](#holiday-period-tightening) below.

What this means for senders: the classifier is a moving target, trained on everything users report. There is no content allowlist to satisfy, because the filter optimizes for what users treat as wanted. This is the mechanism behind the guidance in [content and design for deliverability](https://emailmarketing.net/learn/operations/content-and-design-for-deliverability): avoid spam patterns and optimize engagement, rather than follow a keyword checklist.

## Inbox categorization (the tabs)

Categorization is a **separate decision, made after acceptance**, from the verdict between inbox and spam. A message that has already earned the inbox is then sorted into one of five default categories (in Gmail's "Default" inbox type). **Landing in a category tab other than Primary is not a deliverability failure.** Promotions is the inbox, not spam.

| Category | Google's definition |
|---|---|
| **Primary** | "Emails from people you know and messages that don't appear in other tabs." |
| **Social** | "Messages from social networks and media-sharing sites." |
| **Promotions** | "Deals, offers, and other promotional emails." |
| **Updates** | "Automated confirmations, notifications, statements, and reminders that may not need immediate attention." |
| **Forums** | "Messages from online groups, discussion boards, and mailing lists." |

Mechanics:

- Users turn categories on or off in **Quick Settings**, then **Default inbox type**, then **Customize**. Turning categories off entirely requires switching to a different inbox layout. Because the recipient controls which tabs exist, a sender cannot assume that a given recipient even has a Promotions tab.
- **Senders cannot choose a tab in advance.** Recipients move messages between tabs by **drag and drop** (with an Undo), which teaches Gmail their preference so that it will "sort your email more accurately" over time. This is another override by the individual user.
- Asking engaged readers to drag your mail to Primary is legitimate and effective. See [content and design](https://emailmarketing.net/learn/operations/content-and-design-for-deliverability) for the content factors (personalization, a conversational style, a limited number of images and links) that push mail toward Primary rather than Promotions.

### Category prediction as an ML problem

Source: **"Email Category Prediction"**, by Zhang, Garcia Pueyo, Wendt, Najork and Broder (Google, WWW 2017 Companion). The key points for senders:

- **About 90% of consumer email is generated by machines.** Categorization targets templated mail (shopping receipts, promotional campaigns, booking confirmations, bill reminders), not mail between people.
- The approach uses **template discovery** and the **"causal threads"** of email sequences, because legitimate machine mail follows predictable patterns (an order, then a shipping confirmation, then a delivery alert). Predictable sequences that match their templates classify cleanly; erratic structure does not.
- Comparison of models: **neural networks (multilayer perceptrons, or MLPs, and long short-term memory networks, or LSTMs) far outperform** a Markov chain baseline, with **LSTMs slightly ahead of MLPs**. The applications named include filling calendars, shipment alerts, targeted advertising and spam detection, so categorization feeds back into filtering.

## Per-user personalization signals

Beyond the vote on spam and nonspam, Gmail runs an **importance** model for each user, and spam controls for each user. These are the mechanisms that make placement individual.

**Importance markers** (help topic "Importance markers in Gmail"). Gmail predicts which messages matter to each user from:

- how often the user communicates with the sender (frequency);
- which messages the user opens and replies to;
- **keywords in the emails the user usually reads**;
- interactions: starring, archiving, deleting.

Users see a yellow importance marker, can correct it (which trains the model), and can find flagged mail with `is:important`. A user can turn off predictive marking ("Don't use my past actions to predict which messages are important") or hide the markers entirely. These settings are changed in the browser but apply across the apps.

**Spam controls that feed personalization** (help topic on how Gmail sorts spam):

| User action | Effect on future placement |
|---|---|
| **Report spam** | The message is added to Spam; "as you report more spam, Gmail identifies similar emails as spam more efficiently." Google receives a copy and may analyze it to protect all users. |
| **Not spam** | Recovers the message and teaches Gmail that this sender is wanted (the `manualnonspam` vote). |
| **Block sender** | "Even when you remove their emails from Spam, Gmail still automatically identifies their emails as spam." A hard override for that user. |
| **Add sender to Google Contacts** | "Gmail stops sending their messages to Spam." A hard allow for that user. |
| **Unsubscribe** | Offered inline for senders the user opted in to. A less damaging exit than Report Spam. |
| **Filters** | Rules the user creates can label or prioritize a sender's mail. |

The common thread: **the recipient's own history can override the sender's overall reputation in both directions.** A sender with a strong domain reputation still lands in spam for any recipient who blocked them or repeatedly reported them. A sender with a weak reputation reaches the inbox for recipients who added them to Contacts or clicked Not Spam.

## Holiday-period tightening

Google publicly tightens its filters around periods of heavy fraud. For the **2024 holiday season** (blog post, Dec 2024):

- Gmail blocks **more than 99.9% of spam, phishing and malware**.
- The newly deployed **LLM classifier** and **supervisor model** (see [ML classification](#ml-spam-and-phishing-classification)) launched **before Black Friday**.
- Result: users reported **35% fewer scams** reaching inboxes in the first month of the holiday season, compared with the previous year.
- The holiday scam patterns Gmail strengthened its defenses against: **invoice scams** (fake invoices that invite the recipient to phone a number to dispute them), **celebrity scams** (impersonation or false endorsement), and **extortion scams** (threats that cite a home address or a personal detail).

What this means for senders: filter thresholds are **not constant throughout the year**. A stream with a borderline reputation that reaches the inbox in a quiet month can tip into spam during a tightening period. Build a margin of reputation before Q4, and do not launch new streams or re-warm into the holiday peak.

## Sender takeaways: what actually moves Gmail placement

1. **Authenticate so that your mail can be scored, then build the domain's reputation.** Use SPF and DKIM aligned to the [From domain](https://emailmarketing.net/learn/providers/gmail-sender-requirements). Reputation attaches to that authenticated domain, not to your IP. Unauthenticated mail mostly cannot be scored, and it correlates with spam (about 60% of spam is unauthenticated, compared with about 26% of wanted mail).
2. **The one metric that directly drives the model is the user's spam vote.** Report Spam pushes reputation down; Not Spam and adding the sender to Contacts push it up. Keep the spam rate in [Postmaster Tools](https://emailmarketing.net/learn/postmaster-tools/google-postmaster-tools) below **0.10%**, and never let it reach **0.30%**. Those are the [sender requirement](https://emailmarketing.net/learn/providers/gmail-sender-requirements) thresholds that correspond to this vote.
3. **Reputation changes slowly in both directions.** Days of consistent behavior build it, a single bad blip is absorbed as noise, and recovery after sustained abuse "takes a few days." There is no overnight fix; see [reputation incident recovery](https://emailmarketing.net/learn/operations/reputation-incident-recovery).
4. **Separate your streams.** Even in 2006, transactional mail signed separately from marketing mail scored higher, because recipients want the two streams to different degrees. Segment by [sending domain or subdomain](https://emailmarketing.net/learn/operations/sending-infrastructure-practices) and by content type.
5. **Categorization is not deliverability.** Promotions is the inbox. Improve tab placement with content (personalization, a conversational tone, fewer images and links), not by trying to game the classifier: you cannot choose a tab in advance.
6. **Personalization for each user defeats one-size-fits-all diagnosis.** Importance markers, blocks, Contacts and overrides for individual addresses all apply recipient by recipient, so "where does my mail land at Gmail?" has no single answer, only a distribution. Seed inbox tests sample a handful of fresh accounts with no history of their own, and therefore **cannot represent** what an engaged (or disengaged) real recipient sees. Trust the trends in aggregate reputation and spam rate in Postmaster Tools over any screenshot of a single inbox (see [tracking and measurement distortion](https://emailmarketing.net/learn/operations/tracking-and-measurement-distortion)).

> Note on sources: Gmail help answer **6596** ("Gmail inbox tabs and categories") overlaps with the Gmail tabs guidance in [Content and Design](https://emailmarketing.net/learn/operations/content-and-design-for-deliverability) and with the categorization material above. It was not used separately, because it adds nothing beyond answer 3094499, which is already cited.

## Related articles

- [Gmail Sender Requirements](https://emailmarketing.net/learn/providers/gmail-sender-requirements), the checkable rules and spam rate thresholds that feed this model
- [Gmail SMTP Troubleshooting](https://emailmarketing.net/learn/providers/gmail-troubleshooting), including the messages for reputation blocks
- [Google Postmaster Tools](https://emailmarketing.net/learn/postmaster-tools/google-postmaster-tools)
- [Content and Design for Deliverability](https://emailmarketing.net/learn/operations/content-and-design-for-deliverability)
- [Foundations of Email Deliverability](https://emailmarketing.net/learn/foundations/foundations-of-email-deliverability)
- [ARC](https://emailmarketing.net/learn/authentication/arc), the later fix for the forwarding problem described in the 2006 paper
