If your Google Postmaster domain reputation will not line up with anything else you can see, the usual cause is that you are comparing three separate measurements and treating them as one signal. Postmaster scores your domain and your sending IPs independently, it withholds data when your daily Gmail volume falls under its threshold, and its spam rate uses a denominator you do not control and cannot recreate from your ESP. None of that indicates a problem in your account. Once you read each number for what it actually covers, the contradictions mostly disappear and you are left with two or three numbers worth acting on.
Key takeaways
- Domain reputation and IP reputation are distinct scores. On a shared ESP pool, the IP score is largely not yours to move.
- Reputation is reported as four coarse buckets, so a genuine slide can stay invisible for days and then appear as a cliff.
- Grey or missing panels almost always mean low volume, not a penalty. Postmaster needs consistent daily volume to Gmail before it populates anything.
- Spam rate counts user complaints against inboxed mail, so mail that Gmail already filtered is excluded from the denominator and the ratio can rise while your sending is unchanged.
- Your ESP complaint rate and Postmaster spam rate measure different populations. Expect a gap and investigate only when the shape of the two curves diverges.
- Treat spam rate as the operating metric and the reputation bucket as lagging confirmation.
What Postmaster is actually measuring
Postmaster Tools reports on mail that Gmail accepted from you, broken into daily buckets, with separate dashboards for domain reputation, IP reputation, spam rate, authentication results, encryption, delivery errors and feedback loop identifiers. Google documents the dashboards and the High, Medium, Low and Bad reputation buckets in its Postmaster Tools dashboards help. Two properties of that model cause most of the confusion. First, everything is scoped to Gmail and to the identifier you verified, so it says nothing about Outlook, Yahoo or corporate filters. Second, reputation is a bucket, not a score, which means you cannot see movement inside a bucket.
| Number | What it measures | What it does not tell you |
|---|---|---|
| Domain reputation | Gmail's trust in the authenticated domain across recent sending | Which campaign, segment or subdomain caused a change |
| IP reputation | Gmail's trust in the sending IPs presenting your mail | Anything you can fix, if you sit on a shared ESP pool |
| Spam rate | User spam reports divided by inboxed mail for that day | How much mail was filtered before a user could see it |
| Delivery errors | Share of attempts rejected or temporarily failed, by reason | Whether accepted mail reached the inbox or the spam folder |
Why the reputation panel is grey or empty
Grey bars and empty charts are a volume problem in nearly every account I have looked at. Google populates these dashboards only when your daily volume to Gmail is large enough to be statistically meaningful, and in practice that means a few hundred messages a day before anything appears and a good deal more before the reputation bucket behaves smoothly rather than flipping between states. A store sending two campaigns a week to a 40,000 person list will see data on send days and nothing in between, which looks like a reputation that keeps vanishing.
The operational consequence matters more than the cosmetic one. If you are below the threshold, Postmaster cannot be your feedback mechanism during a domain warmup, because the signal only arrives once you are already at volume. Plan the ramp on your own engagement and bounce data instead, which is the approach I describe in warming a new sending domain.
The spam rate denominator you do not control
Spam rate is the metric to run your sending on, and it is also the one most often misread. It is a ratio of user spam reports to mail that reached Gmail inboxes on that day. Mail that Gmail routed straight to the spam folder is not in the denominator, because nobody had the chance to complain about it in the inbox. That creates a counterintuitive pattern: as filtering tightens against you, your inboxed volume shrinks, and the same absolute number of complaints produces a higher percentage. The number climbs while your list, content and cadence are unchanged.
The thresholds are published. Google's email sender guidelines ask bulk senders to keep the reported spam rate below 0.10% and to avoid ever reaching 0.30%. Read those as operating limits on a rolling basis, not as a single day tolerance. One promotional send that touches 0.12% on a day where you also mailed a dormant cohort is a signal to check the cohort. Three consecutive days above 0.10% is a sending problem that needs the send paused, not explained.
Why your ESP complaint rate will never match
Your ESP reports complaints it received through feedback loops and one-click unsubscribe signals, divided by messages it delivered, across every mailbox provider in the send. Postmaster reports Gmail user reports divided by Gmail inboxed mail. The populations differ, the denominators differ, and the attribution windows differ because complaints arrive over days while your ESP files them against the original send. Two numbers built that differently should not agree.
What you compare is shape, not level. Plot both weekly and look for divergence: if the ESP line is flat and Postmaster climbs, your Gmail cohort is behaving differently from the rest of the list, which usually points at a specific segment or a flow that only Gmail addresses populate. I walk through that reconciliation in more depth in diagnosing a Klaviyo versus Gmail complaint rate gap, and the segment-level hunt in finding the segment burning your domain.
Domain, subdomain and IP are three reputations
Postmaster tracks reputation against the domain that authenticated the message, which in practice is the DKIM d= domain and the SPF domain. Subdomains therefore accumulate their own history, and you have to add each sending subdomain to Postmaster separately or you will simply see no data for it. Operators who moved marketing mail to send.brand.com and left transactional mail on the root domain often have one panel showing Medium and another showing High, and conclude the tool is broken. It is reporting two different identities correctly. The reasoning behind that split is in marketing subdomain or root domain.
IP reputation is the score you are least able to influence. On a shared ESP pool your mail leaves on addresses shared with hundreds of other senders, so a Low IP reputation alongside a High domain reputation is common and is not a defect in your program. It is also the strongest argument for reading domain reputation as your accountable metric. If IP reputation is persistently poor and your volume justifies it, that becomes an infrastructure decision rather than a content one, covered in dedicated versus shared IP.
A reading procedure that holds up
Stop looking at the dashboard daily and build a weekly review instead. The order matters, because each step narrows the next.
- Pull the data, do not screenshot it. The Postmaster Tools API returns daily traffic stats per domain, which lets you keep history past the dashboard window and join it to your own send log.
- Chart spam rate as a seven day trailing average next to daily inboxed volume. The two together explain most single day spikes without further work.
- Join every spike day to what you sent. Campaign, segment, flow, and whether that day included a reactivation or a suppressed-cohort test.
- Check authentication and delivery errors before reputation. A rise in temporary failures or an authentication dip is a concrete cause you can fix; the reputation bucket is downstream of it.
- Read the reputation bucket last, as confirmation that a trend you already identified is being priced in by Gmail.
The discipline here is refusing to act on the bucket alone. A bucket change with no corresponding movement in spam rate, volume or errors gives you nothing to change.
Where to look when Postmaster is not enough
Postmaster covers one provider, so build two more sources before you draw conclusions about your program. DMARC aggregate reports, defined in RFC 7489, give you per-provider volume and disposition from every receiver that sends reports, which is the only cross-provider view of authentication that you get for free. For Microsoft traffic, Smart Network Data Services exposes IP-level complaint and trap data for the addresses you send from. Seeded inbox placement testing fills the last gap, since none of these tools tell you where accepted mail landed.
If DMARC reports themselves look inconsistent between receivers, that is a separate diagnostic path, and I covered it in DMARC alignment failing per recipient.
Trade-offs and what I would do
There is a genuine tension between responsiveness and stability. Reacting to the reputation bucket gives you the earliest possible warning at the cost of frequent false alarms and a program that gets re-engineered every time a coarse indicator wobbles. Waiting for a confirmed spam rate trend gives you fewer, better decisions at the cost of a few days of exposure. I take the second position for any account sending under roughly a million messages a month, because at that volume the bucket is noisy enough that acting on it costs more revenue in cancelled sends than it protects.
Concretely: set your own alert at 0.08% seven day trailing spam rate, well under Google's 0.10% line, and treat that as the moment to investigate rather than the moment to pause. Pause or cut the send when two consecutive days clear 0.20%. Ignore IP reputation on a shared pool unless delivery errors move with it. And if your panels are grey, resist the temptation to read that as good news; it means you have no Gmail signal at all, and your list hygiene work has to stand on your own engagement data instead, which is the situation I describe in cleaning an inherited list.
The honest summary is that Postmaster is a coarse, single-provider, lagging instrument that happens to contain one excellent metric. Use the spam rate, use the API to keep history, and stop trying to make the reputation bucket agree with a number it was never computed from.