InboxRatio

Sources: what we trust, for what, and where we've been burned

Every fact on this site has a defined class of acceptable source, and the classes are ranked. When two sources disagree, the higher rank wins and the disagreement gets noted. This page publishes that hierarchy — and, less comfortably, the list of ways we know data collection goes wrong.

The hierarchy

  1. Our own seed-list placement tests — the only valid source for any deliverability number on this site. No vendor claim, aggregator average or forum thread can put a digit into a Deliverability Score.
  2. The vendor's current public pages — canonical for pricing, features, limits and stated policies. Never canonical for claims about the vendor's own deliverability or reputation. Re-checked within 90 days of any page that cites them.
  3. Our in-app verification — what a feature actually does at the tier we name, checked from inside a real account. Marketing pages and app reality diverge often enough that this is its own class.
  4. Third-party review aggregators (G2, Capterra, Trustpilot) — canonical for user sentiment and recurring themes, read at least 20 recent reviews deep. Their star averages never enter any score.
  5. Independent deliverability mentions — expert studies, practitioner blogs, forum and Reddit threads. These appear in each review's "what others say" section as dated, linked context. Context only.
  6. Demand data (search-volume APIs) — used internally to decide coverage priorities, displayed only as coarse popularity indicators.
  7. Corporate registries, funding databases, incident disclosures — for ownership, stability and security history.

Never valid, in any section: model training data or memory, "we recall reading", unsourced wiki content, sales-rep verbal claims, a vendor's own "99% deliverability" marketing figure quoted as measurement, or "typical for this category" — a heuristic is not a fact.

Cross-verification rules

Single-source claims don't ship. Pricing needs the vendor page plus our market dataset, volume-normalized before comparison. Feature availability needs the marketing page plus an in-app check. Sending limits need the pricing page, the docs and the app — all three routinely disagree, and when they do we report the strictest. Competitor claims inside a review expire after 30 days; everything else after 90. Any cross-source discrepancy above 10% gets flagged and resolved before publication.

Our documented failure modes

This list is the part most sites keep private. We publish it because a methodology that has never caught itself failing hasn't been used. Each entry is now a standing check in our process:

  • Vendor deliverability claims are marketing, not measurement. "99% delivery" counts accepted-by-server, not inbox placement. Never cited as data.
  • A single send is noise. Placement shifts between days; we measure over multiple send days before publishing, and treat one-day results as invalid.
  • New-account throttling. Young accounts get extra filtering. Test accounts must be warmed identically or the comparison is void.
  • Shared-IP pool luck. A result can reflect a pool neighbor's behavior, not the platform. Pool type is recorded with every test.
  • Tabs are not spam. Gmail's Promotions tab is a delivered message. We count inbox, tabs, spam, missing and bounces as five separate outcomes, because collapsing them flatters or damns a service unfairly.
  • ESP pricing scales on two axes — contacts and sends. Comparing prices without fixing a volume produces nonsense; we normalize before any ratio.
  • Marketing pages hide tier gates. Features get verified inside the app at the tier we actually cite.
  • Memory inflates counts. Review counts recalled rather than re-read have been wrong by an order of magnitude. Aggregator pages are re-read every time.
  • Aggregators disagree with each other. When ratings diverge sharply we publish all of them rather than silently averaging.
  • Small brands return no per-country demand data. Reported as "below measurable threshold", never as zero.
  • Vendors rename mid-cycle. Two catalog brands rebranded during our first research pass alone. Proper nouns and domains are re-verified every cycle.
  • Pricing sliders can lie to scripts. One vendor's volume slider moved visually when set programmatically but kept showing prices for the old volume — a 50% understatement a naive scraper would have published. Volume selectors get driven with real input events and spot-checked by hand.

When we hit a new failure mode, it gets added here — the date of the last addition is this page's updated date.

Refresh cadence

Deliverability tests: every 6 months for fully-reviewed services, 9 months for the second tier, plus immediate re-tests on major provider policy changes. Pricing: 90 or 180 days by tier. Features: re-verified with each review refresh. Aggregator readings: on a 0.3-star or 20%-count move. Anything past its cadence shows its age on the page rather than pretending to be current.

How the scores are built from these sources: how we test email deliverability. What every page must clear before shipping: review standards.