This engine was tuned against five invented products. First contact with real catalogues showed a 54% false-positive rate. What came after is most of the work, and it is the only thing that makes a report about a stranger's feed worth anything.
- A real corpus, frozen
- 4830 products from six real feeds and 4188 rows from five Shopify catalogues, sha256-pinned. Every check runs against all of them and its fire rate is recorded.
- A fire-rate ceiling
- If a check fires on more than 20% of the corpus, we stop and read the evidence. At 100% it is a bug, with no exceptions: the two worst false positives this engine ever shipped fired on 4319 of 4319 items. That cleanup took the total from 19,209 findings down to 4,505.
- A control group
- The fixture contains correct products. Zero checks firing on them is a build condition, not a goal: if one fires, the build fails.
- Four network states, not two
- A URL can be verified, not found, unverifiable (timeout, 403, DNS) or never attempted. Only the second can become a finding. That fourth state exists because an earlier version told a merchant “we tried to verify these and could not” about 3987 URLs it had never requested. A false claim about your feed is bad; a false claim about our own work is worse, because you cannot check it.
- Severity measured, not assumed
- 16 checks have their severity confirmed against Google's actual response. The rest are still hypotheses, and the report does not present them as anything else. It also states what it did not look at, and why.
And we never talk about money. Magnitude and mechanism yes; a figure in your currency would be invented, and an invented figure burns credibility faster than a false positive.