How to read aytm's Q1 2026 data quality benchmark figures

aytm logo icon
Posted Aug 24, 2026
Trevor Brown

A removal rate means nothing until you know what it was measured against. Here's every figure in our Q1 benchmark with its base attached, and why the base is the whole story.

Every research platform says its data is clean. The interesting question is what "clean" measures to, and against what.

That's the question aytm's Q1 2026 data quality benchmark report was built to answer. We reviewed 1.1 million survey attempts from last quarter and published what we found, with each figure carrying its own denominator and time period. Here's how to read it.

Start with the denominator

A removal rate on its own tells you very little. Removed from what? Over what period? Two vendors can report the same percentage and mean entirely different things, because one measured against every entry that touched the survey and the other measured against qualified completes.

aytm reports post-survey cleanout at 5.4% of qualified completes for Q1 2026. Qualified completes is the harder base: it's the population that made it all the way through screening and finished the survey, which means the percentage isn't diluted by everyone who never belonged in the study. Every figure in the report ships this way, with the base stated next to it. The denominators differ across metrics, deliberately, because the questions differ, so we never roll them under one number.

Quality is built in across four layers

Data quality isn't a step at the end of fieldwork. aytm's program runs in four layers, each with a defined job.

  • Prevent keeps the survey instrument clean before launch. Survey-design discipline and length control are the cheapest quality work available, because a respondent who stays engaged gives you better data than one who's grinding through question forty. In Q1 2026 that discipline held the median length of interview to 5m 44s, and the abandon rate to 2.6% of survey attempts, aytm's lowest across six quarters.
  • Protect checks identity and sourcing at the entry point, before sample reaches a study. PaidViewpoint is aytm's proprietary, referral-built panel and the identity source of truth, with identity verified at registration and the panel refreshed every quarter. Outside sample is checked against those verified identity records before fieldwork begins, and Traffic Sentinel filters suspicious traffic at the source. In a single Q1 month, third-party duplicate attempts peaked at 46% of respondents in that outside sample, removed before they reached a survey. That figure measures what the screening catches in outside sources, so no duplicate count is claimed for PaidViewpoint itself. It's the clearest argument for owning the panel: you can only check sample against verified identity records if you hold them.
  • Purify runs behavioral detection during and after fieldwork. Data Centrifuge analyzes every respondent across parallel behavioral dimensions at the individual response level, including signals for AI-generated open ends. Post-survey review layers human judgment on top. Together they produced the 5.4% cleanout figure.
  • Prove ships a per-study data quality record, so every figure can be checked against the study it came from rather than an aggregate rollup.

The figure we publish on purpose

One number in the report went up, and we published it anyway.

aytm checks whether a verified panelist's profile matches the demographics they give inside a survey, then publishes how often it doesn't. In Q1 2026 that was 1.6% of verified respondents, up from 0.6% a year ago. The rise is the point: verification coverage widened, so a stricter test found more.

That measurement has a structural precondition. You can only compare a survey answer against a verified profile if the verified profile exists, which means owning the panel and keeping it current. That's why the metric is rare, and it's why we treat publishing it as part of the standard rather than a disclosure risk.

Why most bad data survives cleaning

Greenbook and Rep Data audited 4.1 billion survey attempts in 2025. Roughly a third were fraudulent. About a quarter were inattentive. Close to 70% of that bad data slipped past standard cleaning.

That last figure is the whole problem. If most contamination is invisible to cleaning, a rate with no disclosed base tells you nothing about whether anything was caught. A number with a stated denominator, a stated period, and an audit trail is a measurement. Every figure in our report is open to a methodology audit on request, and the report closes on a companion piece for procurement and insights teams: seven questions that separate measured data quality from asserted data quality, with our own answers in full.

Featured Stories

New posts in your inbox