Documentation

Data quality methodology

This page documents the respondent validation methodology applied to every MX8 Labs study. It is the methodology companion to two existing references: Fraud prevention measures, which lists the public signal categories, and IP Address Hygiene and Exposure Matching, which documents the pipeline in the context of ad-exposure work.

Scope

Quality control runs in four layers, in this order:

  1. Deduplication - is this the same person or device we have already recorded on this study?
  2. Fraud and bot screening - is this a real human, answering in good faith, in the environment they appear to be in?
  3. Identity verification - where the study demands it, can this respondent prove they control a real phone number?
  4. In-survey attention and consistency checks - is this respondent engaging with the questions?
  5. In-field monitoring - is any individual respondent source degrading while the study runs?

Layers one and two run automatically on every study without configuration. Layer three is added where the sensitivity of the study justifies the friction. Layer four is designed into the questionnaire. Layer five is continuous and visible while the study is live.

In the production data summarized in our January 2026 survey fraud analysis, this pipeline excluded 10 to 20 percent of incoming respondents before they reached analysis. That figure describes the observed period, not a guaranteed rejection rate for every study.

The design principle is that fraud is assumed rather than suspected. Screening is applied to every dataset, not triggered when something looks wrong, because contemporary survey fraud is specifically built to not look wrong: plausible completion times, coherent open ends, passed attention checks.

Layer 1: Deduplication

Deduplication uses three concurrent identification signals. A respondent is treated as unique only when all three confirm they have not already been recorded on the study.

SignalRoleFailure mode it covers
First-party session cookiePrimary key for browser-based respondents; identifies repeat visits within and across sessionsThe straightforward repeat entrant
IP addressSecondary signal where cookies are cleared or unavailable; also flags high-density traffic from a shared networkCookie clearing; coordinated entry from one network
Device fingerprintPersistent device identifier that survives cookie deletion and incognito browsingSophisticated duplicates that defeat the first two

The layering matters because each signal is individually defeatable. Cookies can be cleared, IP addresses rotate legitimately on mobile networks and are shared behind household NAT, and fingerprints drift with browser and operating system updates. Requiring agreement across all three is what makes the result resilient to any single signal being spoofed or degraded.

Device intelligence analyzes multiple browser attributes to help identify duplicate and suspicious devices. It is used alongside cookie and IP checks rather than as a standalone decision. The public categories of behavioral, network, bot, device, and mobile checks are listed in Fraud prevention measures.

Layer 2: Fraud and bot screening

Deduplication establishes uniqueness. It does not establish legitimacy: a unique respondent can still be a bot, an automated agent, or a professional fraud operation.

Screening evaluates five families of signal in real time. The full enumerated list is in Fraud prevention measures; in summary:

  • Behavioral - incognito mode, browser tampering and spoofed identity, anti-detect and anti-fingerprinting browsers, cookie tampering.
  • Network - IP geolocation mismatch against browser-reported timezone as a VPN indicator, proxy and data-center detection, IP rotation within a session, duplicate IP density, IP blocklist matching against known botnets and flagged actors.
  • Bot and automation - automated and instrumented browser detection, Android emulator detection, virtual machine detection.
  • Device - raw attribute analysis, remote-access tooling such as RDP or TeamViewer, velocity flagging for devices showing implausible activity across short intervals.
  • Mobile - rooted Android and jailbroken iOS, cloned application detection, Frida instrumentation, recent factory reset.

No single signal is decisive. Each contributes to a weighted composite suspect score, and respondents exceeding the study threshold are excluded automatically. This is the correct structure for the current threat model: individually, incognito browsing or a VPN is unremarkable, and rule-based exclusion on any one of them would reject large numbers of legitimate respondents. It is the combination that carries the signal.

Layer 3: Identity verification

Deduplication and fraud screening operate on signals the respondent emits. Some studies need something stronger: proof that the respondent controls a real, unique phone number.

SMS verification adds that step inside the questionnaire. The respondent enters a phone number, receives a one-time code by SMS, and must enter it correctly before continuing. Because a phone number is costly to acquire at scale and difficult to reuse undetected, this is a materially harder control to defeat than any browser-side signal, and it also serves as a deduplication key in its own right.

The added friction is justified on:

  • High-sensitivity or confidential studies, including anything behind gated stimulus.
  • Studies where duplicate participation would be materially damaging to the result.
  • Studies where a meaningful incentive creates an incentive to game entry.

It is not applied by default, because it raises break-off and narrows the reachable population: respondents without a mobile number, or unwilling to share one, are excluded. That is a coverage cost, and it should be a deliberate trade rather than a standing setting. See Verifying phone numbers.

Where a study is fielded over SMS or voice rather than the web, the phone number is inherently the identifier, and these respondent sources are treated as trusted for fingerprint-based checks accordingly.

Layer 4: In-survey attention and consistency

Automated screening addresses who the respondent is. Questionnaire-level checks address whether they are engaging.

Applied by default, with no configuration:

  • Straight-lining detection on grid questions, which prompts the respondent to review when every row carries an identical answer.
  • Exclusion on duplicate IP address, URL manipulation, and cookie mismatch.

Available to design into the questionnaire:

  • Overt instructional checks - an explicit instruction the respondent must follow, applied either as a warn-and-retry validator or as an immediate termination.
  • Covert consistency checks - logically contradictory answer patterns across questions, which catch respondents who pass overt checks by reading carefully but answer without engaging.
  • Speed traps - elapsed time between any two answered questions is available in survey logic, and a per-question timer can delay advancement, so dwell-time thresholds can be set per question rather than only across the whole interview.
  • Escalation - a running count of failed validations, allowing termination after repeated failures rather than on a single slip.
  • Open-end validation - AI-assisted checking of free-text responses for coherence and relevance to the question asked. See AI-enhanced validation for text questions.

Guidance we give on design: one or two overt checks in a ten to fifteen minute interview is normally sufficient, options in check questions should be randomized so that position cannot be learned, and overt and covert checks should be mixed. Overloading a questionnaire with checks degrades the experience for genuine respondents and raises break-off, which is itself a quality problem.

Worked examples are in Using "Gotcha" Questions in MX8 Labs to Improve Survey Quality.

Layer 5: In-field monitoring

Quality is monitored per source while the study runs, not audited afterward. Live incidence is broken into completed, in progress, terminated, and poor quality for each source separately, so a rise in rejected respondents can be attributed to a specific respondent source and acted on before further spend. See Field reports and the sourcing detail in Respondent sourcing methodology.

Defaults and configurability

Quality controls are on by default and can be relaxed for a configured respondent source. The active controls therefore form part of the study configuration and should be checked when reviewing a dataset.

  • Bot detection is on by default. It can be disabled at the survey level where every entrant is known to be human.
  • The "trust all respondents" toggle is off by default and is set per respondent source. When enabled, respondents from that source are not marked poor quality on fingerprint bot signals or IP and cookie mismatches. Its purpose is to prevent false positives on sources where the standard checks misfire — an interviewer-administered source, or a controlled first-party list.
  • Some source types are inherently trusted and do not present the toggle at all: synthetic, profile synthetic, third-party import, Twilio voice and text, and test links.

Account security is a separate control

Respondent validation and platform access control answer different questions, and a review should keep them apart.

The layers above govern who enters a dataset. Access to the data once it exists is governed separately: platform users authenticate with multi-factor authentication, permissions are configured per user per account through granular roles covering survey editing, reporting, fielding, media, going live, and administration, and only users holding the relevant permission can put a survey live and incur cost. Studies can be hosted in a specific region where data residency requires it.

See Understanding roles and permissions and Local data residency.

What this methodology does not do

It does not correct for coverage or non-response bias. Screening removes invalid respondents. It does not make the surviving sample representative. Composition is addressed by calibration, and the limits of that framework are set out in Weighting methodology.

It does not apply to synthetic respondents. Synthetic Twins and Synthetic Profiles are generated rather than recruited, so human-quality screening is not meaningful for them. They are labeled distinctly throughout the platform, and their inferential treatment is documented in Synthetic data and statistical inference.

It does not claim a zero false-negative rate. The 10 to 20 percent figure above describes exclusions during the published observation period, not a fixed platform rate or a guarantee of what remains. Fraud techniques change, and no screening system detects every invalid respondent.

Privacy

Device intelligence is implemented through standard browser APIs that trigger no permission prompts and do not alter the respondent experience. No personally identifiable information is collected or stored in the fingerprinting process: device signals are hashed into anonymous identifiers that cannot be reversed to identify an individual. Fingerprinting is used exclusively for security and fraud prevention. Data handling is designed to comply with GDPR and CCPA, and studies can be hosted in a specific region where required. MX8 Labs' current certification status is published at trust.mx8labs.com. See Privacy Compliance and Local data residency.