Documentation

Understanding Stat Testing in MX8 Labs Reports

MX8 Labs cross-tab reports can apply statistical tests to table cells and use color to show the result of each comparison:

  • Green: Significantly higher result
  • Red: Significantly lower result
  • Blue: Normal result, no significant difference
  • Transparent: Result is below the reporting threshold (default minimum: 20 respondents)

By default, MX8 Labs uses a residual t-test at a 95% confidence level and does not color cells with fewer than 20 respondents. The residual test identifies cells that are significantly higher or lower than expected from the row and column distribution.

Reports also support row-based and column-based comparisons. The appropriate mode depends on which cells the analysis intends to compare.

Choosing a different configuration
  • To conduct more granular analyses within smaller subsets of data.
  • If a different confidence level (e.g., 90% or 99%) is required to align with specific research standards or regulatory guidelines.
  • To perform comparative analyses focused explicitly on rows or columns rather than overall data distribution.
Adjusting the confidence level

Confidence testing identifies results that are statistically significant compared to other responses. The default confidence level in MX8 Labs is 95%. At that level, a difference is marked only if it is large enough that a difference that size would turn up less than 5% of the time if there were really no difference in the population.

When to change
  • When a higher confidence level (e.g., 99%) is necessary for stricter confidence in findings.
  • When lower confidence levels (e.g., 90%) can highlight potential trends or insights for exploratory analysis.
Using a residual t-test

This approach identifies cells within a report table that are significantly higher or lower than expected from the distribution between rows and columns.

When to use
  • This is for standard exploratory analysis or to identify unexpected anomalies in the data.
  • If a more focused comparative analysis across rows or columns is required, consider using a row - or column-based t-test.
Using a row-based t-test

This test compares each cell against other cells in the same row. Cells marked with labels indicate those cells in the same row that are statistically lower.

When to use
  • When analyzing category performance within individual rows.
  • Applicable when assessing item-level differences clearly across horizontal data.
Using a column-based t-test

This test compares each cell against other cells in the same column. Labels indicate that cells in the same column are statistically lower than those highlighted.

When to use
  • For vertical comparisons across segments or groups.
  • When the intended comparison is between cells in the same column.
Stat testing for median aggregations

Median and Selected count median use respondent-level bootstrap comparisons rather than the analytical standard errors used for means and percentages. The platform repeatedly resamples respondents, preserves their reporting weights and shared comparison context, and recalculates each median to estimate its uncertainty.

For a column-based comparison, the platform evaluates how consistently one segment's median exceeds another segment's median across the paired bootstrap samples. For a residual test, it compares each median with the group of peer medians and uses the variation across bootstrap samples to determine whether the difference is unusually high or low. The report's confidence level controls the threshold for both modes, and the normal minimum-respondent threshold still determines whether a result can be colored.

Row-based t-tests are not available for median rows. Use a column-based or residual test when a report includes Median or Selected count median. Although the report editor retains the familiar statistical-test names, median significance is based on the bootstrap procedure described here rather than an ordinary t-test calculation.

Conjoint and MaxDiff uncertainty

Discrete-choice outputs are calculated from retained posterior draws. Attribute importance, simulated share, above-anchor probability and marginal purchase lift are transformed within each respondent and draw before weighted aggregation. The reported standard error combines variation across draw means with weighted sampling variance using effective sample size. Draws do not increase the respondent base.

Probability above anchor is the weighted proportion of draws with item utility above zero. It is not a p-value, a purchase rate, or a test that one demographic group differs from another. Marginal Purchase Lift is a modeled change in purchase probability after replacing an attribute level; it is distinct from the matched-group Lift report described below.

See Utility and simulated share methodology for the likelihoods, formulas and uncertainty calculation.

Matched-group Lift reports

Lift reports use a separate procedure: Fisher's exact test on target and control outcome counts in each matching iteration, followed by averaging the p-values. P-value threshold filters that mean; a result is retained when its mean p-value is at or below the configured value.

A threshold of 0.95 is an exploratory filter, not a 95% confidence setting. The usual 95% confidence convention corresponds to a 0.05 significance level. The average of matching-iteration p-values is not a pooled test or a correction for multiple comparisons. See Lift measurement methodology and the Lift report guide for interpretation and a worked example.

Selecting a 0.05 filter alone does not establish a confirmatory finding or a pooled 5% false-positive rate. Confirmatory inference needs a method appropriate to the study design, matching procedure and planned comparisons. Specify headline outcomes and X+ thresholds before inspecting results, disclose exploratory comparisons and report negative and inconclusive findings alongside positive ones.

Changing the reporting threshold

MX8 Labs applies a reporting threshold to each response. If the number of respondents for a specific cell falls below this threshold, the results are not colored in the reports.

This is applied at the level of the individual response to the question, not at the question itself, so that we can filter out significantly low responses to an asymmetric question. For example, let's ask people about their ethnicity. We're more likely to get statistically significant results for common ethnicities. Still, we don't want to filter out ethnicity options entirely because of the small number of respondents in a specific minority.

The default threshold for reporting is 20 respondents.

When to change
  • Increase the threshold when cells below a larger minimum base should be left unmarked.
  • Lower the threshold when smaller cells must remain visible, while accounting for their lower precision.