Documentation

Reporting lift against control groups

Lift reports compare survey outcomes between a target group and a matched control group. For a campaign study, the target group might be people who saw an ad, with similar people who did not see it forming the control group.

Use the report to compare outcomes such as awareness, recall or purchase intent. The strength of a causal claim depends on the study design: random assignment supports a different conclusion from observed exposure. Matching accounts for the characteristics you choose, while other differences can remain. See Lift measurement methodology.

Plan the study before fielding

Use rating scales for attitudinal outcomes such as consideration, favorability and purchase intent so you can report the share scoring X or higher (X+). For example, ratings 4 and 5 on a five-point scale form a 4+ outcome. Keep factual purchase questions, unaided awareness and brand selection in formats appropriate to what they measure; not every question should become a rating scale.

Use the same question wording, scale anchors and numeric coding for both groups. Keep higher values consistently more positive for X+ reporting, and define how missing answers and Don't know responses enter the denominator. Do not code Don't know as a high numeric rating that would be included in X+.

Before inspecting results, choose the primary outcome and headline X+ threshold, matching variables, planned brands and segments, and final reporting date or sample size. Other thresholds and interim reads can support exploration, but should not replace the planned headline because they show a larger lift. See Lift measurement methodology for study design and multiple comparisons.

Create or edit a Lift report

On the Reports tab, select Create → Lift, enter the name and audience settings, then Save to open the editor. For an existing Lift report, select it and choose Edit. The editor follows Scope → Lift model → Reporting → Relabelling → Significance. Use Save + Next to save and advance, or Save to finish. Relabelling is unavailable when there are no editable outcome values or tags.

1. Scope

Name the report and choose the included respondent statuses and reporting period. Project reports also let you choose surveys and Require consistent respondent IDs when an ID identifies the same person across those surveys.

Check the scope before interpreting group differences: changing the date window or included statuses changes the population being compared.

Lift report Scope step with the report name, respondent statuses and date window

2. Lift model

Define the target group, its matched control group and the outcomes to measure:

  • Group indicator — the question or dataset tag that divides respondents into groups. For an exposure study, choose the field recording exposure.
  • Target group values — answers that identify the target group. Other answered values form the control group. Confirm these values again after changing the indicator.
  • Matching questions — characteristics that influence the outcome before exposure, such as prior category purchasing, existing customer status or pre-existing need. Include age group, gender or region where relevant to the outcome or confounding. These are used for matching, not as outcomes.
  • Matching iterations — how many matched control groups are built, from 1 to 50. More iterations improve stability but take longer.
  • Outcome questions — the attitudes or behaviors whose lift you want to calculate. When the group indicator is a tag, only questions carrying that tag are available.

For conversion, ask: What made someone more likely to convert even without the campaign? Choose a small, justified set of matching variables, considering factors that affect both exposure and the outcome. Do not match on recall, consideration or intent measured after exposure if the campaign could have changed them. Asking a question earlier in the survey does not make it a pre-exposure measure.

Lift model using an exposed tag, True target value, matching questions, five iterations and selected outcomes

3. Reporting

For every outcome, choose Calculate lift for and Cut by:

Calculate lift forInterpretation
Each valueCompare the rate of each categorical answer
Each value or higher (≥)Compare cumulative numeric thresholds, such as 7 or higher
Each value response (<)Compare responses below each numeric threshold

Only supported calculations are offered for that outcome. If an older saved metric is marked unsupported, choose a supported metric before saving.

For rating-scale outcomes, use Each value or higher (≥) for X+ reporting and lead with the threshold selected in the analysis plan. A 4+ result on a five-point scale combines ratings 4 and 5; it is not the rate selecting exactly 4. Treat additional thresholds as exploratory unless they were included in the planned analysis.

Cut by accepts one question or tag for each outcome. Choose None for an overall result. Different outcomes can use different sources; for example, cut awareness by brand and message recall by creative.

Lift Reporting step with Each value calculations and brand as the Cut by tag for each outcome

4. Relabelling

Use Outcome values to change how individual answers or cut-by values appear in the report. The table shows Outcome question, Original value and Report as. Giving several values the same label combines them, such as combining ratings 4 and 5 into Top two box.

Use Group indicator tag and Cut by tag to relabel the available tags in the same step. These reporting changes do not edit the respondent-facing survey wording. The editor explains when a question has no editable values or has too many values to edit here.

Lift Relabelling step showing original outcome values, editable Report as labels and notices for questions with too many values

5. Significance

P-value threshold controls which results appear. The report calculates a p-value for each matching iteration, averages those values, and keeps results whose Mean P-Value is less than or equal to the threshold. Lower thresholds show fewer results by requiring a lower mean p-value.

ThresholdHow to use it
0.05A stricter filter on the mean p-value; not a pooled 5% significance test
0.10A looser cutoff for directional analysis
0.95A broad exploratory view that includes results with weak statistical evidence

Choose the cutoff before interpreting the results. A threshold of 0.95 does not mean 95% confidence: it allows mean p-values up to 0.95. The usual 95% confidence convention corresponds to a 0.05 significance level for an appropriate test. The report's mean across matching iterations is a summary of those tests, so it should not be read as a pooled test or a probability that the campaign worked. Selecting 0.05 alone does not establish a confirmatory finding; that requires an inference method appropriate to the study design and planned comparisons.

Lift Significance step with P-value threshold set to 0.95 and Include negative lift switched off

The screenshot uses 0.95 to inspect a wider set of results. Their presence in the report does not establish statistical significance. See Sample size and precision for study planning, matching losses and the risks of repeated interim reads.

Enable Include negative lift as a reporting best practice so outcomes where the target group performs worse than the matched control group are included. The screenshot shows this setting switched off; it is not the recommended reporting choice. The p-value filter still applies, so do not treat the visible rows as a complete account of all planned outcomes. Report inconclusive findings as well as positive and negative findings, and distinguish no clear evidence of lift from evidence of no effect. Save the report when its scope, model, calculations and labels are ready.

Read and export the results

Each outcome has its own table, showing the question or tag used to split the results. Read lift alongside the target and control rates and the number of respondents in each group.

Before drawing conclusions:

  1. Compare the usable matched bases with the achieved sample and account for respondents lost through matching.
  2. Check balance on the chosen matching characteristics using the study's matching diagnostics; selecting a variable alone is not evidence of adequate balance. Document any unavailable diagnostics as a limitation.
  3. Check the usable base for every reported brand and segment. A sufficient overall sample does not guarantee sufficient precision within a cut.
  4. State the outcome and X+ threshold, target and control rates, matched bases, reporting window, study design and mean p-value filter with each headline. Document denominator rules, including missing answers and Don't know responses.

Report absolute lift in percentage points. A control rate of 30% and target rate of 36% gives +6 percentage points. Relative lift is a separate calculation: 6 ÷ 30 = +20%. Label it explicitly if used; relative lift is undefined when the control rate is zero.

Separate Lift outcome tables for ad recall, brand awareness and brand consideration, each cut by brand

These results use the exploratory 0.95 threshold shown above. Open a stimulus's lift details to compare Lift, group rates, base sizes and Mean P-Value together.

Clove lift details showing group counts, rates, lift, Mean P-Value and the Download control

For example, Ad Recall Brands shows +7 percentage points with a Mean P-Value of 0.12. It appears at a threshold of 0.95, but would be excluded at 0.05 or 0.10. A large-looking lift alone is not evidence of statistical significance.

Use each table's row and column controls, sorting and widths to arrange the results. The corresponding exports preserve the separate tables and their display choices. Use Downloading data and reports for Excel and CSV workflows.

Supply exposure information in the survey

The tag itself is set in the survey using s.get_exposed_values(...) and an s.tag() block. get_exposed_values returns a list of values MX8 Labs has on file for the current respondent on the named source — commonly brand or creative identifiers. Wrapping the rest of the survey in with s.tag(exposed=...): carries the indicator onto every question, which you can select as the Lift model's Group indicator.

python
from survey import Survey s = Survey(**globals()) exposed_brands = s.get_exposed_values( source="campaign-123", exposed_dimension="brand", ) with s.tag(exposed=bool(exposed_brands)): # The rest of your survey runs inside the block. # Every question carries the `exposed` tag automatically. ... s.complete()

Each returned value also carries frequency and (when available) days_since_exposure tags, so you can build more nuanced indicators — for example, "exposed three or more times in the last 7 days" — instead of a simple exposed/not-exposed split.

For setting up the exposure source that feeds this, see What Are Exposure Sources?.

Limitations

  • Matching is not randomization. Where exposure is observational, unmeasured confounding can remain.
  • Matching reduces usable sample size, and too many matching questions can reduce power.
  • Exposure and group indicators must be recorded reliably.
  • Results describe the matched population and selected reporting window.
  • Small groups and exploratory significance thresholds need careful interpretation.

For the statistical design, see Lift measurement methodology. For exposure-source configuration, see What Are Exposure Sources?.