Two questions decide how large a study needs to be. How precisely do we need to estimate a single number, and how small a difference do we need to be able to detect between two groups? They have different answers, and the second almost always demands more sample than people expect.
This page gives the arithmetic and the practical adjustments used to size a study.
Estimating a single number
For a simple proportion, the 95% margin of error is widest when the result sits near 50% and narrows as it moves toward either extreme. At 50%:
| Base | Margin of error |
|---|---|
| 200 | ±6.9 pts |
| 400 | ±4.9 pts |
| 500 | ±4.4 pts |
| 1,000 | ±3.1 pts |
| 1,500 | ±2.5 pts |
| 2,000 | ±2.2 pts |
Precision improves with the square root of sample, so halving the margin of error requires four times the sample. Increasing the base from 500 to 1,000 reduces the margin of error by 1.3 points. Increasing it from 1,000 to 2,000 reduces it by another 0.9 points.
MX8 Labs reports a weighted margin of error on every result, computed against the effective sample size rather than the raw base, so the figure shown already accounts for the precision cost of weighting. See Weighting methodology.
Detecting a difference between two groups
Comparing two groups (exposed against control, this segment against that one) needs materially more sample than estimating one number, because both sides carry error.
The minimum detectable effect is the smallest true difference a design has a reasonable chance of finding. It is not a property of the base alone: it depends on the base, the confidence level, the power, and the outcome's base rate together. Change any one and the answer moves. The table below holds three of them fixed at the usual conventions — 95% confidence, 80% power, equal groups — and varies the base; the effect of relaxing those conventions is set out below.
| Base per group | Outcome near 50% | Outcome near 20% |
|---|---|---|
| 200 | 14.0 pts | 11.2 pts |
| 300 | 11.4 pts | 9.1 pts |
| 400 | 9.9 pts | 7.9 pts |
| 500 | 8.9 pts | 7.1 pts |
| 750 | 7.2 pts | 5.8 pts |
| 1,000 | 6.3 pts | 5.0 pts |
| 2,000 | 4.4 pts | 3.5 pts |
Turned around, the same arithmetic gives the base each ambition requires:
| Difference you need to detect | Base per group, outcome near 50% | Base per group, outcome near 20% |
|---|---|---|
| 15 pts | 175 | 115 |
| 10 pts | 400 | 250 |
| 8 pts | 615 | 395 |
| 5 pts | 1,570 | 1,005 |
| 3 pts | 4,360 | 2,790 |
| 2 pts | 9,810 | 6,280 |
Moving from a five-point to a three-point minimum detectable difference roughly triples the required sample. Set that threshold before calculating the study size.
Read it against the effect you actually expect. Advertising brand lift on upper-funnel metrics frequently lands in the three to eight point range, which means a design with 300 per group is capable of detecting a strong campaign and incapable of detecting a normal one. A null result from an underpowered study is not evidence that the campaign did nothing; it is evidence that the study could not tell.
Outcomes with lower base rates require smaller percentage-point differences for the same power, so the right-hand column is smaller. Size the study against its primary outcome rather than an average across the full question battery.
The smaller group governs
Precision in a two-group comparison is driven by the smaller group, and much more strongly than intuition suggests. Adding respondents to the larger side helps very little once the imbalance is significant:
| Design | Minimum detectable effect |
|---|---|
| 300 exposed, 300 control | 11.4 pts |
| 500 exposed, 500 control | 8.9 pts |
| 500 exposed, 1,000 control | 7.7 pts |
| 500 exposed, 2,000 control | 7.0 pts |
| 300 exposed, 2,000 control | 8.7 pts |
| 1,000 exposed, 1,000 control | 6.3 pts |
Quadrupling the control group from 500 to 2,000 while the exposed group stays at 500 improves detection by under two points. Raising the exposed group from 300 to 500 with a matched control improves it by two and a half.
In campaign work the exposed group is almost always the constraint, and adding control respondents cannot compensate for a small exposed group. An unbalanced design can improve precision, but the benefit diminishes as the control group grows.
Sizing a lift study
Lift has a sizing problem the tables above do not capture: you do not control who ends up in which group. In a concept test you assign respondents to cells. In a campaign lift study the media decides, and the exposed group is nearly always the scarce one.
The exposed-versus-control design
An exposed-versus-control study has exactly two cells. The exposed cell is the respondents the stimulus reached; the control cell is matched respondents it did not. Every outcome is asked identically of both, and every reported figure is the difference between two estimates, each carrying its own sampling error.
So the sizing question is always "how large does each cell need to be", never "how large does the study need to be". A 3,000-respondent study with 120 exposed respondents in it is a 120-respondent study for every purpose that matters.
For brand lift the cells are concrete. Exposed is the group the campaign reached, established either from ad-server records matched to respondents or by randomized assignment inside the survey. Control is a matched group of respondents the campaign did not reach. Outcomes are the funnel metrics — unaided and aided awareness, message recall, consideration, favorability, purchase intent — each asked the same way of both cells so the difference is attributable to exposure rather than to instrument. How the two cells are constructed, and how strong a causal claim each construction supports, is set out in Lift measurement methodology.
Deriving the working standard
The working default is 500 exposed and 500 matched control. The derivation below shows the assumptions and the resulting detection threshold.
Step 1. The formula. For a difference between two proportions, at significance level α and power 1−β, the base required in each group is:
Note what is inside it. Two of the four inputs are choices about how much risk to accept: α, the chance of calling an effect real when it is not, and β, the chance of missing one that is. They are not properties of the data. Fixing them at the conventional 95% confidence and 80% power, and taking the most demanding case of an outcome near 50%, this reduces to:
where d is the difference expressed as a proportion. Everything below is that one expression evaluated at different ambitions.
Step 2. Set the minimum detectable effect. The working design target is to detect differences of roughly eight to ten percentage points on funnel metrics, while allowing smaller observed differences to be reported as significant when the achieved base supports them. Studies intended to detect smaller effects require larger groups.
Step 3. Solve for n. A ten-point target implies 392 per group. Nine points implies 484. Eight points implies 613. The band that matches the design target is therefore 400 to 600 per group, and 500 sits in the middle of it, giving an 8.9-point detection threshold and significance at 6.2 points.
Step 4. Check the marginal return. The formula's inverse-square shape means sensitivity is bought cheaply at first and expensively later:
| Base per group | Detectable effect | Sensitivity gained per additional 100 |
|---|---|---|
| 100 | 19.8 pts | — |
| 200 | 14.0 pts | 5.8 pts |
| 300 | 11.4 pts | 2.6 pts |
| 400 | 9.9 pts | 1.5 pts |
| 500 | 8.9 pts | 1.0 pts |
| 600 | 8.1 pts | 0.8 pts |
| 800 | 7.0 pts | 0.5 pts |
| 1,000 | 6.3 pts | 0.4 pts |
| 2,000 | 4.4 pts | 0.1 pts |
The curve bends between 300 and 500 respondents per group. Before that range, each additional hundred respondents improves the detection threshold by multiple points; after it, the improvement is less than one point.
Step 4b. Check what the conventions are doing. Because α and β sit inside the formula, the same 500 respondents per group support very different claims depending on the risk tolerance chosen:
| Confidence | Power | Detectable effect at 500 per group |
|---|---|---|
| 99% | 80% | 10.8 pts |
| 95% | 90% | 10.3 pts |
| 95% | 80% | 8.9 pts |
| 90% | 80% | 7.9 pts |
| 80% | 80% | 6.7 pts |
The same detection target requires different sample sizes under different assumptions. An 8.9-point detectable effect needs 500 respondents per group at 95% confidence and 80% power, 394 at 90% confidence, 669 at 90% power, and 744 at 99% confidence. Moving from 95% to 99% confidence increases the required sample by about half; moving to 90% reduces it by about a fifth.
The working defaults are 95% confidence and 80% power. They are conventional choices rather than properties of the data, and a study with different tolerances for false positives and false negatives should calculate its sample against those settings.
Step 5. Leave headroom. Two things consume base after the design is set. Matching discards exposed respondents who have no comparable counterpart, so the analyzed base is smaller than the achieved one. And a single binary subgroup cut halves each cell: 250 against 250 detects 12.5 points, which supports a coarse split and nothing finer.
The conclusion. 500 per cell is the smallest base that reliably distinguishes a working campaign from no effect, sits at the point where additional sample stops paying for itself, and leaves enough headroom to survive matching loss and one coarse cut.
The default can be changed. Increase the base when the study must detect smaller effects, when the analysis plan includes segment-level conclusions, or when the primary outcome sits near 50%, where precision is lowest. A base of 300 per group is appropriate only when an 11-point detection threshold is acceptable.
One caution about the two thresholds. Significance and power answer different questions and are routinely conflated. At 500 and 500 an observed six-point difference will be reported as significant, while a true six-point effect would be missed more often than found. A null result from this design is not evidence that nothing happened; it shows that the design could not reliably detect the effect. Base rates change both figures — at an outcome near 20% the same design is significant at 5.0 points and powered at 7.1 — so size against the primary metric rather than an average across the battery.
A related point concerns the reporting threshold. The lift report's p-value threshold is configurable, and the interface default is intended for exploratory reads rather than confirmatory reporting. At 500 respondents per group, the observed difference needed to clear the threshold changes substantially with it:
| Reporting threshold | Observed difference that clears it |
|---|---|
| p < 0.01 | 8.1 pts |
| p < 0.05 | 6.2 pts |
| p < 0.10 | 5.2 pts |
| p < 0.20 | 4.1 pts |
| p < 0.50 | 2.1 pts |
Use a threshold of 0.05 for externally reported results. Looser thresholds increase the probability of false positives and are more appropriate for explicitly exploratory analysis. See Reporting lift against control groups for where this is configured.
Base rates move both figures too. At 500 per group, 95% confidence and 80% power:
| Outcome base rate | Significant at | 80% power at |
|---|---|---|
| ~50% | 6.2 pts | 8.9 pts |
| ~30% | 5.7 pts | 8.1 pts |
| ~20% | 5.0 pts | 7.1 pts |
| ~10% | 3.7 pts | 5.3 pts |
Getting to 500 exposed is the actual problem
Campaign reach against a general population sample is low, frequently low single digits and rarely more than the mid teens. Sampling a general population and waiting for exposed respondents to appear is an expensive way to build the exposed cell:
| Reach within the audience sampled | Respondents needed to yield 500 exposed |
|---|---|
| 2% | 25,000 |
| 5% | 10,000 |
| 10% | 5,000 |
| 15% | 3,300 |
| 20% | 2,500 |
| 30% | 1,700 |
At 5% reach, a sample large enough to yield 500 exposed respondents also produces about 9,500 unexposed respondents. The exposed group therefore determines the recruitment requirement; the untreated candidate pool is usually much larger.
Increasing the exposed-group yield
Sample the campaign's target audience, not the general population. The relevant reach figure is reach within the sampled audience. A campaign reaching 4% of a national population may reach 25% of the segment it targeted. Aligning the survey's screening criteria to that targeting increases the expected share of exposed respondents.
Recruit from the media itself. An in-ad survey launched from inside the ad unit recruits exposed respondents directly rather than waiting for them to appear in a general sample. The exposed and control groups then come from different frames, so the comparison carries a mode and frame difference on top of the exposure difference. State that in the reporting.
Use an unbalanced design where appropriate. Retaining a larger control group can improve precision when untreated respondents substantially outnumber exposed respondents. A design with 500 exposed and 2,000 control detects about seven points rather than 8.9.
Add a matching buffer
The tables give the base needed in the analysis, and matching happens before the analysis. Respondents who cannot be matched are dropped, so the achieved exposed base has to exceed the target.
How much is lost depends on how different the exposed group's profile is from the unexposed pool. A tightly targeted campaign loses more, because its exposed respondents are harder to match. Plan a buffer above the floor rather than sizing exactly to it, and check the realized matched base in the report rather than assuming. Adding control variables increases the loss, so match on the smallest set that plausibly confounds the outcome.
Narrower exposure definitions shrink the treated group again
Defining exposure as three or more impressions in the last seven days is analytically better than a binary flag and cuts the exposed group substantially, often by more than half. If the analysis plan depends on frequency or recency thresholds, size against the base that survives the threshold, not the base that saw the ad at all.
Per-brand base in combined studies
A combined study changes incidence and comparability. It does not change the arithmetic of precision.
Where each brand in a multi-brand study carries its own campaign, that brand's exposed base is set by its reach within the sample. Brands with higher reach produce larger exposed bases; lower-reach brands may support only directional analysis. Allocate the sample around the brands requiring confirmatory results, and pool related brands into category nets where the analysis calls for a category-level result.
The comparability gained is real and independent of all this: every brand is measured against the same respondents in the same window, which cross-study comparison never achieves.
Interim reads inflate false positives
Reading lift continuously while a campaign runs is operationally valuable and statistically hazardous. Checking significance repeatedly on accumulating data raises the chance of crossing the threshold by luck at some point, even when nothing is happening.
Treat in-flight reads as directional and define the final analysis point in advance, typically at a specified sample size or date. Where an interim decision rests on significance, adjust for repeated testing rather than applying an unchanged 0.05 threshold at every read.
Adjustments that make the real number larger
The tables above are the floor. Four things push the requirement up.
Matching consumes sample. In a lift design, the analyzed base is the matched control and its counterparts, not the achieved sample. A thin untreated pool, or a treated group with an unusual demographic profile, means fewer usable pairs than the raw numbers suggest. See Lift measurement methodology.
Weighting costs precision. Calibration trades bias for variance. The effective sample size after weighting is lower than the raw base, and the weighting diagnostics report the efficiency ratio directly. Where targets strain the achieved sample, efficiency drops and the realized margin of error widens.
Subgroups need their own base. A study powered at the topline is not powered within a segment. If the analysis plan includes cuts by age, market, or category usage, size against the smallest cell that has to carry a conclusion, and use minimum quota lines or a boost source to protect it. See Reaching low-incidence audiences.
Multi-item studies dilute per item. In a combined study covering many brands or concepts, each respondent sees a subset. Two thousand respondents each seeing five of fifty brands yields roughly two hundred per brand — which, from the second table, supports detecting large differences and nothing subtle. Size against base per item.
Multiple comparisons inflate false positives. Testing many outcomes across many stimuli at p < 0.05 will produce some significant results by chance alone. Where a study runs a wide battery, either tighten the threshold or treat the pattern across related metrics as the finding rather than any single cell. See Understanding stat testing.
Worked examples
These examples assume 95% confidence and 80% power. Check each design against the expected incidence and achieved group sizes before fielding.
Concept or message test, two cells. Detecting a 10-point difference in message take-out needs about 400 respondents per cell, or roughly 800 in total. Because assignment is randomized in-survey, this design supports a causal claim.
Standard campaign lift. 500 exposed and 500 matched control is the working default: significant at about six points on a mid-range metric, reliably detecting effects of nine points or more. If the campaign reaches 10% of the audience being sampled, that means fielding around 5,000 respondents to yield the exposed cell, with the control falling out of the same sample many times over. Align the screening criteria to the campaign's targeting and the same 500 exposed can come from 2,000 respondents or fewer.
Low-reach campaign. At 3% reach against the sampled audience, obtaining 500 exposed respondents requires roughly 17,000 respondents by random sampling. Alternatives are to narrow the sample to the campaign's target audience, recruit exposed respondents from inside the ad unit, or use 300 exposed respondents with an 11-point detection threshold.
Unbalanced by design. When reach is low, untreated respondents substantially outnumber exposed respondents. Retaining 2,000 matched controls against 500 exposed respondents improves the detection threshold from 8.9 points to about seven.
Multi-brand tracker. Each brand's exposed base is set by its own reach within the sample, so higher-reach brands produce more precise estimates and lower-reach brands may support only directional analysis. Design the sample around the brands requiring confirmatory results, and use category nets when the intended result is at category level.
Subgroup analysis within a study. A 2,000-respondent study reporting a cut with 300 respondents in it can detect roughly an 11-point difference within that cut, whatever the topline precision. If a subgroup conclusion matters, protect its base with a minimum quota line or a boost source rather than hoping it falls out. See Reaching low-incidence audiences.
Low-incidence B2B audience. The precision calculation is unchanged, but the recruitment requirement is not. At 4% incidence, obtaining 400 completes requires screening roughly 10,000 entrants, so sample sizing and feasibility must be assessed together.
Practical sequence
Define the smallest difference the study needs to detect and size for that rather than for a round number of interviews. Check the resulting recruitment requirement with feasibility at the expected incidence and duration. Soft launch a small batch to confirm incidence and quota behavior before releasing the full sample. When reading the results, consider the achieved base and margin of error alongside the point estimate.
Scope
Every figure on this page assumes 95% confidence, 80% power, and equal groups unless stated otherwise, and all of them shift if those choices shift. See Step 4b. They also assume simple random sampling, and are exact only under that assumption. Online respondent sources are not probability samples, so treat the tables as the best case and the planning floor rather than as a guarantee. Clustering within respondent sources, design effects from weighting, and non-response are not captured here. The framework's limits are stated in Weighting methodology.

