Documentation

Weighted results FAQ

Using the MX8 Labs Research Platform, all your results in cross-tabs or insights are automatically weighted to make up for sample skews.

If you're looking for the mathematical detail behind any of the answers below, see Weighting methodology.

Inspect weights with a Weighting report

On the survey or project's Reports tab, select Create → Weighting, enter the report details and save. Select an existing Weighting report and Edit to adjust its diagnostics.

The editor is one form: set Name, Respondent Statuses, Report window start, Report window end and Histogram bins. Dates use the local timezone shown beneath the fields. Histogram bins range from 1 to 100 and default to 50; more bins show finer detail in the weight distribution. Project reports also select surveys and whether respondent IDs are consistent across them.

Weighting report editor with audience, date window and 50 histogram bins

Save to regenerate the diagnostics for that audience. This form describes existing weights; change weighting targets in the source's weighting settings described below. Cross-tab and Video Overlay editors also offer Unweighted when a particular report should use raw counts.

How do I set up weighting?

Weighting is set up when you configure quotas and weighting targets for your respondent source.

A question is eligible as a weighting variable only when at least 99% of completed respondents answered it and it was asked approximately once per respondent. Questions answered by a smaller subset, such as a branch-specific question, are excluded. For older surveys, refresh the reporting cache before relying on the coverage check; older caches can show a default response rate until rebuilt.

For example, if you field with the US Genpop standard screener, you can define target demographics in the quota tab.

Weighted Results Faq

You can see the weighting setup on the Summary tab:

Weighted Results Faq

Select User Defined to edit the targets in the visual weighting editor. Each card is one weighting table, and every editable cell is a percentage target:

Visual weighting editor showing separate County and Region weighting tables

How should I structure long or complex weighting schemes?

Select Add weighting table for each independent weighting root, then use the table settings button to choose its variables.

Choose the structure that matches your weighting plan:

  • For flat weighting, create one table with one variable, such as age or region. Enter a target for every answer and make sure the targets total 100%.
  • For joint weighting, add two or more variables to one table. Enter a target for every generated combination, such as County × Region.
  • For hierarchical weighting, use the up and down controls in the table dialog to set the variable order. Put the outer level first, followed by each more detailed level.
  • For one-stage weighting, configure only Weighting Scheme. Use this when one calibration objective is enough.
  • For two-stage weighting, configure Pre-Weighting Scheme as well as Weighting Scheme. Pre-weighting runs first and the main stage calibrates its output.

For a flat table, select one variable and save:

Edit weighting table dialog with County selected as its single variable

For a joint or hierarchical table, add more variables and arrange their order. The combination count updates before you save:

Edit weighting table dialog with County followed by Region, producing 112 combinations

How do I weight custom numeric ranges?

Add the numeric question to a weighting table, then enter its buckets in the table dialog as comma-separated integer ranges. For example:

Code
0-5, 6-10, 11+

Edit weighting table dialog showing custom integer age buckets and the option to include missing responses

Both ends of a bounded range are included, so 0-5 contains 0 through 5. A plus range has no upper bound, so 11+ contains 11 and every higher response. After saving the table structure, enter a percentage target for every resulting bucket.

Weighting table with a percentage target for each custom numeric age bucket

Before saving, check that:

  • every range runs from its lower value to its higher value;
  • ranges do not overlap or repeat; and
  • the ranges cover the responses you intend to weight.

Changing a numeric question's buckets updates every weighting table that uses that question. Review the resulting combinations and targets before saving.

How do I give missing responses an explicit target?

In the weighting table dialog, select Include missing responses for [question]. The editor adds Missing response as a category, including within joint and hierarchical tables. After saving the table structure, enter the percentage that respondents without an answer should represent.

If the data contains missing answers and you do not include Missing response, the weighting setup cannot account for those respondents. Use a target of 0 to exclude missing answers from the weighted distribution, or a positive target to include them at the percentage you enter.

Missing responses and numeric ranges can be used together. For example, a numeric table can contain 0-5, 6-10, and Missing response, with a target for each category.

What does the visual editor validate?

Review every table before submitting. The editor validates the whole draft and reports every affected table, including tables with no variables, invalid or duplicate numeric buckets, missing combination values, invalid percentages, or distributions that do not total 100%. In a joint table, check the complete combination grid for empty cells.

Each card shows its issue count and an actionable message above the affected cells. For example, this joint County → Region table identifies all 112 empty targets:

Joint weighting table showing an issue count and an error explaining that 112 values are empty

Invalid changes are not saved, and the survey cannot be submitted until the draft passes validation. Fix every table that shows an issue count, not just the first message at the bottom of the form.

How do you run the weighting?

Weighting runs automatically when the survey closes. It can also run while fieldwork is still open once the minimum base is met, so you can read directional weighted results before close.

MX8 Labs uses an IPFN (Iterative Proportional Fitting) algorithm that iteratively adjusts respondent weights until configured targets are matched within tolerance.

Do you support two-stage weighting?

Yes. A dataset can run up to two stages:

  1. Pre-weighting stage
  2. Main weighting stage

When both are configured, pre-weighting runs first and writes intermediate respondent weights. Main weighting then runs on top of those weights. Final respondent weight is multiplicative:

final weight=pre-weight multiplier×main-weight multiplier\text{final weight} = \text{pre-weight multiplier} \times \text{main-weight multiplier}

This is useful when you need to correct more than one thing in sequence, for example:

  • Stage 1: align a first-party list or panel composition.
  • Stage 2: align reporting outputs to market-level targets.

Can I choose which respondents are used for weighting?

Yes. You can choose the respondent universe used for weighting, and this can differ by stage when two-stage weighting is configured.

Typical choices include:

  • Completes only.
  • Completes plus selected terminated respondents.
  • A constrained subgroup used for frame correction in pre-weighting.

Choose the universe based on the inference you want to support:

  • Use a broader universe when you need representativeness with respect to all eligible entrants.
  • Use completes-only when you need representativeness for the final analyzable sample.
  • Use different universes across stages when one stage is about correcting source composition and another is about final reporting inference.

Why would I use different respondent sets across stages?

Because different stages often answer different methodological questions:

  • Pre-weighting can normalize source or recruitment bias before analysis.
  • Main weighting can calibrate the analysis population to reporting targets.

Using separate respondent sets lets you avoid forcing one compromise universe to do both jobs badly.

What happens if the platform can't weight my data?

Before weighting runs, the platform checks that calibration is safe to attempt. It confirms that every question referenced by targets is present, that there are enough respondents complete across weighting questions, and that every category with a positive target has at least one eligible respondent in the selected respondent set.

If any stage fails, weighting is skipped entirely (all-or-nothing) and the dataset is reported with unit weights. Diagnostics show which check failed so you can fix the root issue (usually by collapsing sparse categories, revising targets, or adding sample).

What's the distribution of my weights?

You can inspect weight distribution by downloading raw data in Excel or SPSS format and charting a histogram.

Where can I see the weighting diagnostics?

Each dataset has a weighting diagnostic report that summarizes how weighting performed: the raw respondent count, effective sample size, weighting efficiency, the minimum, median, mean, and maximum weight, a histogram of the weight distribution, and the eligibility status (with a failure reason) for each configured stage. Use it to spot heavy weighting, a long right tail in the weight distribution, or a stage that was skipped because a target category had no eligible respondents. For the formulas behind each metric, see Weighting methodology.

How do I weight boosts?

"Boost" refers to two different things, and they weight differently. Check which one you have before reading on.

A boost respondent source is a separate respondent source with its own targeting definition, configured to weight back to the primary source. It inherits the survey's weighting configuration, so the extra interviews raise the subgroup's achieved base without moving the overall distribution. It does not create a separate weighted inference frame for that subgroup — it is still one dataset with one calibration.

A boost quota line is a line inside an existing quota group that sets min_respondents with no quota proportion. It raises the base too, but it is excluded from the weighting targets, so the boosted group stays over-represented in the data. That is usually the intent: you wanted the base for sub-analysis, not a change to the headline proportions. If you do want the headline corrected, give the line a quota value as well as a minimum.

If your subgroup is over-represented in the results and you did not expect it to be, you almost certainly have a boost quota line where you wanted a boost respondent source. See Setting up quotas and How to set up respondent sources.

How do I weight nested quotas?

Weighting runs automatically for nested quotas in your configured respondent source.

Can I turn off weighting?

If you set the weighting strategy to None, results are unweighted. Caveat emptor: without proper weighting, your results will not reflect the actual population, and conclusions drawn from them can be misleading. Only do this when unweighted inference is what you actually want — for example, when you're describing the realized sample itself rather than generalizing to a population.

What is effective sample size?

When respondents carry unequal weights, the "information content" of your sample is smaller than the raw respondent count. We report Kish effective sample size for each reporting cell:

neff=(∑wi)2∑wi2n_{\text{eff}} = \frac{\left(\sum w_i\right)^2}{\sum w_i^2}

If all weights are equal, neffn_{\text{eff}} equals the raw count. As weights become more unequal, neffn_{\text{eff}} falls below it. Analytic standard errors and significance tests in weighted cross-tabs use neffn_{\text{eff}} to reflect the precision cost of unequal weights.

Median aggregations use respondent bootstrap comparisons, matched-group Lift tests use respondent outcome counts, and Conjoint/MaxDiff uncertainty also incorporates posterior draws. See Understanding stat testing for these distinctions.

What is weighting efficiency?

Weighting efficiency is the dataset-level summary of precision cost:

Efficiency=neffn\text{Efficiency} = \frac{n_{\text{eff}}}{n}

An efficiency near 1 means weighting has barely moved the needle on precision. Lower efficiency, or a long right tail on the weight histogram, means a small number of respondents are carrying disproportionate weight. That's not wrong — it's the correct response to an imbalanced sample — but it's a signal to check whether the targets are achievable from the realized sample, or whether some categories should be collapsed.

Why don't percentages in my results exactly match my target scheme?

Weighting targets define calibration constraints, but reported bases reflect your selected reporting universe and the realized sample. If incidence and termination patterns differ across groups, reported percentages can differ from intake shares while still being correctly calibrated for the inference frame you configured.

A worked example makes this concrete. Suppose you're running a survey of ESPN subscribers against a US genpop respondent source, and 1000 people enter the experience — 500 men and 500 women. ESPN skews male, so suppose 60% of the women are screened out for not being subscribers. You end up with 700 respondents who completed: 500 men and 200 women.

When the platform weights, it includes the 300 terminated women too — because the calibration is to the entrant population, not just the completes. With that universe, men and women end up evenly weighted, and the report tells you 500 / 700 ≈ 71.4% of ESPN subscribers are men.

That isn't consistent with the 50/50 source you specified upfront, but it is nationally representative — which is the thing that matters for the inference. The mismatch between source proportion and report proportion is the calibration working as intended.