MX8 Labs supports Synthetic Twins, AI-generated respondents trained on your own historical survey data. They are a fast way to explore ideas, triage concepts, and pressure-test survey designs. This page covers a narrower question. When you compute a number from synthetic respondents, whether it is a percentage, a gap between two groups, or a cell flagged "significant," how far can you trust it?
Our position, and the short answer, is that synthetic respondents are a tool for direction, not for inference. We label synthetic data clearly, keep it separate from human responses, and never treat a synthetic sample's row count as a real sample size when reporting significance. (For how inference works on human data, see Understanding stat testing in MX8 reports and Weighting methodology.) The rest of this page works through the questions worth asking before you trust a synthetic number, and indexes the current research so you can judge it yourself. We review it periodically as the literature moves.
1. Do they reflect reality?
This is the question the field started with, and it has the deepest literature. The founding result showed that a language model, conditioned on demographic profiles, can reproduce recognizable patterns in subgroup opinion. Its authors called the idea "silicon sampling."
The same optimism extended to behavior and to replicating known studies.
Look closer, though, and the resemblance frays. Models track the average better than they track nuance, context, or the real spread of human variation, and they carry systematic skew.
The honest reading is "often, on the average, yes, with caveats that grow the closer you look." One distinction matters for the rest of this page. This literature is about whether the answers resemble real people, not about how much to trust a number computed from them. Looking real and being countable as real data are different properties, and the second one is where the trouble lives.
2. Can I treat 1,000 of them as 1,000 people?
Almost everyone makes this leap silently: put the synthetic data in a crosstab, let the software flag what's "significant," and move on. It feels harmless. The answers are plausible, there are a thousand of them, so surely they behave like a thousand respondents.
They don't, and this is the hinge of the whole guide. Every confidence interval and significance test you have ever run leans on one assumption so reliable for human samples that it is easy to forget: that respondents are, roughly, independent draws from a population. Independence is what lets N respondents count as N pieces of information. A thousand answers from one model are a thousand draws from the same trained system, built on the same finite pile of real data. The moment you care about the population rather than the model, they are positively correlated. They are not a thousand independent voices. They are closer to one voice, sampled a thousand times.
Survey statistics has priced this since 1965.
The damage is measurable. The essential warning comes from a study that ran the numbers directly.
Practitioners have reached the same place from the applied side, watching synthetic samples manufacture significance that isn't there.
The likely mechanism is mode collapse: alignment-tuned models return near-identical answers far more often than a real population would, which deflates variance and inflates apparent significance.
3. Is it statistically significant?
This is the question people actually ask. You see a gap between two groups in the synthetic data, a cell lights up "significant," and you want to know whether it is real. Answering it honestly takes two steps: working out how much a synthetic sample is worth, and then whether a valid result can be recovered from it at all.
Start with what a synthetic respondent is worth. If they are not worth one real respondent each, the question is what they are worth. The most direct paper in the field asks it in its title.
The uncomfortable implication is that this number saturates. Because all your synthetic respondents share one model, generating more of them past a point adds rows, not information. The effective sample size climbs toward a ceiling set by how much the model knows, not by how many times you sample it. A cell that lights up "significant" on a large synthetic base is often significant only because the test was handed a row count it should never have trusted.
That leaves the second step: can you recover a valid result anyway? Sometimes, if you keep a real anchor. The idea is to use a small amount of gold-standard human data to correct the abundant machine-generated data and recover honest error bars.
A cluster of 2025-26 work carries it into surveys.
For thresholds and quantiles, such as top-box scores, medians, and tail metrics, the newest and least-settled entrant takes on effective sample size under clustering directly.
The pattern across all of it: none of these methods let synthetic data stand on its own. Each one recovers honest error bars by tethering the synthetic data to real human data. One source of confusion is worth clearing up. "Synthetic data" in the older, official-statistics sense is a different thing, with a mature variance apparatus that works because it is fit to real respondents.
So, is the synthetic result statistically significant? On its own you can't honestly say. A synthetic sample's row count is not its information content, and a valid answer needs a human anchor.
4. What can I reasonably infer today?
The profession has drawn its line, and it is cautious.
As of this writing there is no standard, no accepted formula, and no software convention.
Open questions
Synthetic respondents are, today, a tool for direction, not for inference. Trust the direction, not the decimal. The methods that could change that (see section 3) are real but unsettled, and the questions below are where the field, and this page, will grow. We review this article regularly as part of our documentation review and update it as the literature moves.
- A named design effect for synthetic respondents. Nobody has yet estimated the intra-cluster correlation and reported an explicit design effect for model-generated data. Noonan is the first move, and it is wide open.
- A standard, reportable effective sample size. There is no agreed figure for what to divide a synthetic sample's row count by.
- Validation protocols. How large a human holdout must be, and how often it is needed, to anchor a synthetic read.
- Quantifying bias. Bias has no volume-based cure. You cannot generate your way out of a model being wrong, so the frontier is measuring it, not removing it.
- Reporting conventions. How synthetic-derived numbers should be labeled and caveated in a client deliverable.
Last updated: August 6, 2026. If you've published something we should include, get in touch.

