Documentation

Can I trust this number? Synthetic data and statistical inference

MX8 Labs supports Synthetic Twins, AI-generated respondents trained on your own historical survey data. They are a fast way to explore ideas, triage concepts, and pressure-test survey designs. This page covers a narrower question. When you compute a number from synthetic respondents, whether it is a percentage, a gap between two groups, or a cell flagged "significant," how far can you trust it?

Our position, and the short answer, is that synthetic respondents are a tool for direction, not for inference. We label synthetic data clearly, keep it separate from human responses, and never treat a synthetic sample's row count as a real sample size when reporting significance. (For how inference works on human data, see Understanding stat testing in MX8 reports and Weighting methodology.) The rest of this page works through the questions worth asking before you trust a synthetic number, and indexes the current research so you can judge it yourself. We review it periodically as the literature moves.

1. Do they reflect reality?

This is the question the field started with, and it has the deepest literature. The founding result showed that a language model, conditioned on demographic profiles, can reproduce recognizable patterns in subgroup opinion. Its authors called the idea "silicon sampling."

Peer-reviewed
Argyle, Busby, Fulda, Gubler, Rytting & Wingate (2023). "Out of One, Many: Using Language Models to Simulate Human Samples." Political Analysis 31(3).
Conditioned on personas, LLMs reproduce recognizable subgroup response patterns. This founding "silicon sampling" result speaks to resemblance, not to statistical certainty.
Read →

The same optimism extended to behavior and to replicating known studies.

Peer-reviewed
Horton (2023). "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?" NBER w31122.
LLM "agents" reproduce qualitative economic behaviors. This is evidence of behavioral plausibility, not of measurement.
Read →
Peer-reviewed
Aher, Arriaga & Kalai (2023). "Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies." ICML.
"Turing Experiments" recover the direction of classic findings. They reproduce the effects, but not the spread around them.
Read →

Look closer, though, and the resemblance frays. Models track the average better than they track nuance, context, or the real spread of human variation, and they carry systematic skew.

Peer-reviewed
Santurkar, Durmus, Ladhak, Lee, Liang & Hashimoto (2023). "Whose Opinions Do Language Models Reflect?" ICML.
Introduces OpinionQA and documents systematic bias versus benchmark populations. Models over-represent some groups and flatten others.
Read →

The honest reading is "often, on the average, yes, with caveats that grow the closer you look." One distinction matters for the rest of this page. This literature is about whether the answers resemble real people, not about how much to trust a number computed from them. Looking real and being countable as real data are different properties, and the second one is where the trouble lives.

2. Can I treat 1,000 of them as 1,000 people?

Almost everyone makes this leap silently: put the synthetic data in a crosstab, let the software flag what's "significant," and move on. It feels harmless. The answers are plausible, there are a thousand of them, so surely they behave like a thousand respondents.

They don't, and this is the hinge of the whole guide. Every confidence interval and significance test you have ever run leans on one assumption so reliable for human samples that it is easy to forget: that respondents are, roughly, independent draws from a population. Independence is what lets N respondents count as N pieces of information. A thousand answers from one model are a thousand draws from the same trained system, built on the same finite pile of real data. The moment you care about the population rather than the model, they are positively correlated. They are not a thousand independent voices. They are closer to one voice, sampled a thousand times.

Survey statistics has priced this since 1965.

Foundational
Kish (1965). Survey Sampling. Wiley.
The origin of the design effect and effective sample size, the tools that quantify how much information is lost when respondents aren't independent. Synthetic respondents are that problem at its maximum: one shared "household" the size of the dataset.

The damage is measurable. The essential warning comes from a study that ran the numbers directly.

Peer-reviewed
Bisbee, Clinton, Dorff, Kenkel & Larson (2024). "Synthetic Replacements for Human Survey Data? The Perils of Large Language Models." Political Analysis 32(4).
LLM responses show far smaller variance than real ANES data. A power analysis on synthetic responses wanted about 33 respondents where the real data needed about 300, false precision by nearly an order of magnitude.
Read →

Practitioners have reached the same place from the applied side, watching synthetic samples manufacture significance that isn't there.

Industry
Verian Group (2026). "Synthetic Sample in Social Research: significant limitations of AI-generated responses."
An industry statement of the same failure: synthetic models "flatten variance" until "every difference becomes statistically significant and thus loses its meaning."
Read →

The likely mechanism is mode collapse: alignment-tuned models return near-identical answers far more often than a real population would, which deflates variance and inflates apparent significance.

3. Is it statistically significant?

This is the question people actually ask. You see a gap between two groups in the synthetic data, a cell lights up "significant," and you want to know whether it is real. Answering it honestly takes two steps: working out how much a synthetic sample is worth, and then whether a valid result can be recovered from it at all.

Start with what a synthetic respondent is worth. If they are not worth one real respondent each, the question is what they are worth. The most direct paper in the field asks it in its title.

Preprint
Huang, Wu & Wang (2025). "How Many Human Survey Respondents Is a Large Language Model Worth? An Uncertainty Quantification Perspective." arXiv 2502.17773.
Derives an effective sample size for LLM respondents, the number of independent humans their answers are worth, and builds intervals that widen for human-versus-LLM misalignment.
Read →

The uncomfortable implication is that this number saturates. Because all your synthetic respondents share one model, generating more of them past a point adds rows, not information. The effective sample size climbs toward a ceiling set by how much the model knows, not by how many times you sample it. A cell that lights up "significant" on a large synthetic base is often significant only because the test was handed a row count it should never have trusted.

That leaves the second step: can you recover a valid result anyway? Sometimes, if you keep a real anchor. The idea is to use a small amount of gold-standard human data to correct the abundant machine-generated data and recover honest error bars.

Peer-reviewed
Angelopoulos, Bates, Fannjiang, Jordan & Zrnic (2023). "Prediction-Powered Inference." Science 382.
The foundational recipe: a small human gold-standard sample debiases abundant model output to recover valid confidence intervals and p-values. The lesson that model output does not count one-for-one starts here.
Read →

A cluster of 2025-26 work carries it into surveys.

Peer-reviewed
Krsteski et al. (2026). "Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification." ACL 2026 (Main).
Peer-reviewed. Combines LLM responses with a small human sample, reports the modest effective-sample-size gain that buys, and shows synthesis alone carries 24 to 86% bias.
Read →
Preprint
Ye & Yoganarasimhan (2026). "Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys." arXiv 2604.17267.
Turns the method into a budgeting rule for how many humans versus synthetics to collect, using residual variance after optimally using the LLM rather than raw LLM accuracy.
Read →
Preprint
Tan & Zrnic (2026). "Valid Inference with Synthetic Data via Task Exchangeability." arXiv 2606.13629.
Recovers valid intervals by calibrating, on past tasks where real data existed, how wrong synthetic data tends to be, then widening to cover that gap.
Read →

For thresholds and quantiles, such as top-box scores, medians, and tail metrics, the newest and least-settled entrant takes on effective sample size under clustering directly.

Preprint
Noonan (2026). "The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering." Zenodo (preprint).
Estimates the effective sample size for threshold statistics directly, and is the first to frame the problem as an explicit design effect for correlated, model-generated data. Not yet peer reviewed, and written for ML, so applied here by analogy.
Read →

The pattern across all of it: none of these methods let synthetic data stand on its own. Each one recovers honest error bars by tethering the synthetic data to real human data. One source of confusion is worth clearing up. "Synthetic data" in the older, official-statistics sense is a different thing, with a mature variance apparatus that works because it is fit to real respondents.

Peer-reviewed
Reiter (2003). "Inference for Partially Synthetic, Public Use Microdata Sets." Survey Methodology 29(2).
The other "synthetic data": privacy-preserving microdata with decades-old variance-combining rules. Included to mark the distinction. That machinery does not transfer to LLM respondents, which aren't fit to a known real-data model.
Read →

So, is the synthetic result statistically significant? On its own you can't honestly say. A synthetic sample's row count is not its information content, and a valid answer needs a human anchor.

4. What can I reasonably infer today?

The profession has drawn its line, and it is cautious.

Guidance
AAPOR Task Force on Responsible AI Integration in Survey Research (Rothschild & Marlar, 2026).
Limits synthetic responses to "clearly labeled pretesting, pilot work, or exploratory diagnostics," notes that removing human respondents sits "outside the bounds of total survey error," and states the field "has not yet reached consensus on when, if ever, AI-generated responses can stand in for human ones."
Read →
Guidance
Sarstedt, Adler, Rau & Schmitt (2026). "Your Next Respondent Might Be an LLM: Guidelines for Using Silicon Samples in Marketing Research." NIM Marketing Intelligence Review 18(1).
Thoughtful validation guidance (compare means, variance, and range to human benchmarks) that still stops short of significance testing or an effective-sample-size method.
Read →

As of this writing there is no standard, no accepted formula, and no software convention.

Open questions

Synthetic respondents are, today, a tool for direction, not for inference. Trust the direction, not the decimal. The methods that could change that (see section 3) are real but unsettled, and the questions below are where the field, and this page, will grow. We review this article regularly as part of our documentation review and update it as the literature moves.

  • A named design effect for synthetic respondents. Nobody has yet estimated the intra-cluster correlation and reported an explicit design effect for model-generated data. Noonan is the first move, and it is wide open.
  • A standard, reportable effective sample size. There is no agreed figure for what to divide a synthetic sample's row count by.
  • Validation protocols. How large a human holdout must be, and how often it is needed, to anchor a synthetic read.
  • Quantifying bias. Bias has no volume-based cure. You cannot generate your way out of a model being wrong, so the frontier is measuring it, not removing it.
  • Reporting conventions. How synthetic-derived numbers should be labeled and caveated in a client deliverable.

Last updated: August 6, 2026. If you've published something we should include, get in touch.