Documentation

Can I trust this number? Synthetic data and statistical inference

MX8 Labs supports Synthetic Twins, AI-generated respondents trained on historical survey data. They can be used for exploratory analysis, concept triage, and survey-design testing. This page covers a narrower question: what statistical claims, if any, can be made from a synthetic percentage, group difference, or cell marked as significant?

Our position, and the short answer, is that synthetic respondents are a tool for direction, not for inference. We label synthetic data clearly, keep it separate from human responses, and never treat a synthetic sample's row count as a real sample size when reporting significance. (For how inference works on human data, see Understanding stat testing in MX8 reports and Weighting methodology.) The rest of this page works through the questions worth asking before you trust a synthetic number, and indexes the current research so you can judge it yourself. We review it periodically as the literature moves.

1. Do they reflect reality?

This is the question the field started with, and it has the deepest literature. The founding result showed that a language model, conditioned on demographic profiles, can reproduce recognizable patterns in subgroup opinion. Its authors called the idea "silicon sampling."

Peer-reviewed
Argyle, Busby, Fulda, Gubler, Rytting & Wingate (2023). "Out of One, Many: Using Language Models to Simulate Human Samples." Political Analysis 31(3).
Conditioned on personas, LLMs reproduce recognizable subgroup response patterns. This founding "silicon sampling" result speaks to resemblance, not to statistical certainty.
Read →

The same optimism extended to behavior and to replicating known studies.

Peer-reviewed
Horton (2023). "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?" NBER w31122.
LLM "agents" reproduce qualitative economic behaviors. This is evidence of behavioral plausibility, not of measurement.
Read →
Peer-reviewed
Aher, Arriaga & Kalai (2023). "Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies." ICML.
"Turing Experiments" recover the direction of classic findings. They reproduce the effects, but not the spread around them.
Read →

Look closer, though, and the resemblance frays. Models track the average better than they track nuance, context, or the real spread of human variation, and they carry systematic skew.

Peer-reviewed
Santurkar, Durmus, Ladhak, Lee, Liang & Hashimoto (2023). "Whose Opinions Do Language Models Reflect?" ICML.
Introduces OpinionQA and documents systematic bias versus benchmark populations. Models over-represent some groups and flatten others.
Read →

A cross-domain benchmark examines this skew by comparing model predictions with recorded survey responses.

Preprint
Chen, Zhu & Zheng (2026). "When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses." arXiv 2607.26348.
Benchmarks four LLMs against the General Social Survey and World Values Survey: none beats a simple demographic-lookup baseline at predicting individual responses, and all over-weight demographics, inflating segment differences two- to four-fold — enough to point a team at the wrong segment in about half of U.S. cases.
Read →

The most direct audit to date treats the model as a survey instrument and tests whether it behaves like one.

Preprint
Lukauskas & Šarkauskaitė (2026). "Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents." arXiv 2608.14606.
Audits 37 LLMs against a real employee survey and finds broken latent structure, a +0.84 SD acquiescence bias, and significant mediation effects on placebo pathways — synthetic samples manufacturing significance out of nothing. One workplace dataset, and not yet peer reviewed.
Read →

The honest reading is "often, on the average, yes, with caveats that grow the closer you look." One distinction matters for the rest of this page. This literature is about whether the answers resemble real people, not about how much to trust a number computed from them. Looking real and being countable as real data are different properties, and the second one is where the trouble lives.

2. Can I treat 1,000 of them as 1,000 people?

Almost everyone makes this leap silently: put the synthetic data in a crosstab, let the software flag what's "significant," and move on. It feels harmless. The answers are plausible, there are a thousand of them, so surely they behave like a thousand respondents.

They don't, and this is the hinge of the whole guide. Every confidence interval and significance test you have ever run leans on one assumption so reliable for human samples that it is easy to forget: that respondents are, roughly, independent draws from a population. Independence is what lets N respondents count as N pieces of information. A thousand answers from one model are a thousand draws from the same trained system, built on the same finite pile of real data. The moment you care about the population rather than the model, they are positively correlated. They are not a thousand independent voices. They are closer to one voice, sampled a thousand times.

That intuition now has a formal counterpart.

Preprint
Dale, Rodu & Baiocchi (2026). "Synthetic Data, Information, and Prior Knowledge: Why Synthetic Data Augmentation to Boost Sample Doesn't Work for Statistical Inference." arXiv 2603.18345.
An information-theoretic argument that generated respondents add no Fisher information beyond the prior knowledge already baked into the generator — the formal version of "1,000 synthetic interviews are not 1,000 people." Treats augmentation generally rather than survey estimators specifically.
Read →

Survey statistics has priced this since 1965.

Foundational
Kish (1965). Survey Sampling. Wiley.
The origin of the design effect and effective sample size, the tools that quantify how much information is lost when respondents aren't independent. Synthetic respondents are that problem at its maximum: one shared "household" the size of the dataset.

The damage is measurable. The essential warning comes from a study that ran the numbers directly.

Peer-reviewed
Bisbee, Clinton, Dorff, Kenkel & Larson (2024). "Synthetic Replacements for Human Survey Data? The Perils of Large Language Models." Political Analysis 32(4).
LLM responses show far smaller variance than real ANES data. A power analysis on synthetic responses wanted about 33 respondents where the real data needed about 300, false precision by nearly an order of magnitude.
Read →

Practitioners have reached the same place from the applied side, watching synthetic samples manufacture significance that isn't there.

Industry
Verian Group (2026). "Synthetic Sample in Social Research: significant limitations of AI-generated responses."
An industry statement of the same failure: synthetic models "flatten variance" until "every difference becomes statistically significant and thus loses its meaning."
Read →

The likely mechanism is mode collapse: alignment-tuned models return near-identical answers far more often than a real population would, which deflates variance and inflates apparent significance.

One nuance is worth holding onto: part of the collapse appears to be an artifact of how the model is asked, not only of what it knows.

Preprint
Maier et al. (2025). "LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings." arXiv 2510.08338.
Maps free-text LLM answers onto Likert scales via embedding similarity and recovers realistic response distributions across 57 consumer surveys, at roughly 90% of human test-retest reliability. Distributional realism, though, says nothing about independence across synthetic respondents or valid standard errors.
Read →

3. Is it statistically significant?

This is the question people actually ask. You see a gap between two groups in the synthetic data, a cell lights up "significant," and you want to know whether it is real. Answering it honestly takes two steps: working out how much a synthetic sample is worth, and then whether a valid result can be recovered from it at all.

Start with what a synthetic respondent is worth. If they are not worth one real respondent each, the question is what they are worth. The most direct paper in the field asks it in its title.

Preprint
Huang, Wu & Wang (2025). "How Many Human Survey Respondents Is a Large Language Model Worth? An Uncertainty Quantification Perspective." arXiv 2502.17773.
Derives an effective sample size for LLM respondents, the number of independent humans their answers are worth, and builds intervals that widen for human-versus-LLM misalignment.
Read →

The uncomfortable implication is that this number saturates. Because all your synthetic respondents share one model, generating more of them past a point adds rows, not information. The effective sample size climbs toward a ceiling set by how much the model knows, not by how many times you sample it. A cell that lights up "significant" on a large synthetic base is often significant only because the test was handed a row count it should never have trusted.

That leaves the second step: can you recover a valid result anyway? Sometimes, if you keep a real anchor. The idea is to use a small amount of gold-standard human data to correct the abundant machine-generated data and recover honest error bars.

Peer-reviewed
Angelopoulos, Bates, Fannjiang, Jordan & Zrnic (2023). "Prediction-Powered Inference." Science 382.
The foundational recipe: a small human gold-standard sample debiases abundant model output to recover valid confidence intervals and p-values. The lesson that model output does not count one-for-one starts here.
Read →

The most complete treatment for social research formalizes the human-machine mix directly.

Peer-reviewed
Broska, Howes & van Loon (2025). "The Mixed Subjects Design: Treating Large Language Models as Potentially Informative Observations." Sociological Methods & Research 54(3).
Formalizes mixing human and LLM "silicon" subjects via prediction-powered inference and introduces the PPI correlation — a single reportable statistic that converts LLM informativeness into effective sample size and tells you how to split budget between humans and machines. Presumes every design still includes a human arm.
Read →

A cluster of 2025-26 work carries it into surveys.

Peer-reviewed
Krsteski et al. (2026). "Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification." ACL 2026 (Main).
Peer-reviewed. Combines LLM responses with a small human sample, reports the modest effective-sample-size gain that buys, and shows synthesis alone carries 24 to 86% bias.
Read →
Preprint
Ye & Yoganarasimhan (2026). "Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys." arXiv 2604.17267.
Turns the method into a budgeting rule for how many humans versus synthetics to collect, using residual variance after optimally using the LLM rather than raw LLM accuracy.
Read →
Preprint
Chen, Guo & Li (2026). "Power Analysis for Prediction-Powered Inference." arXiv 2603.16041.
Closed-form power and sample-size formulas, with an accompanying R package, for sizing the gold-standard human sample a rectified synthetic study needs — the required human count drops by roughly the model's R². Worked examples are biomedical rather than survey research.
Read →
Preprint
Ye, Lyu & Tao (2026). "Allocating Human Oversight in AI-Enabled Analytics." arXiv 2604.12497.
An online rule that spreads a fixed human-validation budget across survey questions as the model's reliability reveals itself, cutting wasted human sample severalfold versus uniform allocation. Its guarantees are asymptotic in budget.
Read →
Preprint
Tan & Zrnic (2026). "Valid Inference with Synthetic Data via Task Exchangeability." arXiv 2606.13629.
Recovers valid intervals by calibrating, on past tasks where real data existed, how wrong synthetic data tends to be, then widening to cover that gap.
Read →

For thresholds and quantiles, such as top-box scores, medians, and tail metrics, the newest and least-settled entrant takes on effective sample size under clustering directly.

Preprint
Noonan (2026). "The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering." Zenodo / arXiv 2608.21262 (preprint). · arXiv →
Estimates the effective sample size for threshold statistics directly, and is the first to frame the problem as an explicit design effect for correlated, model-generated data. Not yet peer reviewed, and written for ML, so applied here by analogy.
Read →

The pattern across all of it: none of these methods let synthetic data stand on its own. Each one recovers honest error bars by tethering the synthetic data to real human data. One source of confusion is worth clearing up. "Synthetic data" in the older, official-statistics sense is a different thing, with a mature variance apparatus that works because it is fit to real respondents.

Peer-reviewed
Reiter (2003). "Inference for Partially Synthetic, Public Use Microdata Sets." Survey Methodology 29(2).
The other "synthetic data": privacy-preserving microdata with decades-old variance-combining rules. Included to mark the distinction. That machinery does not transfer to LLM respondents, which aren't fit to a known real-data model.
Read →

So, is the synthetic result statistically significant? On its own you can't honestly say. A synthetic sample's row count is not its information content, and a valid answer needs a human anchor.

4. What can I reasonably infer today?

The profession has drawn its line, and it is cautious.

Guidance
AAPOR Task Force on Responsible AI Integration in Survey Research (Rothschild & Marlar, 2026).
Limits synthetic responses to "clearly labeled pretesting, pilot work, or exploratory diagnostics," notes that removing human respondents sits "outside the bounds of total survey error," and states the field "has not yet reached consensus on when, if ever, AI-generated responses can stand in for human ones."
Read →
Guidance
Sarstedt, Adler, Rau & Schmitt (2026). "Your Next Respondent Might Be an LLM: Guidelines for Using Silicon Samples in Marketing Research." NIM Marketing Intelligence Review 18(1).
Thoughtful validation guidance (compare means, variance, and range to human benchmarks) that still stops short of significance testing or an effective-sample-size method.
Read →

A useful vocabulary is emerging for what separates defensible practice from wishful practice.

Preprint
Hullman, Broska, Sun & Shaw (2026). "This human study did not involve human subjects: Validating LLM simulations as behavioral evidence." arXiv 2602.15785.
Splits current practice into heuristic validation (adjust prompts until outputs match) and statistical calibration (combine with human data under formal guarantees), and argues only the latter supports confirmatory claims. Stops short of prescribing holdout sizes or cadence.
Read →

As of this writing there is no standard, no accepted formula, and no software convention.

Open questions

Synthetic respondents are, today, a tool for direction, not for inference. Trust the direction, not the decimal. The methods that could change that (see section 3) are real but unsettled, and the questions below are where the field, and this page, will grow. We review this article regularly as part of our documentation review and update it as the literature moves.

  • A named design effect for synthetic respondents. Nobody has yet estimated the intra-cluster correlation and reported an explicit design effect for model-generated data. Noonan is the first move, and it is wide open.
  • A standard, reportable effective sample size. There is no agreed figure for what to divide a synthetic sample's row count by. Broska, Howes & van Loon's PPI correlation is the first reportable candidate, though it presumes a mixed human-machine design.
  • Validation protocols. How large a human holdout must be, and how often it is needed, to anchor a synthetic read. Chen, Guo & Li's power formulas and Ye, Lyu & Tao's allocation rule are first answers to the sizing and placement halves.
  • Quantifying bias. Bias has no volume-based cure. You cannot generate your way out of a model being wrong, so the frontier is measuring it, not removing it.
  • Reporting conventions. How synthetic-derived numbers should be labeled and caveated in a client deliverable.

Last updated: September 11, 2026. If you've published something we should include, get in touch.