A validity framework for synthetic research.
The Fallacy: Matching surface-level human survey results (marginal fit) is easy to fake, since AI personas often output plausible headline numbers while masking incoherent internal belief systems.
Human-Calibrated Consistency: Real humans reproduce their answers only ~91% of the time, so valid synthetic personas must target human-calibrated variance rather than static, 100% agreement.
Three Pillars of Validity: True simulation accuracy requires evaluating internal validity (stable internal beliefs), construct validity (authentic psychometric covariance across hidden traits), and external validity (replicating real-world causal effect sizes on held-out human data).
Beyond the Industry Baseline: While most providers evaluate models on two basic distributional benchmarks, rigorous enterprise testing should use eight validity tests to ensure reliable predictions for high-stakes decisions.
Our newsletters cover the latest MRS events, policy updates and research news.