
If there was one thing everyone agreed on at our recent MRS and RSS event, Beyond the Hype: Synthetic Data in Practice, it was this: synthetic data is no longer a future concept.
Organisations are experimenting with it, clients are asking about it, and researchers are under pressure to understand both its potential and its limitations. What remains less certain is where it genuinely fits in the research toolkit, and where the hype still outweighs the evidence.
Across an afternoon of candid presentations and debate, researchers, practitioners, suppliers and client-side teams shared what they are testing in the real world. The result was not one neat story, but a range of approaches, outcomes and unanswered questions.
One important starting point was that “synthetic data” has become an umbrella term. Data imputation, fusion and augmentation sit alongside digital twins, persona bots and synthetic respondents, yet they rely on different techniques, require different validation and solve different problems.
So the useful question is not simply, “Does synthetic data work?” It is, “Which type, for which use case, and under what conditions?” The speakers’ findings showed just how much the answer can change depending on the problem being tackled.
NielsenIQ’s James Pitcher examined whether large language models could outperform established statistical approaches when imputing missing survey data. His results prompted a wider discussion about what generative AI is good at, where conventional methods retain an advantage, and why methodological fit matters more than novelty.
The session offered a useful warning against reaching for AI simply because it is available. But the detail behind the comparison, including how the approaches performed across different study designs, question types and levels of missing data, is where the most valuable learning lies.
Synthetic boosting, generating additional records to enlarge a dataset, has an obvious appeal: the prospect of smaller samples, lower costs and faster results. In practice, the discussion revealed a much more complicated picture.
Simon Harris from MMR Research shared extensive testing of whether small samples could be boosted to offer the confidence and decision-making reliability of larger human samples. His findings raised important questions about seed data, representation and what adding synthetic records can, and cannot, change.
This led to one of the day’s most thought-provoking challenges: are organisations turning to synthetic data to solve research problems, budget problems, or both? The speakers did not offer an easy answer, but their evidence gives research buyers plenty to consider.
If the event had an unofficial strapline, it was “validate everything”. Optimists and sceptics alike stressed that plausible-looking outputs are not proof of quality.
Tabita Razaila from STRAT7 shared year-on-year benchmarking against real-world holdout data, while Ipsos discussed applying synthetic methodologies within large-scale tracking programmes. Their examples showed why headline results may not tell the whole story, particularly once segmentation, behavioural drivers and trends are examined more closely.
A recurring tension was how teams can move quickly without lowering the evidential bar. Several examples showed that validation is not a single final check, but something that needs to be designed around the intended decision, the available benchmark and the risks of getting it wrong.
What should robust validation look like? Which comparisons are most revealing? And how can researchers spot an output that appears convincing but is not decision-ready? The presentations explored these questions in much greater depth than a short summary can capture.
Ipsos’ David Priestley and Chris Moore began their case study with a deceptively simple question: what problem are we actually trying to solve? A desire to use synthetic data is not, by itself, a use case.
Their work within a complex multi-country tracking study illustrated how the structure and quality of the source data shape what is possible. Sparse variables, intricate routing and thousands of interconnected fields introduce challenges that a model cannot simply wish away.
The team also drew a valuable distinction between synthetic data as a “data printer” and as a modelling tool. Their account of what happened when they simplified the problem offered one of the afternoon’s most practical lessons, and one best understood through the full case study.
The conversation shifted when James Tyrrell from Virgin and Leanne Tomasevic from Electric Twin presented their work with synthetic audiences. Rather than focusing on creating more records, their approach uses digital representations built from rich attitudinal and behavioural data to explore ideas and concepts quickly.
This changed the debate from statistical accuracy alone to practical usefulness. Could synthetic audiences help teams decide which ideas deserve further investigation? Where might they support early exploration, and where must direct customer research remain central? The session offered a different lens on value, without presenting synthetic approaches as a replacement for people.
By the closing panel, the debate had moved beyond methodology to decision-making. What confidence is required? What is the cost of being wrong? And what does responsible use look like in different contexts?
One of the most interesting reframings was that the comparison may not always be synthetic insight versus traditional research. Sometimes it may be synthetic insight versus no customer input at all. That opens up new possibilities, but also demands clearer boundaries around what the output should inform.
The discussion also showed why broad claims for or against synthetic data are unhelpful. The same approach may be useful for one stage of a project and unsuitable for another. Context, purpose and the consequences of error all matter, which is why the speakers’ differing perspectives are worth hearing in full.
Nobody left believing synthetic data was either a miracle solution or a passing fad. The picture was more balanced: some applications show promise, others remain experimental, and every approach depends on a clear question, strong source data and appropriate scrutiny.
For researchers, the reassuring message was that expertise still matters. Judgement, critical thinking, methodological rigour and an understanding of human behaviour remain essential, particularly as the technology develops.
The event did not close the debate. It made it more useful. To hear the evidence behind the conclusions, explore the contrasting case studies and decide where you stand, watch the full recording.
Couldn’t join us on the day, or want to revisit the discussion?
MRS members can access the recording for free here: LINK
Non-members can request access to purchase the recording for £30 + VAT by here: LINK
From augmentation and synthetic audiences to digital twins and the future role of researchers, the recording lets you hear the debate first-hand, examine the evidence and make up your own mind about what comes next.
Our newsletters cover the latest MRS events, policy updates and research news.
0 comments