Skip to content

AI Research Ethics: A Practical Guide for Simulated Respondent Research

Presenting simulated-respondent findings to a board, an investor, or a launch committee is a disclosure decision, not just a data decision. Get it wrong and a good study becomes a credibility problem the moment someone asks who was actually surveyed.

Decision path: four triggers requiring disclosure (shared outside company, made public, hands off a decision, mixed with real data) versus internal-only uses where disclosure matters less.
Disclose simulated data the moment it leaves the team, goes public, hands off a decision, or mixes with real-respondent findings.

The decision this guide is about

Presenting undisclosed or overstated simulated findings as real customer data causes lasting credibility damage when it surfaces. A launch, price, or message built on biased or unvalidated calibration data can misdirect budget when nobody checks the finding first.

The safer path: an approach that discloses its simulated nature, states uncertainty honestly, and can be checked against real-human outcomes before it justifies a launch, pricing, or messaging decision. Subconscious's experiments report causal effects with confidence intervals rather than a single plausible-sounding answer, and a team can move from a simulated study to real-human validation without changing the causal question. It doesn't resolve every question in this guide: consent, anonymization, and the health of the real-participant ecosystem remain data-privacy and market-structure questions outside what any research platform can solve.

The published validation evidence shows our best configuration reaches 87% of the measured human ceiling on one study: 0.832 rank correlation against the published human result, where two independent samples of real humans reach 0.959. Across all 43 studies that pass design filters the mean is 0.73, detailed in the causal fidelity paper. This is corpus-level evidence, not a guarantee for a new market or decision.

When must you disclose that data is simulated?

Disclosure is mandatory in these cases:

Disclosure matters less when the work stays inside the team:

The principle: anyone who might act on simulated research data has a right to know it's simulated. A team that later has to admit an undisclosed AI panel sat behind a major call won't get a second chance at that credibility.

Accuracy and misrepresentation

A simulated respondent's answer can sound entirely reasonable without being right. The obligation is to represent it for what it is: a modeled estimate, not a verified report of what real customers think.

Responsible framing:

Framing to avoid:

How bias enters a simulated panel

A simulated panel is built on data, and data carries the biases of its source: over-represented demographics or historical patterns carry through to the panel.

Bias typeHow it entersPractical risk
Selection biasCalibration data (e.g., CRM records) includes only customers who purchasedThe panel reflects survivors, not the people who considered and rejected the product
Demographic biasInterview transcripts or source data skew toward one gender, age group, or geographyThe panel carries the same skew, especially risky when the research is meant to represent a diverse population
Confirmation biasThe panel is built to represent what the team already believes about customersThe research becomes a mirror of existing hypotheses instead of a check on them

Mitigation:

Impact on real research participants

When simulated respondents take over a large share of work that once relied on real participants, the participant-recruitment market contracts. Downstream effects can include:

This isn't a case against simulated research, but it matters for any organization that wants real-respondent infrastructure to stay available: lean on simulated methods too heavily and the ecosystem a team occasionally needs can break down.

Privacy in building a simulated panel

Calibrating a panel on customer data brings privacy questions with it, GDPR among them. Worth weighing:

Companies based in Europe, or serving European customers, don't get to treat these as optional: they're legal obligations that fall to the team's own data governance, not something a research platform can settle.

A practical framework

Before building a panel

During research

When presenting findings

Industry standards are still forming

The market research industry is developing standards for simulated research: professional bodies drafting guidelines, academic institutions studying accuracy, and regulators watching closely. Teams that adopt disciplined practices now will be ahead when formal standards arrive.

The opportunity is real: faster, more accessible first-pass evidence for teams that previously couldn't afford to test every decision. The risk is just as real: used carelessly, simulated research produces bad decisions and credibility damage that sets back the whole method. Rigor about disclosure and accuracy is what makes that value durable enough to keep using.

Limitations

This framework addresses disclosure, accuracy, and bias in using simulated respondents. It does not resolve consent and anonymization questions in the underlying data, which are legal and governance questions specific to each organization, and it does not restore the real-participant ecosystem on its own. Simulated panels, including experiments run on Subconscious, remain a modeled estimate until checked against real-human outcomes. See the leaderboard for method-by-method results.