Skip to content
Subconscious

AI Research Ethics: A Practical Guide for Simulated Respondent Research

Presenting simulated-respondent findings to a board, an investor, or a launch committee is a disclosure decision, not just a data decision. Get it wrong and a good study becomes a credibility problem the moment someone asks who was actually surveyed.

Decision path: four situations where this guide recommends disclosing simulated data (shared outside the company, made public, hands off a decision, mixed with real data) versus internal-only uses where disclosure matters less.
This guide recommends disclosing simulated data when it leaves the team, goes public, hands off a decision, or mixes with real-respondent findings.

The decision this guide is about

Presenting undisclosed or overstated simulated findings as real customer data causes lasting credibility damage when it surfaces. A launch, price, or message built on biased or unvalidated calibration data can misdirect budget when nobody checks the finding first.

The safer path: an approach that discloses its simulated nature, states uncertainty honestly, and can be checked against real-human outcomes before it justifies a launch, pricing, or messaging decision. Subconscious's experiments report estimated effects with uncertainty intervals rather than a single plausible-sounding answer, and a team can separately scope a recruited human study of the same alternatives. This guide leaves some questions open: consent, anonymization, and the health of the real-participant ecosystem remain data-privacy and market-structure questions outside what any research platform can solve.

The July 2026 working paper reports parameter-rank agreement across a defined replication corpus, including mean Spearman correlation of 0.73 across 43 design-filtered studies. That is not an accuracy percentage or an outcome guarantee for a new decision. The working paper is not peer reviewed. Public paper.

When must you disclose that data is simulated?

This guide recommends disclosure in these cases; applicable legal, contractual and professional obligations must be assessed separately:

Disclosure matters less when the work stays inside the team:

The principle: anyone who might act on simulated research data should be told it is simulated. A team that later has to admit an undisclosed AI panel sat behind a major call won't get a second chance at that credibility.

Accuracy and misrepresentation

A simulated respondent's answer can sound entirely reasonable without being right. The obligation is to represent it for what it is: a modeled estimate, not a verified report of what real customers think.

Responsible framing:

Framing to avoid:

How does bias enter a simulated panel?

A simulated panel is built on data, and data carries the biases of its source: over-represented demographics or historical patterns carry through to the panel.

Bias typeHow it entersPractical risk
Selection biasCalibration data (e.g., CRM records) includes only customers who purchasedThe panel reflects survivors, not the people who considered and rejected the product
Demographic biasInterview transcripts or source data skew toward one gender, age group, or geographyThe panel carries the same skew, especially risky when the research is meant to represent a diverse population
Confirmation biasThe panel is built to represent what the team already believes about customersThe research becomes a mirror of existing hypotheses instead of a check on them

Mitigation:

What impact does simulated research have on real participants?

Replacing human fieldwork could change demand for participant recruitment. The following are possible effects to investigate, not a demonstrated forecast of market contraction:

This isn't a case against simulated research, but it matters for any organization that wants real-respondent infrastructure to stay available: lean on simulated methods too heavily and the ecosystem a team occasionally needs can break down.

Privacy in building a simulated panel

Calibrating a panel on customer data brings privacy questions with it, GDPR among them. Worth weighing:

Determine whether personal data and the GDPR’s territorial scope apply, then assess purpose, lawful basis, minimization and applicable rights. Consent is one possible lawful basis; using simulated respondents does not remove obligations attached to identifiable source data. GDPR, Articles 3, 5, 6 and 17.

A practical framework

Before building a panel

During research

When presenting findings

What does the professional code say?

The 2025 ICC/ESOMAR International Code is a professional self-regulatory code. It binds ESOMAR members and the associations that adopt it; it is not universal law. Article 7(e) addresses notifying the client and keeping human oversight. Article 9(b) addresses disclosure when significant published findings rely on AI or synthetic data, human oversight, and providing technical information on reasonable request. Check the code's own text and your contracts, because they set the duties that apply to your work.

The opportunity is real: an early read on options before a team commits to a full study. The risk is just as real: used carelessly, simulated research produces bad decisions and credibility damage that sets back the whole method. Rigor about disclosure and accuracy is what makes that value durable enough to keep using.

Limitations

This framework addresses disclosure, accuracy, and bias in using simulated respondents. It does not resolve consent and anonymization questions in the underlying data, which are legal and governance questions specific to each organization, and it does not restore the real-participant ecosystem on its own. Simulated panels, including experiments run on Subconscious, remain a modeled estimate until checked against real-human outcomes. The leaderboard reports aggregate parameter-rank results and their limits.