AI Research Ethics: A Practical Guide for Simulated Respondent Research
Presenting simulated-respondent findings to a board, an investor, or a launch committee is a disclosure decision, not just a data decision. Get it wrong and a good study becomes a credibility problem the moment someone asks who was actually surveyed.
The decision this guide is about
Presenting undisclosed or overstated simulated findings as real customer data causes lasting credibility damage when it surfaces. A launch, price, or message built on biased or unvalidated calibration data can misdirect budget when nobody checks the finding first.
The safer path: an approach that discloses its simulated nature, states uncertainty honestly, and can be checked against real-human outcomes before it justifies a launch, pricing, or messaging decision. Subconscious's experiments report estimated effects with uncertainty intervals rather than a single plausible-sounding answer, and a team can separately scope a recruited human study of the same alternatives. This guide leaves some questions open: consent, anonymization, and the health of the real-participant ecosystem remain data-privacy and market-structure questions outside what any research platform can solve.
The July 2026 working paper reports parameter-rank agreement across a defined replication corpus, including mean Spearman correlation of 0.73 across 43 design-filtered studies. That is not an accuracy percentage or an outcome guarantee for a new decision. The working paper is not peer reviewed. Public paper.
When must you disclose that data is simulated?
This guide recommends disclosure in these cases; applicable legal, contractual and professional obligations must be assessed separately:
- Sharing the work with anyone outside the company, including investors, partners, or regulators
- Putting findings in front of the public, whether that's a blog post, a press release, or an industry report
- Handing over a decision for someone else to independently weigh the evidence behind
- Mixing simulated results into the same analysis as real-respondent data
Disclosure matters less when the work stays inside the team:
- Forming hypotheses that no one outside the team will see
- Narrowing a list of concepts before a real study gets funded
- Running practice exercises like sales roleplay or a stakeholder simulation
The principle: anyone who might act on simulated research data should be told it is simulated. A team that later has to admit an undisclosed AI panel sat behind a major call won't get a second chance at that credibility.
Accuracy and misrepresentation
A simulated respondent's answer can sound entirely reasonable without being right. The obligation is to represent it for what it is: a modeled estimate, not a verified report of what real customers think.
Responsible framing:
- "In this segment, our simulated research panel points toward a positive response."
- "Pricing came up repeatedly as a concern across the simulated panel."
- "Simulated customer scenarios point to two likely objections, X and Y."
Framing to avoid:
- "Customers say they want this." (Suggests real customers were the ones asked.)
- A bare sentiment percentage with no defined method behind it. (Implies quantitative rigor the underlying research doesn't support.)
- "The study proves this direction." (Claims validation without naming the evidence or method.)
How does bias enter a simulated panel?
A simulated panel is built on data, and data carries the biases of its source: over-represented demographics or historical patterns carry through to the panel.
| Bias type | How it enters | Practical risk |
|---|---|---|
| Selection bias | Calibration data (e.g., CRM records) includes only customers who purchased | The panel reflects survivors, not the people who considered and rejected the product |
| Demographic bias | Interview transcripts or source data skew toward one gender, age group, or geography | The panel carries the same skew, especially risky when the research is meant to represent a diverse population |
| Confirmation bias | The panel is built to represent what the team already believes about customers | The research becomes a mirror of existing hypotheses instead of a check on them |
Mitigation:
- Diversify calibration data sources rather than relying on one channel.
- Deliberately include perspectives underrepresented in the source data.
- Regularly compare simulated responses to real customer feedback to catch drift.
- Document data sources and known limitations behind each panel.
What impact does simulated research have on real participants?
Replacing human fieldwork could change demand for participant recruitment. The following are possible effects to investigate, not a demonstrated forecast of market contraction:
- The supplemental income people earn as professional respondents dries up
- Platforms built around recruiting participants see fewer requests
- The pipelines that reach real respondents wither from disuse
- Finding real participants gets harder for a team that still needs them
This isn't a case against simulated research, but it matters for any organization that wants real-respondent infrastructure to stay available: lean on simulated methods too heavily and the ecosystem a team occasionally needs can break down.
Privacy in building a simulated panel
Calibrating a panel on customer data brings privacy questions with it, GDPR among them. Worth weighing:
- Consent. Does the consent attached to the underlying data extend to this use? A transcript gathered under a general "research purposes" banner might or might not stretch to cover calibrating a simulated panel.
- Anonymization. Is the panel built from aggregated, anonymized data, or does it represent identifiable individuals? Modeling a named customer raises different questions than modeling "enterprise buyers in the fintech sector."
- Data minimization. Is the team using only the data necessary for calibration, or feeding in everything available? If GDPR applies, its minimization principle applies either way.
- Right to deletion. Can the team comply if a customer whose data fed calibration invokes their right to erasure?
Determine whether personal data and the GDPR’s territorial scope apply, then assess purpose, lawful basis, minimization and applicable rights. Consent is one possible lawful basis; using simulated respondents does not remove obligations attached to identifiable source data. GDPR, Articles 3, 5, 6 and 17.
A practical framework
Before building a panel
- Audit data sources. Know what goes in, whether it was properly consented, and whether it's missing demographic groups or carrying other biases.
- Define the use case. What decision will this inform? Does it need real-respondent rigor, or is a simulated pass appropriate first?
- Establish disclosure norms as a team, in writing, before anyone needs to decide in the moment.
During research
- Label everything. Put "Simulated Panel Research" in the document title from the start; a footnote is too easy to miss.
- Watch for confirmation bias. If the panel tells the team exactly what it wanted to hear, treat that as a flag to probe further, not as confirmation.
- Document limitations. Every output should state what the research can and cannot tell the team.
When presenting findings
- Disclose by default, unless there's a specific reason not to (internal ideation, informal exploration).
- Present accurately, with language that reflects the data's nature rather than implying rigor or validation it doesn't have.
- For high-stakes decisions, recommend real-participant validation as an explicit follow-up step rather than letting a simulated finding carry more weight than it should.
What does the professional code say?
The 2025 ICC/ESOMAR International Code is a professional self-regulatory code. It binds ESOMAR members and the associations that adopt it; it is not universal law. Article 7(e) addresses notifying the client and keeping human oversight. Article 9(b) addresses disclosure when significant published findings rely on AI or synthetic data, human oversight, and providing technical information on reasonable request. Check the code's own text and your contracts, because they set the duties that apply to your work.
The opportunity is real: an early read on options before a team commits to a full study. The risk is just as real: used carelessly, simulated research produces bad decisions and credibility damage that sets back the whole method. Rigor about disclosure and accuracy is what makes that value durable enough to keep using.
Limitations
This framework addresses disclosure, accuracy, and bias in using simulated respondents. It does not resolve consent and anonymization questions in the underlying data, which are legal and governance questions specific to each organization, and it does not restore the real-participant ecosystem on its own. Simulated panels, including experiments run on Subconscious, remain a modeled estimate until checked against real-human outcomes. The leaderboard reports aggregate parameter-rank results and their limits.