How to frame a causal behavioral experiment
A useful experiment starts with a decision, not a broad request for insights. Define what the team may change, whose behavior matters, which alternatives need comparison, and the context in which the decision will occur.
Step 1: Name the intervention contrast
Write a question about a change the team can make and the response it will measure. State whether that response will be generated by a model, chosen by a person in a hypothetical task, or observed in a live setting.
For example:
- How does fuel efficiency affect car choice?
- Which product claim changes purchase intent among a defined buyer segment?
- Which message changes support for a proposed policy?
Avoid questions so broad that no experiment can distinguish one action from another. “What causes car buying?” may help open a discussion, but a study needs explicit alternatives and an observable outcome.
Who should the experiment include?
Specify the population whose response matters to the decision. Useful characteristics may include profession, income, age, current behavior, or another trait tied to the study.
Define eligibility and traits needed for population matching or planned subgroup analysis. Record the dates and sources behind those traits. Confirm any current interface limits when configuring the study; the number of fields is not a general rule for a valid population definition.
Synthetic or simulated participants compare aggregate patterns, but should not be treated as exact replicas of individuals. Human baselines and validation remain important, especially when the decision affects vulnerable groups or carries high stakes.
What should the experiment compare?
List the alternatives, attributes, or claims the experiment will compare. Each attribute needs concrete levels that participants can evaluate.
For a product study, these might include:
- Product concepts
- Features or claims
- Price points
- Packaging or message options
Select attributes using customer evidence, qualitative pretesting, and the decision to be made. Check feasible levels and whether the design can separate their effects. An attribute's effect is a question for the study rather than something already known at setup.
Discrete-choice experiments compare choices among stated alternatives. Their attribute-and-level designs should support the parameters the study intends to estimate (ISPOR experimental-design report). A hypothetical human choice or a generated choice is not an observed sale. Specify the design and endpoint before naming the output a conjoint, MaxDiff, pricing, or portfolio result.
When and where should the experiment take place?
Define the time and place that frame the audience's decision.
- When: Use a specific year such as 2024 or a broader period such as the early 80s when the historical context matters. Past and present periods are easier to ground than unsupported future conditions.
- Where: Choose the country, market, or other geographic scope relevant to the audience.
Time and place are part of the experiment, not decorative context. A response that is plausible in one market or period may not transfer to another.
Review the design before running it
In a hypothetical onboarding study, compare the current self-service offer with the same offer plus an assisted setup session. Randomly assign eligible buyers to the two descriptions and measure their stated selection. This can estimate an effect on that selection task; a separate live test would be needed to measure actual renewal or purchase. Before running the study, check:
- Are the intervention and baseline explicit, including the status quo where relevant?
- Are eligibility, population coverage, and the response source specified?
- What is randomized: a person, account, task, or generated draw? How are versions assigned, and can repeated responses or spillovers affect the comparison?
- Are the decision period, geography, and information shown to participants defined?
- Is the endpoint generated choice, hypothetical human selection, or an observed action, and how will it be recorded?
- Are the estimator, units, uncertainty assumptions, and independent validation criteria agreed before results are read?
Resolve unclear assignment, population, or outcome definitions before interpreting the comparison. Randomization addresses the tested intervention contrast; it does not automatically establish representative coverage or predict sales in another setting.
See how this framing step fits into the full testing process on How We Work, or bring a live decision to a demo to work through it directly.