How to Sequence a Year of Research Into One Experiment Roadmap: 10 Steps
A research leader who runs one study at a time re-answers the same question every quarter. A launch decision needs a segment read, so a study gets commissioned. A pricing decision needs another read three months later, and nobody connects it to what the first study already showed about the same buyers. Phase 2 gets designed without phase 1's evidence, and the program can never show which action moved which outcome over a year of decisions.
Build one prioritized program around decisions, dependencies, evidence requirements, and owners. Earlier findings can refine a later study when there is a real dependency; independent questions can run in parallel. The ten steps below are a proposed planning workflow.
1. Define the target goal
Start with the business outcome and the action the program can change. For an app team, a goal might be to improve retained active use among a specified segment, with a named measurement period and minimum worthwhile change. "Which features matter?" is a research question within that goal rather than the outcome itself.
2. Break the goal into research questions
Each experiment should answer one piece of the goal, and together they should cover it. For a customer loyalty goal, that could break down into three questions: which messaging drives repeat purchases, whether easier navigation improves retention, and which payment methods lift checkout completion. Rank the questions by impact and feasibility to set the run order.
3. Design each experiment around a specific hypothesis
Give each study a testable question, baseline and alternatives where appropriate, and a named endpoint. Choose enough independent variation to identify the required effects; a fixed one- or two-variable limit is not a general rule. Predeclare the analysis and decision threshold. ISPOR's experimental-design guidance explains why identification matters before efficiency.
4. How do you map the experimental path into phases?
A hypothetical annual plan could use the following sequence. Quarter labels and deadlines are planning assumptions, not delivery estimates:
| Planning period | Decision and owner | Evidence and prerequisite | Advance, adapt, or stop |
|---|---|---|---|
| Q1, before feature budget approval | Product lead selects candidate features | Research owner reviews usage records and human task comprehension; choice screening can narrow defined offers | Retain candidates when evidence is unresolved; stop concepts that are infeasible |
| Q2, before onboarding release | Product and engineering leads choose navigation | Human usability checks followed by a live randomized activation test; requires a working prototype and valid instrumentation | Revise task failures; ship only against the predeclared activation criterion |
| Q3, before campaign booking | Marketing lead chooses messages | Randomized message comparison with a named response; requires stable offers and channel constraints | Require applicable validation and sufficient uncertainty bounds; keep baseline if inconclusive |
| Q4, before pricing approval | Pricing lead selects an offer | Structured choice evidence plus a live renewal or purchase comparison where required; requires agreed offers and billing feasibility | Use the margin and retention decision rule; stop or adapt if behavioral validation fails |
5. Design for flexibility
Set review points and allowable adaptations before results arrive. A surprising result can change priorities or motivate a new study, but changing endpoints or repeatedly testing until a favorable result appears requires an updated analysis plan and fresh validation.
6. Test product, pricing, and GTM actions before committing capital
Map each roadmap question to the method and response it needs. A Subconscious generated-choice comparison can screen stated alternatives. Human usability work can inspect navigation and comprehension; live randomized studies can measure activation, retention, or checkout completion. Descriptive records can establish a baseline. None of these endpoints should be substituted for another merely because they appear in the same roadmap.
7. Standardize data collection and analysis across phases
Use a common evidence record rather than require identical data across different questions. For each study, record population, dates, source, assignment, alternatives, endpoint and units, estimator, uncertainty assumptions, and validation. Compare numerical effects only when those definitions support comparison; otherwise use each result for its own decision.
8. How do you review and adapt after each phase?
After a batch of experiments closes, check whether the results actually answered the research questions for that phase, and whether the next phase's plan still makes sense in light of them. Assumptions sometimes turn out wrong, or a phase surfaces a question nobody planned for; refining the plan mid-program is the point of running it as a sequence rather than a single study.
9. How do you keep one central record of hypotheses, experiments, and results?
Use the team's existing shared research record to connect hypotheses, designs, sources, results, decisions, and owners. Include inconclusive and contrary findings, changes to the plan, and evidence that no longer applies. A new dashboard is unnecessary if the existing record makes those dependencies clear.
10. Synthesize the program back into the original goal
Synthesize the program against the original business goal. State what each study established, what action followed, and which later outcome was actually observed. Keep unresolved questions visible and distinguish forecasted benefit from measured business impact.
Where this breaks down
A dependency-based program fails when a prerequisite is missing, evidence has become stale, or results from different endpoints are combined as if they measured one effect. Recheck the intended use of each study. A generated contrast can inform a screening decision; actual activation, retention, or purchase requires evidence on those behaviors.
Confidence intervals, segment-level heterogeneity, and automated recommendations are not standard outputs of every study; they depend on how a given phase is configured. Treat this roadmap as the discipline for sequencing decisions, and treat each phase's specific configuration as the place to confirm what that phase will and won't report.
To design the causal experiments for a specific phase, or to see how the sequencing and synthesis work in practice, book time to walk through a roadmap.