Skip to content
Subconscious

Why Per-Test Bayesian Model Loops Stop Scaling

Many modest Bayesian A/B, ABC and ABCD datasets can spend substantial time in repeated compilation, warmup and adaptation. Vectorizing them may help, but the gain depends on dataset sizes and posterior geometry. Benchmark the actual workload before changing the pipeline.

The PyMC Labs HelloFresh case reports model redesign and unpooled vectorization. The author's forum clarification explains why the advantage depends on many small tests and can differ for large per-test datasets.

Where the original pipeline broke down

HelloFresh ran large numbers of A/B, ABC, and ABCD test campaigns at once, each fit as its own independent Bayesian model. Diagnosing the existing models turned up strong correlations between some posterior parameters and high autocorrelation in the MCMC chains, both of which slow convergence and cut the effective sample size a team gets for a given compute budget.

That combination of scale and inefficient sampling meant overnight batch runs took long enough that teams either waited past the point where results were still actionable, or limited how many tests they could run concurrently. Neither option scales with test volume.

Why did the team fix the model before optimizing speed?

The team's first move was a correctness pass, not a performance optimization. They proposed a structurally different Bayesian model that removed one parameter compared to the original formulation, used revised priors tailored to the problem, and applied the same structure consistently across A/B, ABC, and ABCD tests.

The new priors also matched domain knowledge: for an ABC test, most prior mass sat on scenarios where the variants' conversion rates were close together, not far apart, while still spreading roughly evenly across the 0–1 range. Under the new model, MCMC chains showed well-behaved mixing, and the correlation and autocorrelation problems that had been slowing convergence were resolved.

Before trusting the new model's output, the team ran parameter recovery simulations: known conversion probabilities were used to generate simulated data, and the posterior distributions were checked against those true values. For an ABC test, the posteriors were correctly centered on the ground truth. That result supports recovery for the reported simulation. It does not prove structural identifiability, calibration across parameter regimes or agreement with the real conversion process; those questions need separate checks.

At this stage, the corrected model alone was already faster: roughly 1.2x for A/B tests and roughly 2x for ABC and ABCD tests. Worthwhile, but the team judged it insufficient given the batch's scale.

Why didn't tuning knobs solve the speed problem?

Two conventional speed levers were tried next, and both fell short. Cutting MCMC tuning steps from the default 1000 down to 100 saved only about 0.1 seconds per A/B test, not enough at HelloFresh's batch scale. Avoiding model recompilation by defining one PyMC model with data swapped in through pm.Data containers, rather than rebuilding a model object per dataset, improved engineering hygiene but produced a negligible speed gain in practice.

Both attempts optimized how a single test was fit. The bottleneck was somewhere else: how many separate fits the pipeline ran.

What was the structural fix that solved the scaling problem?

The breakthrough was reframing how the batch was fit, not tuning any individual fit further. Instead of looping over datasets and running MCMC separately for each test, the team constructed one large unpooled model that fit all datasets simultaneously.

In this unpooled structure, each test keeps its own statistically independent parameters (nothing borrows strength across tests), but every test shares the same model structure, so the pipeline runs one warmup and adaptation phase and evaluates gradients across all tests in a single vectorized computation instead of restarting the sampler for each one. The pipeline stopped looping over small models entirely; PyMC compiled and sampled one model encoding every test at once.

The primary PyMC Labs case reports runtime improvements and parameter-recovery demonstrations for this workload. The forum provides a workload-dependence caveat, rather than the complete numerical benchmark. These are the reported results of that engineering engagement, not a Subconscious performance claim.

What generalizes beyond this one pipeline

The unified model also supported two-arm, three-arm, and four-arm comparisons under one consistent prior structure, so A/B, ABC, and ABCD tests could run through the same system instead of separate bespoke pipelines. That consistency simplified implementation and interpretation. Parameter recovery supplied a check on the tested configurations; broader accuracy and calibration claims require evidence across the intended test conditions.

The general lesson transfers past this specific stack: when Bayesian test volume grows, look for the structural bottleneck, how many independent models a pipeline is fitting, before assuming the answer is faster hardware or a lighter-weight per-test model.

Where this fits a causal-testing decision

This case study describes one vendor's engineering work on one customer's pipeline; it is not a Subconscious capability, customer result, or benchmark, and Subconscious does not claim to reproduce this exact pipeline, tooling, or result for any customer. What it illustrates is the argument behind treating causal experimentation as software infrastructure rather than a one-off analysis script: a testing system that has to run at scale needs structural design decisions, not just faster per-test loops. Teams evaluating a causal-testing platform can ask for diagnostics and a benchmark of their actual workload. For a modeled response study, assess agreement against an independent comparator that matches the audience, intervention and endpoint; confirm who will recruit participants, collect outcomes and report the comparison. Discuss the study and computing requirements.

Four-step chain: correlated posteriors and autocorrelation cause slow convergence, which makes runtime scale with test count, forcing a choice: throttle tests or ship late results.
The batch didn't get slow from scale alone, each model's own sampling inefficiency multiplied across every test in it.

Limitations

The engineering evidence comes from the HelloFresh/PyMC Labs case and applies to its workload. Conversion observations from a randomized live test are already measurements of actual people and can identify a treatment effect under the assignment, population and outcome-window assumptions. Additional checks are needed for untested populations, changed conditions or simulated outcomes; calling the data historical does not make a randomized test correlational.

For general background on the underlying test design, see A/B testing.

Five workload checks: diagnose per-test sampling, check a revised model, verify parameter recovery, compare loop and unpooled vectorization, and retain the faster valid configuration.
Benchmark the actual batch; the HelloFresh result for many modest datasets does not select the best configuration for every workload.