AI-Generated Models That Run vs. Models You Can Trust
An AI-generated model deserves a closer look only after it clears a viability gate checking convergence and usable posteriors, and then scores well on a documented quality rubric covering completeness, model fit, and adherence to best practice. Those two checks are necessary. They are not sufficient: a domain expert still has to confirm the model answers the decision in front of you.
The decision: trust the output, or gate it first
An AI coding agent hands back a statistical model. It ran. No errors, a finished script, a results object. Does that mean the model is ready to inform a pricing decision, a risk score, or a clinical read?
A benchmark study of Claude Code building Bayesian models offers a direct answer: no. Code that executes and a model fit to inform a decision are two different claims. A model with divergent MCMC chains, a mis-specified prior, or label-switching in a mixture model can look identical to a sound one until someone checks the diagnostics. The failure does not show up as a bug report. It shows up later, as a wrong call that nobody can trace back to its source.
What did the benchmark measure?
Christopher Fonnesbeck of PyMC Labs published the benchmark on 26 February 2026 (PyMC Labs, "Measuring Reliability in AI-Assisted Bayesian Modeling"). It tested two conditions with Claude Code: a base agent working from general training knowledge, and the same agent with a written domain-knowledge "skill" injected into its context, covering current best practice for model specification, parameterization, sampling configuration, and convergence diagnostics. Both conditions used identical prompts, tools, and CLI flags, and each task ran three times per condition, 30 runs in total.
Five tasks escalated in difficulty: a hierarchical model, ordinal regression, a stochastic-volatility model, a Gaussian mixture model, and a sparse variable-selection model. Each targets a point where a wrong implementation choice, not a syntax error, produces code that runs but gives unreliable answers.
Rather than score every run on one number, the study split evaluation into two stages:
| Stage | Question it answers | What it checks |
|---|---|---|
| 1. Viability gate (pass/fail) | Did the model produce output worth examining at all? | Sampling completed with a usable posterior, convergence diagnostics inside an acceptable range, and no degenerate estimates |
| 2. Quality scoring (0–5 per criterion) | Among the models that pass, how good is this one? | Completeness of the output, convergence quality, whether the model choice fits the problem, adherence to modern best practice, how many rewrite cycles it took, and how many turns it used |
That two-stage design is the transferable lesson, independent of which tool produced the model. A single quality average hides failure: if a third of runs never clear the viability bar, their absence quietly inflates the reported average for everyone left standing.
Where base knowledge runs out
A benchmark that reports only its wins is marketing copy. This one also names where the two conditions tied: both conditions passed the hierarchical model and ordinal regression tasks at the same rate, because those patterns are common enough in training data that domain augmentation adds little.
The gap opened on the harder, less-documented tasks:
- On stochastic volatility, the unaided agent produced no viable run in three attempts (0%), reaching for a manually built autoregressive parameterization that never converged. With the domain skill, two of three runs (67%) were viable.
- On the horseshoe variable-selection task, the unaided agent passed one of three runs (33%); with the skill, all three passed.
- On the Gaussian mixture model, viability rose from 67% to 100%, and the convergence-quality criterion (scored 0 to 5) rose from 2.7 to 5.0. The author attributes the gain to applying an ordered transform that prevents label-switching, rather than sorting the draws after the fact. That mechanism is the author's explanation for a three-run comparison, not a separately tested cause.
Across all 30 runs, the viability pass rate was 60% without the skill and 93% with it. With three runs per cell, treat the task-level numbers as illustrations. The author's own summary is more careful than a pattern claim: the skill made the agent more consistent about reaching for the parameterization a domain expert would already use.
What does the code actually differ on?
On the sparse variable-selection task, the two conditions produced structurally different models. The two snippets below are parameterization fragments from those fuller models. Without the domain document, Claude built a standard horseshoe prior with a centered parameterization:
with pm.Model() as model:
tau = pm.HalfCauchy("tau", beta=1)
lambda_i = pm.HalfCauchy("lambda", beta=1, shape=n_predictors)
sigma_beta = tau * lambda_i
beta = pm.Normal("beta", mu=0, sigma=sigma_beta, shape=n_predictors)
eta = pm.math.dot(X, beta)
trace = pm.sample(2000, tune=1000, target_accept=0.95, random_seed=42,
chains=4, return_inferencedata=True)With the document available, it built a regularized horseshoe with a slab component and a non-centered parameterization:
with pm.Model(coords=coords) as model:
X_data = pm.Data("X", X_scaled, dims=("obs", "features"))
tau = pm.HalfStudentT("tau", nu=2, sigma=1)
lam = pm.HalfStudentT("lam", nu=5, dims="features")
c2 = pm.InverseGamma("c2", alpha=1, beta=1)
z = pm.Normal("z", mu=0, sigma=1, dims="features")
lam_tilde = pt.sqrt(c2 / (c2 + tau**2 * lam**2))
beta = pm.Deterministic("beta", z * tau * lam * lam_tilde, dims="features")
idata = pm.sample(draws=1000, tune=1000, chains=4, nuts_sampler="nutpie",
random_seed=42, target_accept=0.95, init="adapt_diag")These are parameterization fragments, not complete runnable models: the likelihood, observed data and other setup are omitted, so sampling them as shown would not fit anyone's data. The PyMC Labs post supplies the study context, and the pymc-modeling skill it describes is public in pymc-labs/python-analytics-skills. The post does not link a separate repository of the benchmark's complete model code. The second fragment adds a slab term that keeps shrinkage from producing implausibly large coefficients, decouples the coefficients from the shrinkage scale for easier sampling geometry, and labels its dimensions so the diagnostics are interpretable later.
The lesson generalizes past this one benchmark
This is a third-party benchmark of one coding agent building Bayesian models with the PyMC package, not a Subconscious study, and its specific pass rates and costs describe that setup, not any general guarantee. What generalizes is the discipline: any AI-assisted analysis that will inform a real decision needs an explicit pass/fail gate before a quality score, and a documented rubric for what "good" means once something clears that gate.
That is the same discipline behind causal experiment design: a result is not useful because a model produced a number, it is useful because the design, the diagnostics, and the uncertainty around that number can be checked. Subconscious's own method runs randomized discrete choice experiments on a simulated population and reports estimated effects with uncertainty rather than a single point estimate, so a decision-maker has something to check before acting on it. Those effects are modeled stated-choice effects. How it works describes the design, and the replication leaderboard shows the published fidelity evidence (rank agreement on choice parameters across studies, from a working paper that is not peer reviewed).
Where this comparison stops applying
A gate is trustworthy only when its blind spot sits in the open next to it. This section names that blind spot directly: a viability gate catches models that fail outright or produce degenerate estimates. It does not catch a model that passes every diagnostic while answering the wrong business question, and no automated check replaces a domain expert reviewing whether the model specification matches the decision it will inform. The benchmark's own quality scoring needed a rubric written by someone who knew what "appropriate" looked like for each task; a gate without that judgment behind it is just a lower bar dressed up as assurance.
Next step
Treat "it ran" as the start of the checklist, not the end of it. For how a documented, checkable process applies to causal decisions specifically, see how Subconscious runs an experiment, or book a decision review to walk through a model your team is about to rely on.