Compiling Code Is Not Validating a Model
A generated PyMC model can compile while using the wrong likelihood, an unjustified prior or an unsupported causal interpretation. Compilation checks whether model construction succeeds; it does not establish sampling convergence or scientific validity.
- Check construction and sampling separately.
- Review the observation process before choosing a likelihood.
- Test prior sensitivity and predictive behavior.
- Match external validation to the decision and endpoint.
What did ModelCraft demonstrate?
Bernard Mares, Allen Downey and Alexander Fengler describe ModelCraft in PyMC Labs' July 1, 2025 walkthrough. The prototype compiles candidate models in a sandbox without sampling, returning tracebacks for revision. Its capture-recapture example initially failed with:
TypeError: HyperGeometric.dist() missing 1 required positional argument: 'k'What does the example specify?
The revised example used 25 tagged bears, a second capture of 20, and 4 tagged recaptures, with a population upper bound of 500. These are tutorial inputs, not a field estimate. This model-construction fragment preserves the example's prior and likelihood; it is not a complete sampling script:
import pymc as pm
n1, n2, k_observed, N_max = 25, 20, 4, 500
with pm.Model() as capture_model:
N = pm.DiscreteUniform("N", lower=max(n1, n2), upper=N_max)
recaptures = pm.HyperGeometric(
"recaptures", N, n1, n2, observed=k_observed
)The source's sampling example requests 3,000 draws after 1,000 tuning steps. A sampling request is not evidence of convergence. The authors also report that requiring type annotations produced more convoluted categorical constructions for a separate Think Bayes problem; removing that instruction favored a concise Dirichlet formulation. Complexity alone does not prove mathematical invalidity.
What does each check establish?
| Check | What it can establish | What remains open |
|---|---|---|
| Model construction | Arguments, imports and shapes accepted on that path | Whether distributions fit the observation process |
| Sampling diagnostics | Evidence about exploration and convergence | Whether a converged model is scientifically appropriate |
| Prior and posterior predictive checks | Whether generated data reproduce relevant features | Whether omitted mechanisms matter to the decision |
| Sensitivity analysis | How conclusions change under alternative assumptions | Which assumptions deserve substantive support |
| External comparison | Agreement on a specified population and endpoint | Transfer to other populations or interventions |
How should you review the capture-recapture assumptions?
A hypergeometric observation model treats the second capture as sampling without replacement from a population containing a fixed number of marked animals. Ask whether the population remained closed between captures, marks were retained and recognized, and capture probabilities were compatible with that sampling assumption. Movement, tag loss or systematically different recapture probabilities can undermine the interpretation even when the code runs.
The upper bound also matters. A uniform prior truncated at 500 is an assumption about plausible population size, not a compiler finding. Check whether posterior mass accumulates near the bound and whether alternative defensible bounds or priors change the decision. Verify the likelihood's support against observed counts before sampling.
After fitting, inspect chain mixing, convergence diagnostics and effective sample sizes suitable for the sampler. Use prior predictive checks to examine assumptions before conditioning on observations, and posterior predictive checks to compare generated and observed data. Simulated-data recovery tests can expose implementation problems when the generating truth is known; they cannot prove that real data follow that generating model.
Where does causal validity enter?
A likelihood and a sampler do not establish an intervention effect. A causal interpretation needs an identification argument supported by the design and assumptions, whether randomized assignment or an appropriate observational strategy. A model can predict well while giving the wrong answer to an intervention question.
For model-generated choice data, an assigned contrast is identified inside the configured model under the design assumptions. Generated choice is modeled stated choice. Human or live-market transfer needs evidence with a matched population, task and endpoint; no compiler pass supplies that evidence. Intervals from a fitted simulation remain conditional on its assumptions.
Before using generated code for a consequential recommendation, have a reviewer inspect the data-generating assumptions and the intended decision, with the code and diagnostics available. The methods and validation hub explains related choice-model checks; the evidence record should be read for its measured endpoint rather than as certification of an arbitrary generated model.