What a Rigorous Causal-Inference Pipeline Checks Before You Trust Its Output
A causal-modeling vendor's output is only as trustworthy as the engineering discipline behind the pipeline that produced it. Before a technical or data-science leader bases pricing, messaging, or go-to-market decisions on a vendor's causal effects, the question worth asking is not "does the model run," but "which stage of verification can this vendor actually show me."
The decision: trust the pipeline, or demand more proof first
Basing a budget or go-to-market decision on causal output from an opaque inference pipeline risks discovering, only after the money is spent, that gradients were never checked, uncertainty was never quantified, or the model's internal structure was never inspected. The fix: ask each vendor to demonstrate, not just describe, the stage of verification their pipeline actually passes.
A general pattern for turning a simulator into an inspectable model
The PyMC example gallery and a related PyMC Labs tutorial illustrate wrapping JAX functions for probabilistic inference. They provide computational examples, rather than validation of any vendor's causal model.
Step 1: Specify the computation and reference behavior
Write down the generative process, inputs, outputs and assumptions. Keep a reference implementation or known outputs for checks. Compilation and caching can improve performance, but an interpreted implementation can also compute the specified model correctly.
Step 2: Obtain gradients, and check them numerically
If the sampler needs gradient-based inference, the pipeline needs a way to compute how the output changes with respect to each input. JAX's automatic differentiation can supply this directly for functions already written in JAX, without hand-deriving gradient expressions. Before that gradient is trusted, it gets checked: a numerical gradient check compares the analytic gradient against a finite-difference approximation and confirms they agree.
Step 3: What does keeping the model graph inspectable show?
Once the wrapped function sits inside a model, that model can be rendered as a graph, showing exactly which prior feeds which deterministic node, which likelihood is observed, and how they connect. This is not decoration. It's the difference between a pipeline a reviewer can audit and one they have to take on faith.
Step 4: Support the execution backend actually used
A wrapped PyTensor operation can run through the standard backend. To compile the whole graph to JAX, provide the required JAX dispatch implementation. Registration is a requirement for that execution path, rather than a universal prerequisite for correct inference. Compare outputs and derivatives for the paths the pipeline uses.
Step 5: Confirm the output matches under both execution paths
The source compares log-probabilities and gradients across standard and JAX execution. Agreement checks that the implementations compute the same quantities for the tested inputs. It does not prove all numerical cases are correct or that the sampler converged. Check multiple chains, R-hat, effective sample size and divergences separately.
Do some samplers require a manual gradient?
The source example includes a case where a more complex simulator, such as a small neural network, is wrapped without implementing a manual gradient method at all. JAX-native samplers can differentiate through the entire compiled model graph automatically, so a manual gradient implementation becomes unnecessary for that specific class of sampler. That's a valid implementation choice, but it does not remove the verification requirement: the numerical gradient check in Step 2 still matters whenever a manual gradient is written, and the automatic differentiation path still has to be confirmed against the model's actual log-probability, the way Step 5 does.
Apply the checklist to a behavioral-experimentation pilot
Subconscious is a causal behavioral platform. It uses causal experimentation and discrete-choice-style modeling to test product, pricing, messaging, and go-to-market actions before teams commit capital. Where a study design supports it, results can include causal effects with credible intervals or other uncertainty quantification. That is the general family this checklist describes, a causal-modeling pipeline that needs its own verification stages, not a claim that Subconscious's internal tooling matches any specific library in the example above. Read more about the method.
What can this checklist not establish?
This is a discipline checklist, not a product claim. It describes what rigorous causal-inference engineering looks like in general; it does not confirm which specific technical stack any vendor, including Subconscious, runs internally, and it is not a substitute for a buyer's own technical due diligence. Credible-interval and uncertainty output should be described as available where the underlying study design supports it, not treated as a universal guarantee attached to every result a vendor produces.
Turn vendor claims into an evidence request
Request the reference outputs, applicable gradient checks, model graph and evidence for the execution paths actually used. Then ask about sampler diagnostics, predictive performance and causal identification. Randomization is one route to identification; observational designs need their own assumptions and checks. Discuss a study or review the methodology.