DML Quantifies Representation Error in Causal Estimates
Cross-fitted DML turns AI-learned controls into an inferential target, separating valid representation-dependent estimates from conditions for causal coverage.
Underlying Paper
Pragmatic DML with AI-Learned Representations
Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We study when this approach is valid and develop a practical framework for causal inference with learned representations. For a broad class of estimands, an imperfect representation distorts the target causal parameter by the product of two representation errors: one in the outcome regression and one in the balancing weight (or Riesz representer). This yields three constructive results. First, cross-fitted double machine learning (DML) provides valid Wald inference for the representation-dependent target. When representation errors are small, the same interval covers the causal parameter, and it can even attain the semiparametric efficiency bound. Second, fold-wise representation learning (or fine-tuning) is compatible with DML inference for the causal parameter. To this end, we develop convex- and star-aggregation pipelines for learning and combining representations. Third, when representation errors are substantial, we can provide interpretable sensitivity regions and root-$n$ inference for their endpoints. In a multi-modal demand application, seven representation-specific estimates and their star aggregate all imply a negative near-unit elasticity for rank-based price response, and the result remains robust over the reported sensitivity grid.
Text, images, and other high-dimensional covariates increasingly enter causal studies only after an AI model compresses them into representations. That workflow is convenient, but it changes the inferential question: conditioning on an imperfect representation need not identify the causal parameter defined with the original covariates. The paper studies that gap rather than assuming it away. Its central result is that, for a broad class of estimands, the representation-induced distortion is governed by the product of an outcome-regression error and a balancing-weight, or Riesz-representer, error.
Core Contribution
The authors distinguish two targets that are often conflated. A representation-dependent target is the parameter identified after replacing rich covariates with a learned representation. The causal target is the parameter one would seek under the relevant original information set. Cross-fitted double machine learning provides Wald inference for the former under its usual orthogonality logic; coverage for the latter additionally depends on the two representation errors being sufficiently small.
That framing is useful because it does not present learned embeddings as automatically adequate controls. A representation can predict outcomes well while failing to preserve the variation required for balancing, or it can support balancing while losing outcome-relevant information. The product characterization makes the failure mode more specific than a generic warning about omitted variables: either component can make the causal discrepancy material, while a small error in one component can attenuate the effect of a larger error in the other.
Technical Approach
The paper places the representation inside a cross-fitted DML pipeline. Nuisance functions are learned on training folds and evaluated on held-out folds, limiting the direct reuse of observations for representation construction and estimating-equation evaluation. The authors argue that this yields valid Wald inference for the representation-dependent target. When the relevant representation errors vanish quickly enough, the same interval covers the causal parameter; under stronger conditions it can attain the semiparametric efficiency bound.
A second contribution concerns representations that are learned or fine-tuned during estimation rather than fixed in advance. The proposed fold-wise learning setup keeps that adaptation compatible with DML inference. The paper also develops convex and star aggregation pipelines for combining candidate representations. Rather than selecting one embedding as definitively correct, aggregation allows the estimating procedure to form a representation from several candidates while retaining an inferential analysis tied to the resulting target.
The third branch handles the case that matters most in applied work: representation quality may not be good enough to dismiss bias. The authors construct interpretable sensitivity regions indexed by the two representation errors and report root- inference for the endpoints. This replaces an unsupported point-identification claim with a range whose movement can be inspected as assumptions about representation adequacy change.
Results and Analysis
The empirical application is a multimodal demand setting using rank-based price response. The paper reports seven representation-specific estimates and one star aggregate. All eight estimates imply a negative elasticity close to minus one, and the conclusion persists across the reported sensitivity grid. The consistency across representations is meaningful evidence that the substantive finding is not driven by one particular learned feature map.
Still, the application supports a narrower conclusion than universal validation of learned representations for causal adjustment. Agreement among seven candidates and their aggregate is reassuring within this demand problem, but it does not establish that text, image, or multimodal embeddings will preserve the needed outcome and balancing information in another domain. The sensitivity analysis is the more transferable part of the contribution: it exposes how a causal interpretation depends on representation errors instead of treating the embedding as an innocuous preprocessing step.
For practitioners using foundation-model features in econometric pipelines, the practical message is disciplined. Report the representation-dependent DML estimate as such, use fold-wise learning when the representation is adapted to the data, compare credible representations or aggregates, and show sensitivity regions when the bridge from representation to causal control cannot be justified. The paper supplies an inference framework for those choices; its empirical evidence is concentrated in one multimodal demand application.
Evidence Box
moderateKey Claims
- •Representation bias equals the product of outcome and balancing-representer errors
- •Cross-fitted DML gives Wald inference for the representation-dependent target
- •Fold-wise learning and representation aggregation can support causal inference
- •Sensitivity regions provide root-n inference when representation errors remain material
Key Results
- •7 representation-specific estimates imply negative near-unit rank-based price elasticity
- •1 star aggregate also implies negative near-unit rank-based price elasticity
- •8 reported estimates remain consistent over the paper's sensitivity grid
- •Root-n inference is developed for sensitivity-region endpoints
Limitations & Caveats
- •Empirical evidence is concentrated in one multimodal demand application
- •Causal coverage requires representation errors to be sufficiently small
- •Sensitivity conclusions depend on the specified outcome and balancing-error grid
- •Agreement across 7 representations does not validate learned controls in other domains