Inventory A/B Tests Show Predictable Interference Bias
A pairwise randomization over items and time separates carryover from capacity crowding, reducing bias mechanisms that affect switchback and item-level designs.
Underlying Paper
Experimental Designs for Multi-Item Multi-Period Inventory Control
Randomized experiments, or A/B testing, are the gold standard for evaluating interventions, yet they remain underutilized in inventory management. This study addresses this gap by analyzing A/B testing strategies in multi-item, multi-period inventory systems with lost sales and capacity constraints. We examine two canonical experimental designs--switchback experiments and item-level randomization--and show that both suffer from systematic bias due to interference: temporal carryover in switchbacks and cannibalization across items under capacity constraints. Under mild conditions, we characterize the direction of this bias in different scenarios. Motivated by two-sided randomization, we propose a pairwise design over items and time and analyze its bias properties. Controlled stochastic simulations verify the theoretical predictions, and trace-driven experiments on real-world fresh-retail data show that the same mechanisms persist in realistic environments with stockout substitution.
A/B tests are difficult in inventory systems because the unit being randomized is not isolated. Today's stocking decision changes tomorrow's leftover inventory, and one SKU's allocation can crowd out another SKU when products share a capacity constraint. This paper studies that interference directly for multi-item, multi-period lost-sales inventory control, where the intervention is a forecasting improvement and the outcome is profit.
The authors compare three designs: switchback experiments over time, item-level randomization, and a pairwise design that randomizes across both items and periods. The central result is not that one design is always best. It is that the direction of bias depends on what the forecasting intervention changes: the mean error or the dispersion of errors.
Core Contribution
The paper's contribution is a bias map for inventory experiments under capacity and carryover. In Scenario 1, the treatment reduces downward mean forecast bias: both treatment and control underforecast demand, but treatment underforecasts less. In Scenario 2, treatment keeps the forecast centered at the same mean but reduces error dispersion, formalized through convex order.
That distinction matters because the same experiment design can change sign across scenarios. Under mean-bias improvement, switchbacks are downward biased because treated periods inherit different leftover inventory from earlier periods. Item-level randomization can be upward biased under tight shared capacity because treated items crowd out control items. Pairwise randomization combines the two channels, so its bias is bounded above by item-level randomization but can still become negative when carryover dominates.
For dispersion reduction, the asymptotic picture changes. Under a mean-field regime with many items and proportional capacity, the paper proves that the global treatment effect is nonnegative, switchbacks have nonnegative asymptotic bias, item-level randomization is asymptotically unbiased, and pairwise randomization inherits the same asymptotic bias as switchbacks.
Technical Approach
The inventory model is a myopic base-stock system with lost sales, item-level margins, ordering costs, holding costs, and a shared capacity constraint. The stocking decision is determined by estimated demand parameters and a common KKT multiplier. Capacity coupling enters through that multiplier; temporal interference enters through leftover inventory carried into the next period.
For Scenario 2, the authors introduce a mean-field limit with items and capacity growing proportionally. Forecasts take the form , with treatment errors smaller in convex order than control errors. Assumptions impose bounded forecast errors, independence across items and assignments, bounded moments and parameters, and deterministic limits for the capacity multiplier. Lemma 2 shows that the random KKT multiplier converges to a deterministic , which removes the main finite-system capacity complication in the limit.
The resulting base-stock level converges to an affine response,
This is the technical hinge: once the shared capacity term becomes deterministic, the remaining treatment-control difference comes from the forecast-error distribution. That is what lets the authors use convex-order comparisons to prove the sign of the global effect and the asymptotic biases.
Results and Analysis
The experiments are best read as mechanism checks rather than as an operational claim that a particular design will dominate in every retailer. In controlled simulations, demand follows with . The reported setups use , , marginal treatment probability , 300 global-treatment/global-control replications, and 300 design replications.
For Scenario 1, control and treatment forecasts use and , so treatment is less downward biased. The paper reports results under capacity factors 0.90, 0.92, and 1.20. The simulated global treatment effect is positive in all regimes. Switchback estimates are negative biased, item-level estimates overstate the effect under tight capacity, and pairwise randomization remains below item-level randomization; when carryover dominates crowding, pairwise randomization can become negative biased.
For Scenario 2, control errors are drawn from and treatment errors from , with both centered at the true mean. The large-system simulations match the theory: the treatment improves expected profit, switchbacks are upward biased, and item-level randomization is approximately unbiased across tight, medium, and loose capacity regimes.
The trace-driven section uses FreshRetailNet-50K with a 90-day training window and a 7-day evaluation horizon. Demand is recovered with a TimesNet-based component, forecasts come from naive, DLinear, and Temporal Fusion Transformer models, and the simulator includes stockout substitution. The paper reports substitution ratios of about 19%–33% in Scenario 1 and 5%–20% in Scenario 2, which is enough to show that cross-item interference is not just a synthetic artifact. The evidence supports the paper's main claim about bias direction, but the empirical layer is still a simulator over recovered demand rather than a field experiment on live inventory decisions.
Limitations
The analysis depends on stylized lost-sales dynamics, myopic base-stock decisions, bounded parameters, independence conditions, and mean-field limits. The trace-driven evaluation adds realism through retail data and substitution, but prices, ordering costs, holding costs, and capacity factors are synthesized rather than observed as full operational primitives. The paper therefore gives useful design guidance for experimenters, especially teams testing demand forecasts under shared capacity, but it does not replace a deployment-specific interference audit.
Evidence Box
moderateKey Claims
- •Switchback experiments are biased by inventory carryover
- •Item-level randomization is biased by cross-item capacity crowding
- •Pairwise randomization separates temporal and cross-sectional interference channels
- •Bias direction depends on whether forecasts improve mean error or error dispersion
Key Results
- •Scenario 1 simulations use N=3000, T=60, p=0.5, and 300 design replications
- •Scenario 1 uses forecast shifts Δ(0)=-0.50 and Δ(1)=-0.05, with positive simulated GTE across capacity factors 0.90, 0.92, and 1.20
- •Scenario 2 uses ε(0)∼Unif[-30,30] and ε(1)∼Unif[-0.2,0.2], with item-level randomization approximately unbiased in 3 capacity regimes
- •Trace-driven FreshRetailNet-50K experiments use a 90-day training window, 7-day evaluation horizon, and substitution ratios of 19%–33% in Scenario 1 and 5%–20% in Scenario 2
Limitations & Caveats
- •No live randomized field experiment on operational inventory decisions
- •Trace-driven experiments rely on recovered latent demand rather than fully observed demand
- •FreshRetailNet-50K lacks full economic inputs, so prices and costs are synthesized
- •Theory requires bounded parameters, independence assumptions, and mean-field asymptotics