Outlier Choices Rarely Move Means but Flip Marginal Decisions
A pre-registered reanalysis of 358 behavioral-science meta-analyses shows median mean shifts of at most 0.047 Cohen’s d but decision reversals up to 15.9%.
Underlying Paper
Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses
Meta-analysts routinely face estimates that look too large or extreme. Yet, how to handle them is left to the reviewer's judgment. The methods for detecting such estimates are well known. What is missing is an informed assessment of how much alternative handling choices might change a meta-analysis' conclusions. We fill this gap by analyzing the effects of four pre-registered handling treatments across 358 behavioral science meta-analyses with at least ten estimates. Each outlier handling treatment is estimated by two estimators (random effects and unrestricted weighted least squares), and compared to the 'do-nothing' baseline on three outcomes: the pooled effect, statistical significance, and whether the effect reaches the smallest effect size of interest (|d| >= 0.20). Our entire analysis and comparison pipelines were pre-registered. Alternative outlier handling treatments have little effect on the meta-analysis mean as the median absolute change in Cohen's d is at most 0.047 and often much less. Yet, at least one of these four treatments in combination with one of these estimators reverses the statistical significance of 11.5% of meta-analyses and the smallest-effect-of-interest assessment in 15.9%. Winsorizing has the least effect and DFBETAS the most. Categorical changes are found almost entirely among results already close to the decision boundary; strongly significant results essentially never change. These findings give applied meta-analysts, methods specialists, and reviewers a reference point for how much this under-reported choice matters and provide yet another reason for meta-analysts to publicly pre-specify their methods and handling treatments.
Meta-analysts often see unusually large or influential estimates, but the choice of what to do with them is still treated as a discretionary analysis detail. That matters because meta-analysis results are frequently converted into categorical claims: statistically significant or not, large enough to matter or not. This paper asks a narrow empirical question: across a large set of behavioral-science meta-analyses, how much do standard outlier and influence-handling decisions change the conclusion compared with doing nothing?
Core Contribution
That design separates two different notions of “mattering.” On the continuous scale, the paper finds that outlier handling usually moves the pooled mean only slightly. On the decision scale, the same movement can still matter when a result is already near a threshold. The paper’s main value is this distinction: small median shifts can coexist with nontrivial rates of conclusion reversal.
Technical Approach
The authors treat the original meta-analysis result as the baseline and then rerun the analysis after alternative handling decisions. The abstract identifies winsorizing as the least disruptive treatment and DFBETAS as the most disruptive one, showing that the four treatments differ in how often they reverse decisions. The comparison is repeated across both estimators, which is useful because the outlier decision and the meta-analytic estimator are not independent choices in practice.
The study also emphasizes design discipline. The analysis and comparison pipelines were pre-registered before the results were evaluated, and the treatment comparisons were defined in advance. That does not make the design immune to all modeling choices, but it reduces the concern that the headline rates were selected after inspecting the results.
Results and Analysis
The central quantitative result is conservative on the pooled mean and less conservative on decisions. Across the four treatments, the median absolute change in Cohen’s d is at most 0.047, and often smaller. For many applied conclusions, that is not a large movement. It suggests that outlier handling is rarely a way to transform an average effect into a very different average effect.
The categorical results are the warning. Across 358 meta-analyses, at least one treatment-estimator combination reverses statistical significance in 11.5% of cases. The same setup reverses whether the pooled effect reaches the smallest effect size of interest in 15.9% of cases. The practical reading is that outlier handling is usually not decisive, but it is decisive often enough that hiding the choice is not defensible.
The distribution of reversals matters. The authors report that categorical changes occur almost entirely among results already close to the decision boundary, while strongly significant results essentially never change. That is exactly where sensitivity analysis has the highest editorial value: not to relitigate stable findings, but to flag results whose status depends on a reasonable handling rule. Winsorizing produces the fewest reversals, while DFBETAS produces the most, so a single generic statement that “outliers were checked” would conceal a meaningful methods choice.
Limitations
The study does not identify which treatment is correct for any given meta-analysis. It measures sensitivity to pre-specified treatments, not ground-truth bias removal. The source domain is also behavioral science, with included meta-analyses constrained to those with at least ten estimates, so the exact rates should not be transferred mechanically to fields with different study designs or effect-size distributions. The useful takeaway is procedural: pre-specify outlier and influence handling, report the central finding with and without the treatment, and treat any threshold crossing as part of the result rather than an implementation detail.
Evidence Box
strongKey Claims
- •Outlier handling usually has little effect on pooled Cohen’s d
- •Decision reversals concentrate near statistical or practical thresholds
- •Winsorizing changes conclusions less often than DFBETAS
- •Pre-specification and transparent reporting reduce discretionary sensitivity
Key Results
- •358 behavioral-science meta-analyses with at least 10 estimates each
- •Median absolute change in Cohen’s d at most 0.047 across handling treatments
- •Statistical significance reversed in 11.5% of meta-analyses for at least one treatment-estimator combination
- •Smallest-effect-of-interest status reversed in 15.9% of meta-analyses for at least one treatment-estimator combination
Limitations & Caveats
- •Does not determine which outlier or influence treatment is correct
- •Evaluation limited to behavioral-science meta-analyses with at least 10 estimates
- •Categorical reversal rates depend on chosen thresholds for significance and |d| ≥ 0.20
- •Treatment effects are summarized across meta-analyses rather than validated against known true effects