Humans and LLMs Split Theory-Building Strengths
A blinded comparison of 25 LLMs and 73 scientists finds stronger individual AI performance but more prediction-efficient, aggregable human theories.
Underlying Paper
Artificial intelligences and human scientists exhibit complementary strengths in theory building
We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.
Theory building is not a single capability. A useful account must generate an explanation, make predictions about observations not yet seen, and change when those predictions fail. This study compares large language models with human social scientists on those linked tasks in academic discourse about gender and race inequality. Its central result is uneven rather than absolute: the models generally perform better as individual contributors, while people produce a more diverse set of theories whose aggregation yields larger improvements in predictive accuracy.
Core Contribution
The paper treats theory formulation, prediction, and revision as separable empirical activities rather than using fluency or rater preference as a proxy for scientific reasoning. It compares 25 LLMs with 13 senior researchers and 60 doctoral scholars, then evaluates both the theories they write and the predictions those theories produce for empirical patterns.
The distinction between elaboration and predictive efficiency is the important contribution. Independent raters, blinded to source, rated AI-generated theories more highly, and the models introduced more theoretical paths and latent variables. Yet that extra structure was not associated with more accurate predictions. The authors characterize part of the complexity as ornamental: it can make an account look richer without supplying information that improves out-of-sample empirical expectations.
Human contributors had a different advantage. Their theories were more diverse, and aggregating them produced greater gains in prediction accuracy. That makes the paper less a ranking of human versus machine intelligence than an argument for complementarity: models may be useful for producing and extending candidate explanations, whereas diverse human proposals remain valuable when a group needs a better collective forecast.
Technical Approach
The evaluation spans four task-by-effect conditions. Contributors first supplied explanations and predictions before seeing machine-learning-derived evidence, then revised their theories after that evidence was available. This design lets the study assess not only whether a theory sounds plausible, but whether its predictions correspond to observed empirical patterns and whether its author responds to error.
For revision, the authors quantify the change in a contributor's explanation as cosine distance between embeddings of the pre-ML and post-ML explanations. They relate that distance to pre-ML prediction error, defined as one minus pre-ML accuracy, and compare humans and GenAI using pooled observations across the four tasks. The reported analysis includes ordinary least-squares fits, Spearman correlations, bootstrap 95% confidence intervals, Welch's t-tests, and a mixed-effects model with contributor and task-by-effect structure.
That operationalization has a clear practical benefit: it distinguishes merely producing a different post-evidence answer from revising more when one's earlier predictions were wrong. It also supplies a concrete test for whether responsiveness to new evidence is calibrated rather than indiscriminate.
Results and Analysis
The paper reports that AIs outperformed most individual humans on most of the studied tasks. Their theories were more extensively elaborated and received higher blinded quality ratings, while they were also significantly more likely to revise after seeing new evidence. Figure 5 directly examines the latter claim: it compares mean pre-to-post explanation distance, plots revision against prior prediction accuracy for both groups, and fits a mixed-effects model of revision against prior error.
The more consequential result is the interaction between error and revision. Human scientists revised selectively in a way sensitive to their prior prediction errors; the paper contrasts this with broader AI revision. The visual analysis pools one observation per contributor across four tasks, so it supports a behavioral comparison within this experimental setting, not a claim about all scientific disciplines or real multi-year research programs.
The evidence supports the narrower conclusion that LLMs can be strong individual participants in structured theory-building exercises. It does not support replacing scientific communities with a single model. Higher ratings and greater elaboration are weak substitutes for predictive accuracy, and the human aggregation result shows why diversity matters: a collection of individually imperfect theories can be more useful than a set of similar, polished answers. For practice, the implication is to assign roles rather than declare a winner—use models to generate, extend, and reconsider candidate explanations, then preserve independent human judgment when combining predictions.
Limits of the Comparison
The domain is limited to discourse on gender and race inequality, and the tasks are structured studies rather than open-ended scientific programs. The reported evidence also evaluates theories through selected empirical effects and a particular embedding-based measure of textual revision. Those choices make the comparison tractable, but they leave unanswered whether the same division of strengths holds in fields with different data, causal methods, incentives, or forms of experimentation.
Evidence Box
strongKey Claims
- •LLMs outperform most individual human contributors on most theory-building tasks
- •Human theory diversity yields larger gains from aggregation
- •AI theories are more elaborated without commensurate predictive accuracy
- •Humans revise theories selectively according to prior prediction error
Key Results
- •25 LLMs were compared with 13 senior researchers and 60 doctoral scholars
- •Theory revision was evaluated across 4 task-by-effect conditions
- •Figure 5 uses 95% bootstrap confidence intervals for group revision comparisons
- •Prior error was defined as 1 minus pre-ML prediction accuracy
Limitations & Caveats
- •Evaluation is confined to academic discourse on gender and race inequality
- •The four structured task conditions do not reproduce long-horizon scientific research
- •Theory revision is measured through cosine distance between explanation embeddings
- •The abstract reports relative performance without providing task-level effect sizes