Scientists Report Seven Hours Saved Each Week With AI

Triangulating 15 million Gemini interactions, 2,600 specialized models, and a 600-plus scientist survey links AI use to nearly seven weekly hours saved.

Editorial Desk·September 26, 2026·4 min readmoderate

Underlying Paper

AI in Science: Early Insights

Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements-- LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.

arXiv:2609.28504Submitted: Sep 25, 2026v1

AI’s effect on research is usually discussed through demonstrations, funding announcements, or isolated laboratory studies. This paper instead assembles three observational sources to ask where AI is already entering scientific work: 15 million Gemini interactions, an inventory of more than 2,600 specialized AI models, and a survey of more than 600 scientists. Its central finding is not that a single model replaces scientific judgment, but that general-purpose language models and specialized systems are being used at different points in the research process—and that the resulting time savings may expose new bottlenecks.

Core Contribution

The authors introduce a scientific-task taxonomy used to map both human interactions with Gemini and the capabilities of specialized AI systems. That common frame lets them compare tools that are usually counted separately: a scientist asking a general model for help with analysis or code, and a field-specific model making a prediction, generating data, or classifying an input.

The comparison supports a complementarity claim. Gemini interactions cluster around broadly reusable work such as analysis, coding, and manuscript preparation, while specialized models are concentrated in domain-specific predictive, generative, and classification tasks. That distinction matters because adoption statistics alone could otherwise imply that a general LLM is displacing scientific software. The paper’s evidence instead suggests a division of labor across the workflow.

Technical Approach

The study combines behavioral traces, a model inventory, and self-reports rather than treating any one source as a complete measure of scientific AI use. Gemini activity is used as a proxy for LLM use, including API interactions; the specialized-model inventory is mapped to scientific domains and tasks; and the survey captures reported frequency, time savings, use cases, and downstream constraints. Figures on domain representation, task representation, and geography make clear that the analysis is designed as a broad descriptive map rather than a controlled productivity experiment.

Figure 8 compares the distribution of scientific tasks in LLM interactions with specialized AI models. Its purpose is analytical rather than promotional: the contrast shows why aggregate counts of “AI for science” conceal different kinds of work. General models appear suited to tasks that transfer across fields, whereas specialized systems encode narrower scientific functions.

Figure 8. Distribution of Science Tasks in LLM Interactions vs. Specialized AI Models

The paper also uses the survey to move beyond tool availability. Nearly half of respondents report using some form of AI every day, and respondents report that saved time is primarily reinvested in research. The authors connect that pattern to task interdependencies: accelerating one stage does not automatically accelerate discovery if validation, experimentation, or review becomes the limiting stage.

Results and Analysis

The clearest quantitative result is the survey estimate of nearly 7 hours saved per week from AI use. If the respondents’ reports reflect actual practice, that is a material shift in research capacity, particularly because the reported destination of saved time is more research rather than less work. But it is not a causal estimate of output: self-reported time savings do not establish whether papers, experiments, or discoveries improve by a corresponding amount.

The scale of the descriptive evidence is a strength. The Gemini sample contains 15 million interactions, the specialized-model inventory contains over 2,600 systems, and the survey includes over 600 scientists. Those sources provide more coverage than a single-field case study and reveal broad disciplinary representation. They do not, however, identify the full population of scientists or all AI use: Gemini is only one LLM ecosystem, specialized-model coverage depends on the inventory, and survey participants may differ from non-users.

The paper’s most useful interpretation concerns bottlenecks. Scientists report a growing backlog of untested hypotheses and demand for verification as AI makes some upstream tasks easier. That is a credible warning against reading time savings as an automatic increase in scientific output. The practical implication is that laboratories, funders, and tool builders may need to invest in experimental validation and output checking alongside model deployment. The evidence supports an early map of adoption and perceived productivity, not a settled estimate of AI’s long-run effect on science.

Evidence Box

moderate

Key Claims

  • •General LLMs and specialized AI models serve complementary scientific tasks
  • •AI use is broad across scientific disciplines and research activities
  • •AI saves researchers time that is largely reinvested in research
  • •Faster upstream work shifts constraints toward hypothesis testing and verification

Key Results

  • •15 million Gemini interactions mapped to scientific domains and tasks
  • •More than 2,600 specialized AI models inventoried across disciplines
  • •More than 600 scientists surveyed about AI use and research workflows
  • •Nearly 7 hours of weekly time savings reported by surveyed scientists

Limitations & Caveats

  • •Gemini interactions are a proxy for LLM use rather than a census of scientific AI activity
  • •Weekly time savings are self-reported rather than experimentally measured productivity
  • •The survey has more than 600 respondents but may not represent all scientific fields or institutions
  • •No causal estimate links AI use to publications, validated findings, or long-run scientific output

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.