What is it about?

Large language models such as ChatGPT are increasingly used by researchers, raising an important question: do these tools actually increase scientific productivity? A recent study (Kusumegi et al., 2025) addressed this question by identifying the first time a researcher's paper appeared to contain LLM-assisted writing and treating that date as the moment the researcher adopted an LLM. We show that this approach can create a misleading pattern. Researchers who publish more papers in a given month are automatically more likely to have at least one paper identified as LLM-assisted. As a result, the estimated "adoption date" is itself related to productivity. Using arXiv publication data, placebo tests and simulations, we show that this mechanism can produce an apparent increase in productivity even when LLM use has no effect at all. Our findings highlight the importance of carefully separating how technology adoption is measured from the outcome researchers are trying to explain.

Featured Image

Why is it important?

Understanding whether AI tools increase scientific productivity has important implications for researchers, universities and science policy. Our results show that measuring LLM adoption from researchers' publication output can introduce a statistical bias that looks like a productivity effect. More broadly, the study illustrates a general problem for empirical research: when the timing of an intervention is determined using the outcome being studied, standard event-study methods can produce convincing but potentially spurious patterns. Careful research design and placebo tests are therefore essential when evaluating the effects of rapidly adopted technologies such as generative AI.

Perspectives

The rapid adoption of generative AI creates exciting opportunities for studying how new technologies affect knowledge production, but it also creates important measurement challenges. LLM use is difficult to observe directly, so researchers often have to infer adoption from indirect signals. Our study shows that when the estimated timing of LLM adoption depends on researchers' publication output, adoption and productivity can become mechanically linked, potentially generating misleading event-study patterns. Importantly, our results do not show that LLMs have no effect on scientific productivity. Rather, they show that this particular empirical design cannot, on its own, reliably identify such an effect. More broadly, our findings highlight the need for careful treatment definitions, identification strategies, and placebo tests when evaluating the consequences of generative AI and other rapidly adopted technologies.

Thomas Renault
COMUE Universite Paris-Saclay

Read the Original

This page is a summary of: Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics, Proceedings of the National Academy of Sciences, August 2026, Proceedings of the National Academy of Sciences,
DOI: 10.1073/pnas.2618638123.
You can read the full text:

Read

Contributors

The following have contributed to this page