Large Language Models (LLMs) such as ChatGPT are transforming how scientists conduct and validate research, offering promise as tools to improve scientific reproducibility. However, computational reproducibility and error detection remain expensive and labor-intensive. We experimentally test how collaboration between researchers and LLM assistants influences the reproduction of quantitative social science findings across different levels of AI autonomy. We randomly assigned 288 researchers to 103 teams working under three conditions: human-only, AI-assisted (using ChatGPT as a collaborative tool), or AI-led (ChatGPT operating with minimal human oversight). Teams reproduced published results from leading social science journals, detected coding errors, and proposed robustness checks.
PNAS
AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science
Journal Article
Reference
Hammar, Olle and Joakim Jansson (2026). “AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science”. PNAS 123(22), e2524747123. doi.org/10.1073/pnas.2524747123
Hammar, Olle and Joakim Jansson (2026). “AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science”. PNAS 123(22), e2524747123. doi.org/10.1073/pnas.2524747123
Authors
Olle Hammar, Joakim Jansson