Search and Filter

Repository Feedback

Your feedback helps us improve the repository's content relevance and usability. Please share your thoughts to help us better serve researchers and practitioners.

Submit feedback

Submit a research study

Contribute to the repository:

Add a paper

Analyzing Students' Statistics Writing Before And After The Emergence Of Large Language Models

Authors
Sara Colando,
Erin Franke,
Gordon Weinberg,
Alex Reinhart
Date
Publisher
arXiv
The ability to communicate statistical results to domain experts and stakeholders is an important goal of the undergraduate statistics and data science curriculum. However, as large language models (LLMs) have become more accessible, a major concern is that students are offloading important cognitive tasks to generative AI. Using a corpus of over 1,600 undergraduate students' data analysis reports from 2021 to 2025, we show how students' writing style and verb usage have become more similar to that of LLMs. This shift is most pronounced in the first and fifth quintiles of students' reports, which roughly map onto the introduction and conclusion sections, respectively. At the same time, we demonstrate that students' writing style has become more similar to that of statistics experts with the addition of LLMs. We end by discussing the implications of our findings for statistics and data science educators. In particular, we propose alternative modes of assessment that still emphasize statistical thinking, such as targeted writing assignments for structuring a report introduction.
What is the application?
Who is the user?
Who age?