Across the social sciences, many studies use cross-sectional designs that reveal associations but are generally unable to support direct causal claims, yet authors of such articles may make or imply causal claims anyway. Here, to examine the prevalence of such ‘overreaching’ causal language, we analysed 194,631 cross-sectional articles using large language models. Over the period 1980–2024, an average of 46% of articles contained causal language in their titles or abstracts, where the annual rate has risen almost threefold since 2000 from 20% to 60%. To examine the effects of such language, we conducted a human-subjects experiment (N = 1, 105), finding that readers frequently indicate abstracts with this phrasing provide causal evidence but that methodological labels (β = −0.4, 95% confidence interval −0.56 to −0.19) and associational wording (β = −0.3, 95% confidence interval −0.43 to −0.07) reduce this tendency. Experiments with five LLMs revealed that model summaries of these articles (N = 100 each) can amplify causal overstatement, removing hedges and introducing causal claims where articles used strictly associational phrasing; however, prompting caution diminishes this pattern.
That is from a recent paper by Calvin Isch, Timothy Dörr, Neil Fasching, Grace Jennings & Duncan J. Watts. Note that Isch is on the job market this year, working with Tetlock and Watts.
Facts Only
* 194,631 cross-sectional articles were analyzed.
* An average of 46% of articles contained causal language in titles or abstracts over the period 1980–2024.
* The annual rate of causal language increased from 20% in 1980 to 60% in 2024.
* A human-subjects experiment found readers frequently indicated that abstracts with specific phrasing provided causal evidence.
* Methodological labels (e.g., $\beta = −0.4$) and associational wording (e.g., $\beta = −0.3$) reduced the tendency to believe causal claims.
* Experiments with five LLMs demonstrated that model summaries could amplify causal overstatement by removing hedges and introducing causal claims.
Executive Summary
Many studies in the social sciences frequently utilize cross-sectional designs, which are capable of revealing associations but generally cannot establish direct causal relationships. Despite this methodological limitation, authors often employ causal language in their titles or abstracts. An analysis of 194,631 cross-sectional articles published between 1980 and 2024 found that an average of 46% contained causal language. This rate has increased significantly over time, rising from 20% in 1980 to 60% in 2024.
Experiments involving human subjects indicated that readers often perceive abstracts using specific phrasing as providing causal evidence. However, methodological labels (like $\beta = −0.4$) and associational wording (like $\beta = −0.3$) were found to reduce this tendency among readers. Furthermore, when large language models summarized these articles, they showed an ability to amplify causal overstatements by removing hedges and introducing causal claims, though prompting caution reduced this amplification.
Full Take
Sentinel — Human
The text reads like a synthesis of empirical research methods applied to language use, exhibiting a formal yet context-aware style typical of academic commentary rather than pure LLM output.
