Full text loading...
Abstract
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms—statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.
Facts Only
* Large language models are entering psychological research as tools and objects of inquiry.
* Many studies apply human instruments to LLMs without establishing output reliability or interpretability.
* Robust AI psychological research requires integrating psychometric validation (what a score means) and causal inference standards (what results warrant).
* A dual-validity framework is developed where evidentiary demands scale with scientific ambition across tool use, behavioral characterization, human simulation, and cognitive modeling.
* Text classification may require only accuracy and reliability.
* Claiming an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence such as construct validity and experimental controls.
* Progress depends on developing computational analogs of psychological constructs.
Executive Summary
Full Take
Sentinel — Human
This text presents a sophisticated argument about the methodological pitfalls of applying human psychology to LLMs, proposing a specific dual-validity framework for future research.
