Large language models can generate investment recommendations instantly. They can process client profiles, summarize market conditions, and produce polished rationales that sound personalized. But there is an important question underneath the fluency. Are these systems actually integrating the full client profile when they make decisions? This paper introduces a troubling possibility. LLMs may only appear to personalize advice while actually relying on a small number of dominant signals, especially self-reported risk tolerance. The result is a subtle but important failure mode. Recommendations can look sophisticated and individualized while still ignoring much of the information that suitability standards require advisors to consider.
One Size Fits None: Heuristic Collapse in LLM Investment Advice
- Jillian Ross and Andrew W. Lo
- Airxiv, 2026
- A version of this paper can be found here
- Want to read our summaries of academic finance papers? Check out our Academic Research Insight category
Key Academic Insights
LLMs often reduce complex advice to simple heuristics
The paper introduces the concept of heuristic collapse. This occurs when an LLM appears to use a full client profile but in practice bases its recommendations primarily on one or two salient variables. In investment advice, the dominant feature is usually self-reported risk tolerance. Other critical inputs such as age, income, liquidity needs, debt, and investment horizon receive much less weight.
The problem is difficult to detect from individual outputs
A key insight is that heuristic collapse is not obvious when reading a single recommendation. Individual responses can still sound coherent, personalized, and financially reasonable. The failure only becomes visible when analyzing recommendations systematically across large populations of clients and measuring how sensitive allocations are to different inputs.
The study uses synthetic investor profiles to test personalization
The authors generate 1,000 synthetic client profiles using Latin hypercube sampling to ensure broad and low-correlation variation across characteristics like age, income, debt, savings, investment timeline, and risk tolerance. They then ask frontier GPT models to construct portfolios from a standardized set of investment products.
Risk tolerance dominates allocation decisions
Across most models, portfolio recommendations are overwhelmingly driven by self-reported risk tolerance. Aggressive clients receive equities. Conservative clients receive fixed income. In many cases, risk tolerance accounts for 57% to 88% of the predictive weight behind recommendations, while age, liquidity, and income play much smaller roles.
Better diversification does not guarantee better personalization
Portfolios generated with web search are often more diversified, but not necessarily more tailored to individual clients. This distinction matters because fiduciary suitability standards require both diversification and personalization. Advice that looks professionally diversified may still fail to reflect the client’s unique circumstances.
GPT-4o and GPT-5 models behave differently
The paper finds important behavioral differences across models. GPT-4o produces more client-specific recommendations in the baseline setting but deteriorates when forced to use web search. By contrast, the GPT-5.4 family generally benefits from web search, producing more grounded and detailed rationales. However, even these models still exhibit significant heuristic collapse.
LLMs exhibit strong round-number heuristics
Another interesting finding is that LLMs strongly favor round-number allocations such as 10%, 20%, or 25%, even when no such constraint exists. In some models, nearly all allocation percentages are multiples of five. Web search reduces this tendency somewhat, but the bias remains persistent.
Practical Applications for Investment Advisors
Do not confuse fluency with personalization
An LLM may generate polished and convincing recommendations while still failing to integrate important client characteristics. Advisors should evaluate whether recommendations truly reflect the client’s broader financial situation.
Risk tolerance questionnaires are not enough
Suitability requires balancing both willingness and capacity to take risk. Clients often misreport risk tolerance, especially during emotionally charged market environments. Human oversight remains essential.
Use LLMs as assistants, not fiduciaries
Current models may be useful for summarization, scenario generation, or drafting explanations. But the evidence suggests they should not yet be treated as fully autonomous financial advisors.
Be cautious with tool-augmented AI systems
Adding web search or external tools does not automatically solve reasoning problems. In some cases, it may improve diversification while reducing client specificity.
How to Explain This to Clients
“AI systems can generate investment recommendations that sound highly personalized. But this paper shows that many models still rely heavily on simple shortcuts, especially a client’s stated risk tolerance.That means the advice may look customized while ignoring other important factors like age, income, debt levels, liquidity needs, or time horizon.The key lesson is that good financial advice requires holistic judgment. Technology can help support the process, but human oversight still matters when evaluating what is truly suitable for an investor.”
The Most Important Chart from the Paper
Figure 1 shows that LLMs exhibit substantial heuristic collapse across most asset classes, though the degree varies by model and asset class. For GPT-4o, equities (FC = 0.780, R2 = 0.882) and tax-advantaged accounts (FC = 0.704, R2 =0.843) show the highest heuristic use. This indicates that a small number of client features — most prominently self-reported risk tolerance — dominate allocation decisions.
The results are hypothetical results and are NOT an indicator of future results and do NOT represent returns that any investor actually attained. Indexes are unmanaged and do not reflect management or trading fees, and one cannot invest directly in an index.
Abstract
Large language models are increasingly deployed as advisors in high-stakes domains — answering medical questions, interpreting legal documents, recommending financial products — where good advice requires integrating a user’s full context rather than responding to salient surface features. We investigate whether frontier LLMs actually do this, or whether they instead exhibit heuristic collapse: a systematic reduction of complex, multi-factor decisions to a small number of dominant inputs. We study the phenomenon in investment advice, where legal standards explicitly require individualized reasoning over a client’s full circumstances. Applying interpretable surrogate models to LLM outputs, we find systematic heuristic collapse: investment allocation decisions are largely determined by self-reported risk tolerance, while other relevant factors contribute minimally. We further find that web search partially attenuates heuristic collapse but does not resolve it. These findings suggest that heuristic collapse is not resolved by web search augmentation or model scale alone, and that deploying LLMs as advisors requires auditing input sensitivity, not just output quality.
About the Author: Elisabetta Basilico, PhD, CFA
—
Important Disclosures
For informational and educational purposes only and should not be construed as specific investment, accounting, legal, or tax advice. Certain information is deemed to be reliable, but its accuracy and completeness cannot be guaranteed. Third party information may become outdated or otherwise superseded without notice. Neither the Securities and Exchange Commission (SEC) nor any other federal or state agency has approved, determined the accuracy, or confirmed the adequacy of this article.
The views and opinions expressed herein are those of the author and do not necessarily reflect the views of Alpha Architect, its affiliates or its employees. Our full disclosures are available here. Definitions of common statistics used in our analysis are available here (towards the bottom).
Join thousands of other readers and subscribe to our blog.
Facts Only
* Jillian Ross and Andrew W. Lo authored the paper "One Size Fits None: Heuristic Collapse in LLM Investment Advice."
* The research was published via Airxiv in 2026.
* The study used 1,000 synthetic client profiles generated via Latin hypercube sampling.
* Variables tracked included age, income, debt, savings, investment timeline, and risk tolerance.
* GPT-4o and GPT-5.4 models were tested for portfolio construction.
* Findings show risk tolerance accounts for 57% to 88% of the predictive weight in recommendations.
* Analysis indicates LLMs favor round-number allocations in multiples of five.
* The study tested the impact of web search augmentation on model outputs.
* GPT-4o baseline recommendations deteriorated with web search, while GPT-5.4 rationales improved.
* Heuristic collapse was specifically identified in equities (FC = 0.780, R2 = 0.882) and tax-advantaged accounts (FC = 0.704, R2 = 0.843) for GPT-4o.
Executive Summary
Frontier large language models exhibit "heuristic collapse" when providing investment advice, meaning they rely on a few dominant signals—primarily self-reported risk tolerance—while ignoring broader client context such as age, liquidity needs, and debt. This creates a failure mode where recommendations appear polished and personalized but lack the holistic integration required by fiduciary suitability standards.
The phenomenon is difficult to detect in individual outputs because the generated rationales remain coherent and financially reasonable. Systematic analysis reveals that while web search augmentation can increase portfolio diversification and improve the grounding of rationales in newer models like GPT-5.4, it does not resolve the underlying reliance on simple heuristics. Consequently, these systems function more effectively as assistants for summarization and drafting than as autonomous fiduciaries, as they struggle to balance a client's willingness to take risk with their actual financial capacity.
Full Take
This research employs a robust methodology, using Latin hypercube sampling to prevent correlation bias among synthetic profiles and applying interpretable surrogate models to quantify the weight of specific inputs. A peer reviewer would likely note that while the sample size (1,000) is sufficient for statistical trends, the use of synthetic profiles rather than real-world messy data may overlook how LLMs handle contradictory or ambiguous human inputs. Furthermore, the "GPT-5.4" reference suggests a forward-looking or specific internal versioning that would require precise documentation for replication.
The data confirms a significant gap between "fluency" (the ability to sound like an expert) and "reasoning" (the ability to integrate multi-factor constraints). The authors' claim that heuristic collapse persists despite model scaling or tool augmentation is well-supported by the predictive weights (57%-88%) attributed to a single variable. This challenges the assumption that more data or "grounding" via web search automatically leads to better personalization.
The real-world implication is a potential "suitability trap": advisors might trust an LLM because the output looks professional, while the system is actually ignoring the client's debt or time horizon. This shifts the risk from the model's output quality to its input sensitivity.
Bridge Questions:
1. Would the degree of heuristic collapse change if the prompt explicitly mandated a weighted scoring system for each client variable?
2. How do these findings change when LLMs are integrated with structured financial planning software versus acting as standalone chat interfaces?
Counterstrike Scan: A bad actor would use this narrative to argue that AI is fundamentally incapable of professional reasoning to protect human-advisor monopolies. However, this work is structured as a technical audit of specific failure modes rather than a broad condemnation of the technology.
