Ground Truth.
AI, checked against the source.

News · 2026-08-01

AI financial advice works when you hand it a full financial plan

Researchers collected financial questions from 1,000 demographically representative US adults, then simulated whether hypothetical households following the resulting chatbot advice would end up better off across a full lifetime of income, taxes, unemployment and retirement. The advice moves outcomes toward the authors' life-cycle-finance benchmark - but the version that works is not a well-worded question. It is a complete financial-planning harness with fixed assumptions and arithmetic constraints.

Key facts

The framing everywhere else is "AI gives surprisingly good financial advice if you ask the right questions." That is true and badly misleading, because of what "the right questions" turns out to mean.

Under ordinary prompts - the kind real people wrote - the advice is decent and conventional: save more during working years, hold diversified equities, reduce equity exposure with age, draw down in retirement. That is genuinely useful, because it is exactly the guidance many households never receive.

The failures are subtler and more revealing than the successes. The model cuts spending too sharply after a simulated job loss even when the household is sitting on liquid savings - the opposite of what the savings are for. It leans on round-number heuristics for saving and withdrawal rates. And it largely lets portfolios drift with market returns rather than rebalancing them, which is the single most mechanical piece of portfolio maintenance there is.

The authors' improved prompt is where the story turns. It is not a phrasing discovery. It supplies age, job status, income history, unemployment benefits, account balances, wealth and spending; fixes tax and Social Security assumptions; specifies return assumptions; instructs the model to act in the client's best interest; constrains the household to a single person with no dependents or bequest motive; and forces machine-readable numeric output subject to a budget identity and account-balance constraints.

That is a small financial-planning system with complete inputs and guardrails. The honest translation is that the model is a competent reasoning component inside a harness someone else built - not an adviser you can summon with better wording. And even with the harness, the rebalancing inertia does not go away.

Which sets up the finding with the most uncomfortable implications. Prompts written by respondents with lower financial literacy, by women, and by people without prior experience using AI for advice produced materially different simulated retirement outcomes. Much of the gap traces to the questions people asked - what information they volunteered, what they thought to ask about. But a separate test, in which identical prompts were labelled with different genders, found the model changing its recommendations. The authors are careful here: that could reflect legitimate inference about unobserved needs, or learned stereotypes, and the study does not resolve which. The defensible claim is that AI advice can amplify existing prompt-literacy and information inequalities - the people who most need good advice are the ones least equipped to extract it.

The analogy is a brilliant accountant who answers exactly what you ask and volunteers nothing. Bring a complete file and you get excellent work. Bring a vague question and you get a competent answer to the wrong question, delivered with the same confidence.

One source discrepancy is worth flagging for anyone citing this. The MIT Sloan write-up says prompts were sent to GPT-5.2, GPT-5.6 or Gemini 3 Flash; the linked full paper documents GPT-5.2 and Gemini 3 Flash throughout, with no GPT-5.6 evaluation found. Do not describe it as a documented three-model comparison.

A parallel result sets the ceiling on the institutional side. A paper asking whether large language models can execute parent orders tests something much narrower than trading: a client has already decided to buy or sell, and execution means slicing the order across time to limit adverse prices. Its method is a constrained scheduler around a standard time-weighted strategy - a planner reads recent minute-level price and volume history and lays out a coarse allocation, and an executor nudges each minute's quantity within bounded distance of the ordinary amount. The best variant improves execution by roughly one basis point, about one hundredth of one percent of traded value, in a one-month backtest on Shenzhen snapshot data with randomly generated orders. There are no live orders, no broker integration, no fees, no queue position, no hidden liquidity and no market-impact response. The authors call live trading future work, and they are right to.

The honest caveat covers both papers. Neither observes a real outcome. One simulates fictional households with perfect data; the other backtests synthetic orders in one market over one month. Community discussion on Hacker News landed on the sharpest objection: an expert-quality prompt may simply move the hard work onto the user, and the study never compares AI advice against a human adviser, a knowledgeable friend, or ordinary web research. Several commenters made the point benchmarks structurally cannot capture - that a human adviser's real contribution is behavioural, keeping a client on a sensible plan when markets fall.


Primary source, verified: read the paper →

Key questions

Did anyone actually get richer from AI financial advice?

No. The study simulates hypothetical households following model advice through simulated incomes, taxes, unemployment and retirement. It measures distance from the authors' life-cycle-finance benchmark, not real household outcomes, and it does not compare the advice to a human adviser.

What is the 'right question' that improves results?

Not a phrasing trick. The authors' academic prompt supplies age, employment, income history, benefits, balances and spending, fixes tax and return assumptions, constrains the household situation, and forces machine-readable numeric output subject to a budget identity.

Where does the advice consistently fail?

It cuts spending too hard after job loss even when the household has liquid savings, leans on round-number savings and withdrawal rules, and lets portfolios drift after market moves instead of rebalancing. The better prompt does not fix the rebalancing problem.
Cite this

APA

Ground Truth. (2026, August 1). AI financial advice works when you hand it a full financial plan. Ground Truth. https://groundtruth.day/news/ai-financial-advice-works-when-you-hand-it-a-full-financial-plan.html

BibTeX

@misc{groundtruth:ai-financial-advice-works-when-you-hand-it-a-full-financial-plan,
  title  = {AI financial advice works when you hand it a full financial plan},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/ai-financial-advice-works-when-you-hand-it-a-full-financial-plan.html}
}

Topics: finance · llm-evaluation · bias · consumer-ai · research

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.