A research team led by Taha Choukhmane at MIT Sloan, with co-authors from MIT and Stanford, did something unusually rigorous with the "can AI give financial advice?" question: they had 1,000 ordinary adults write their own prompts asking leading AI models for advice, then simulated how people aged 22 to 89 would fare over their whole lifetimes if they actually followed it. The paper won the Swiss Finance Institute Outstanding Paper Award. Two findings matter for anyone in the advice business.
Finding one: the advice was surprisingly good
"We were somewhat surprised by how good the advice was," Choukhmane said. The AI guidance pushed people toward higher savings rates, stock market participation, diversification, and age-appropriate risk reduction: for many users, following it would have built substantial savings buffers. The caveats are real but specific: the models handled economic shocks like unemployment poorly, let portfolios drift rather than rebalancing, and leaned on oversimplified savings rules.
Finding two: the quality depended on who was asking
This is the finding that should stop an advice practice mid-scroll. Outcomes were not equal. In the simulations, women and less financially literate users ended up with roughly $50,000 (4%) less wealth by age 60 than other users. People with no prior AI experience ended nearly $100,000 (6%) behind. And about two-thirds of the gender gap came not from the model discriminating, but from different prompting styles: the model gives different advice to "Where should I invest starting with $50?" than to a question that carries age, income, savings, risks and assumptions.
When the researchers replaced user prompts with structured "academic prompts" containing explicit financial detail, advice quality jumped for everyone. As Choukhmane put it: "regular people are not writing prompts the way a finance professor is."
The uncomfortable implication for consumer chatbots
Read plainly, the study says a consumer chatbot amplifies the very gaps professional advice exists to close. The people who most need good financial guidance (less financially literate, less tech-experienced) get the worst of it, precisely because getting good output requires knowing what a complete financial picture looks like. There was even an unprompted product tilt: specific providers appeared in responses (Vanguard in 6% of answers, iShares in 3.4%) although almost no prompts mentioned them.
What this means if you run an advice practice
We read this research as the strongest argument yet for the division of labour we build around: the context is the product. The difference between mediocre and excellent AI output in this study was not the model; it was whether the question carried the full, structured picture. That is exactly what a professional practice has and a consumer does not: the client file, the risk profile, the fee schedule, the history, the regulatory frame.
- A consumer types a one-line question at a chatbot and gets generic guidance shaped by their own blind spots.
- A practice runs agents that are handed the complete, structured context every time: which is why Tallify, our product for financial planners, works from the adviser's own meeting notes, client records and fee schedules, never from a blank prompt, and every output crosses the adviser's desk before it counts.
The same logic explains the global adoption pattern: the FPSB's survey of 6,000 planners shows the profession deploying AI for admin and analysis while keeping advice human. MIT's simulations show why that boundary is not caution for its own sake: unassisted AI advice is good on average and uneven exactly where it hurts, and the fix (structured, complete context, professional review) is the thing a practice can systematise and a consumer cannot.
Want AI working from your practice's full context instead of a blank prompt? That is what we build.
Book a free call