Financial advisors who assume their clients' AI chatbots can handle basic money questions may want to think again. A new benchmarking study from Saturn, a technology firm serving more than 750 advisory firms and 6,500 advisors, found that popular AI models gave incorrect or incomplete answers to financial queries 57% of the time.
The Artificial Authority report, released Tuesday, tested 18 AI models—including free and paid versions of ChatGPT, Gemini, Claude, and CoPilot—across 121 questions spanning debt management, student loans, mortgages, pensions, inheritance tax, and retirement planning. Each question was run five times, yielding more than 10,000 responses. The average accuracy rate across 17 current models was just 43%, according to the UK-based firm.
Performance deteriorated sharply on complex queries. On "hard" questions—multi-part scenarios involving interacting tax rules and precise figures—accuracy fell to 12%, meaning mistakes occurred 88% of the time. Even on the easiest questions, which required no calculation, accuracy averaged only 54%.
The gap between free and paid models was stark. Free models failed 63% of the time, compared with 49% for paid versions. The worst-performing free model, Claude Haiku 4.5, produced wrong or substantially incomplete answers 82% of the time. Even the best free model, ChatGPT-5.6 Luna (max), still returned incorrect or incomplete responses 56% of the time. The best overall performer was Claude Opus 5 (reasoning), a paid Anthropic model, which achieved a pass rate of 61%—meaning it still failed nearly four out of ten times. On hard questions, even this top model erred 67% of the time.
The report's authors flagged an equity concern: consumers least able to afford professional advice are the most likely to rely on free AI tools, yet those tools are the least accurate. This dynamic, they argue, risks compounding financial inequality rather than narrowing it.
Saturn's methodology required each response to meet a strict all-pass standard—every required figure, rule, warning, and deadline had to be present. The most common failure type (37.1% of errors) was an incomplete answer not captured by other categories. Omitting a required figure, limit, or deadline accounted for another 18.4% of failures. More alarming were fabricated rules (6.1% of failures) and out-of-date regulatory guidance (2.3%). The report noted that wrong answers were delivered in the same authoritative, fluent tone as correct ones, giving consumers no clear signal of unreliability.
The study comes amid growing reliance on AI as a substitute for professional advice. An EY report from April 2026 found that 49% of respondents had used AI in the prior six months to assist with saving or investing, 21% for product recommendations, and 18% for budgeting or trading support. For advisors, the Saturn data offers a frame of reference as clients increasingly arrive having already consulted AI tools. The findings echo other industry research, such as a T. Rowe Price study on DC advisors deploying AI, which suggests AI's role in advice is growing but not yet reliable.
Saturn's research also aligns with broader concerns about AI in financial services. A separate warning from Macabacus, a New York-based Microsoft 365 productivity platform, highlighted that most finance firms have sent clients an AI-generated error. The Saturn report concludes that the most freely available models are also the most error-prone, a dynamic that could have serious financial consequences for consumers who rely on them.
For wealth management professionals, the takeaway is clear: AI chatbots are not yet a substitute for professional judgment. As clients bring AI-generated answers to meetings, advisors should be prepared to verify and correct. The study's findings on hard questions—where even the best model failed two-thirds of the time—underscore the need for human oversight in complex financial planning.


