All projects
AI Product Design Wealth Management Enterprise SaaS Trust & Explainability

The future of AI & wealth advisory.

Led product design for an AI tool that drafts investment advice for financial advisors. I added the things advisors needed before they would trust it: where every number came from, how sure the AI was, and the ability to edit the wording before it reached a client.

Case Study Goldman Sachs AI SaaS · Jun 2025 — Present
Role Product Designer — AI SaaS Platform
Timeline Jun 2025 — Present
Platform Web SaaS · Advisor + Client portals
Domain AI · Wealth Management · Compliance

Before we start — what this is

30 seconds · no finance background needed

01 Wealth management

The part of a bank that looks after money for rich clients.

02 The advisor

A real person at Goldman Sachs assigned to those clients. They decide what a client should do with their money, then explain it to them.

03 The tool I designed

An AI that writes the first draft of that advice. It reads the client's holdings and drafts a recommendation in seconds instead of hours.

04 The catch

The advisor's name goes on the advice, not the AI's. Every piece of it is kept on file for seven years, and the advisor answers for it if it is wrong.

the client’s money holdings, goals the AI drafts the advice the client acts on it the advisor checks it and signs their name this step was broken

Where the advice comes from, and who answers for it

The AI showed its answer and nothing else. No sources, no reasoning, no way to check it. Advisors were being asked to stake their name on something they could not inspect — so half of them never used it.

At a glance Full read · 15 min  |  This panel · 40 sec

The problem

An AI wrote investment advice for financial advisors but showed no working. Advisors had to put their name to advice they could not check, so half of them never started using it.

What I did

Attached three things to every recommendation: how sure the AI is, a link to the data behind each claim, and text the advisor can edit before a client sees it.

The result

68%Advisors using it 67%Advice accepted +16ptTake-up −60%Support tickets

The Stakes

Only half of new advisors ever started using the tool. The AI gave sound advice, but it showed no working, so advisors could not check it before putting their own name to it.

Outcomes — two product releases

Advisors using it 52% → 68% · AI advice accepted 38% → 67% · Support tickets down ~60%

+16pt new-advisor activation lift over two product releases (52% → 68%)
+29pt AI recommendation acceptance lift after citation + confidence patterns shipped
~60% estimated reduction in "why did the AI suggest this" support tickets post-release
3 trust patterns shipped — confidence visibility, citations, editable outputs
01
Overview

An AI tool that advisors did not trust enough to use

Goldman Sachs builds AI tools for the people who manage money for wealthy clients. The AI drafts investment recommendations, sums up a client's portfolio, and writes client-ready notes in seconds instead of hours. The technology did its job. Almost nobody used it.

I joined the team that owns this tool. My work covered how advisors get started with it, how the AI's advice is presented, and what sits between what the AI writes and the advisor who has to sign their name to it. In this industry a wrong recommendation is not only a bad experience. Every piece of advice is filed and kept for seven years, and the advisor answers for it.

Headquarters

New York, USA

Founded

1869

Industry

Investment Banking · Wealth Mgmt

AUM

~$2.8T (public, 2024)

Platform Users

~3,000 advisors (internal SaaS)

Compliance Surface

SEC · FINRA · internal IB controls

02
Problem

Advisors stopped using the AI because they could not check its work

When I joined, only 52% of new advisors ever got going with the tool. The AI presented its advice as finished, confident text. It gave no reasoning, no source for any figure, and no way to change the wording before it went to a client. Advisors were being asked to put their name to something they could not inspect.

Problem statement First-time platform users (wealth advisors, 5–15 years tenure) failed to complete activation because AI-generated recommendations lacked source provenance, confidence signalling, and edit affordances — resulting in 48% of new advisors abandoning within their first two weeks, documented via internal product analytics across the first two release cohorts.

"I'm not going to send a client a recommendation I can't defend to compliance. The AI gives me a clean answer, but I can't see where it came from. So I just... don't use it."

— Senior wealth advisor, week-2 platform exit interview
03
Research

Three kinds of evidence pointed at the same cause

I treated the low take-up as a research question first. The easy reading was that the screens confused people, which points at a layout fix. The evidence pointed somewhere else. Advisors did not doubt the answer. They doubted anything they could not trace back to a source.

01
Activation Funnel Analysis
Analytics showed 48% of new advisors dropped out between seeing their first AI recommendation and accepting one. The drop sat inside a single 90-second window right after the advice appeared. People were not getting lost. They were hesitating.
02
Advisor Shadowing (n=12)
I sat with 12 advisors during real client reviews. Eight said, without being asked, that they could not see how the AI had got there. Four kept their own spreadsheets to work the recommendation out again by hand before using it. Nobody rebuilds an answer they already trust.
03
Industry Benchmark
An industry survey found 73% of financial advisors turn down AI recommendations that show no sources, even when the recommendation is correct. This was not a Goldman problem. It was a pattern across the category that the product had never been designed for.

What the numbers could not tell me

Analytics showed where advisors dropped out, but not why any single recommendation was turned down. So I worked with the AI team to add a one-tap “this was not useful” button with an optional reason. Within three weeks we had reasons for more than 1,200 rejections. The top three were no visible source (38%), cannot edit before sending (24%), and does not say how sure it is (19%). Four out of five rejections were about trust, not about the quality of the advice.

04
Options Considered

Three options. I ruled out two.

Option A — Rejected
Full automation — AI auto-applies recommendations to client portfolios
The quickest way to make the AI look useful, and a non-starter. The rules require a person to sign off on advice, so the firm could not let software send recommendations on its own. Compliance said so on day one. The tool did not need to work around advisors. It needed to make them faster.
Option B — Rejected
AI as a "suggestion bot" — vague, low-confidence outputs only
The safest option. Only ever show soft suggestions and let the advisor do the real work. I ruled it out because the pilot feedback was blunt: advisors called it too vague to bother with. Advice that commits to nothing saves nobody any time, so people simply skipped it.
Option C — Chosen
AI as auditable co-pilot — confident outputs with full provenance + edit affordances
Show every recommendation with three things attached. How sure the AI is, as a number. Links to the actual data behind each claim. And editable text, so the advisor can change the wording before a client sees it. The AI does the heavy lifting and the advisor stays responsible. This was the version compliance, advisors, and the AI team could all live with.
05
Decision & Tradeoff

We made it slower on purpose. Here is what that cost.

Every recommendation now takes 2.4 seconds longer to appear. That is the price of working out a confidence score, fetching the sources behind each claim, and laying out the editable fields. The product team pushed back on the delay. Here is how we defended it.

+ Gained
Advisors trusted it. Acceptance of AI advice went from 38% to 67%, and the share of new advisors who got started rose 16 points. Support tickets asking why the AI suggested something fell by an estimated 60%.
− Lost
2.4 seconds of waiting on every recommendation. Against the old version the AI feels slower. An advisor producing 30 recommendations a day waits about 70 seconds longer in total.
Why accepted
A slower tool that gets used beats an instant one that does not. Four out of five rejections were about trust and none of them were about speed. Fixing the thing that was actually blocking people was the only sensible call.

The three things we shipped

How sure the AI is. Every recommendation carries a score from 0 to 100 with a high, medium, or low band. Advisors can filter to high confidence only when they want to move quickly. The number comes from the model's own output rather than from a guess.

A source on every claim. Each statement links to the portfolio data, market signal, or client preference behind it. Click a link and the raw data opens beside it. This became the most-used feature in the release, with 84% of active advisors opening sources daily.

Text the advisor can change. Advisors can edit any AI output before a client sees it. Edits are saved, carry the editor's name, and survive if the AI rewrites the text. Compliance required this. It also turned out to be the part advisors valued most.

06
Wireframes

Concept sketches: the structure, not the real screens

The shipped product is covered by an NDA. These are redrawn concept sketches made to explain the design thinking. They are deliberately rough, use made-up numbers, and carry none of the real visual design or client data.

BEFOREAFTER SEND the answer, and nothing else. no sources, no reasoning, no way to edit before sending CONF. 82 SEND EDIT score in the first place the eye lands · sources underlined under each claim · the text can be changed before it goes

01 — The change: an answer becomes a checkable answer

Everything the advisor needed in order to put their name to the advice was missing. The redesign attaches three things to the same card: how sure the AI is, where each claim came from, and the ability to change the wording. None of them alters the advice itself.

HOW SURE THE AI IS 82 / 100 LOWMEDIUMHIGH band as well as number, so it reads at a glance FILTER THE QUEUE HIGH ONLY ALL an advisor in a hurry can work only the ones the AI is most sure about the number comes from the model’s own output, not from a guess. it is the first thing on the card because trust gets decided before the advice is read.

02 — How sure the AI is

In pilot sessions advisors scanned each card from the top right, so that is where the score sits. They judge whether to trust it before reading the advice rather than after. The band matters as much as the number, because a figure on its own gives nothing to compare against.

FIRST TRY — PILLSSHIPPED — UNDERLINES HOLDINGSRISK PROFILE read as category tags. rarely clicked. the working was there and nobody saw it. click → the raw data opens beside the claim it supports the convention they have read in documents for twenty years. about three times the clicks on first sight.

03 — Sources: the version that failed, and the one that worked

This is the detail I would have got wrong on taste alone. Pills looked tidier and tested badly, because a rounded shape already means “category” to anyone who uses software. Underlines carry twenty years of meaning that says “there is something behind this”.

GETTING STARTED — WATCH, PRACTISE, DO 1 · WATCH 2 · PRACTISE 3 · DO a ready-made example.the score and sources explained. a made-up client.editing carries no risk here. a real client, at last.first real data on screen. the drop-out had one location: the first recommendation an advisor ever saw. so nothing real is on screen until they have already seen the trust layer work twice.

04 — Getting started: teaching trust before asking for it

Nearly half of new advisors dropped out at their very first recommendation. Putting a real client on that screen asks someone to trust a system they have never seen work. Two rehearsals first, and the real thing third.

07
Design Decisions

One question sat behind every screen: could the advisor defend this to compliance?

The recommendation card, explained

The confidence score sits top right, not at the bottom.

In pilot sessions advisors scanned each card starting at the top right. Putting the score in the first place they looked meant they judged whether to trust it before reading the advice rather than after.

Sources are underlined, not styled as tags.

The first version put sources in rounded pills. It tested badly. Advisors read pills as category labels and rarely clicked them. Plain underlines, the same convention they have seen in documents for twenty years, got about three times the clicks on first sight.

Getting started: teaching trust before asking for it

The drop-out had one clear location, the first recommendation an advisor ever saw. So we rebuilt the opening. The first screen walks through a ready-made recommendation and explains the score and the sources before any real data appears. The second gives the advisor a practice recommendation on a made-up client, where editing carries no risk. Only the third screen uses their own clients. Watch, then practise, then do.

The design system

I built and looked after the shared Figma library the recommendation card lives in. It holds reusable parts, rules for using them, and clear labelling of which elements are written by the AI and which by the advisor, a distinction that matters when the record is checked later. Handover to engineering got quicker once it was in place, and later features were faster to design and faster to review.

CONF. 82 SENDEDIT 1. trust is judged here first 2. the working, under the claim 3. the advisor still owns the wording
One card, annotated — the three things an advisor needs before signing their name
08
Outcome

Advisors using it: 52% → 68%. AI advice accepted: 38% → 67%.

What changed

New advisors who got started went from 52% to 68%, a rise of 16 points, measured in the firm's own analytics across the two releases after the redesign. The gain was largest among advisors with 5 to 15 years of experience, the same group the research had flagged as hardest to convince.

AI advice accepted went from 38% to 67%, a rise of 29 points, counted per recommendation across everyone using the tool. When advisors could check the working, they accepted the advice at nearly twice the rate.

Support tickets fell by an estimated 60%. Tickets tagged as the AI being unclear dropped in the first quarter after release, based on the ticketing system's own tagging and checked against a manual read of 200 tickets. The comparison uses matching 90-day windows before and after.

What that is worth to the business

The platform serves around 3,000 advisors inside the firm. A 16-point rise means roughly 480 more advisors actively using it. Each advisor looks after about $200M for clients, a range the firm publishes, so between them those advisors manage somewhere near $96B. The design did not bring that money in. It removed the thing stopping the AI from helping with money the firm already had.

"The redesign didn't change the AI model. It changed whether advisors believed the AI model. That's the entire game in regulated industries."

— Product lead, internal release retro
09
What's Next

Next: showing confidence differently to different advisors

67% acceptance is good, and the remaining third is informative. The reasons tell a split story. Advisors with 15 or more years call a single confidence number too definite and would rather see a range. Advisors with under five years turn advice down because they want more explanation, not less. One design cannot serve both.

My theory: show experienced advisors a range, such as 78 to 84%, and show newer advisors a single number with a plain-English explanation beside it. I expect that to add another 8 to 12 points of acceptance. The tracking is already in place, split by experience, along with a plan to roll it back if acceptance among senior advisors drops instead.

15+ YEARS — WANTS A RANGEUNDER 5 YEARS — WANTS THE WHY 78 – 84 a single figure reads as falsely precise 82 the number, plus a sentence saying what drove it one design cannot serve both, so the next release tests two tracking is already split by experience, with a rollback ready
Why the last third of rejections needs a different answer for each group
10
Constraints

A regulated firm, an AI nobody can see inside, and two releases to work with

This is a bank, so the usual consumer app freedoms do not apply. Every decision had to pass three tests before shipping: compliance, whether anyone could explain what the AI had done, and whether advisors would trust it. Time was not the real limit. How much I could change inside two releases was.

I protected scope and quality. Time was what I let move.

Compliance reviews take as long as they take, and the AI model itself could not be changed. So the thing I let slip was the schedule. The work shipped across two releases instead of one, with the confidence score first and the sources and editing second. Cutting scope was risky, because half a trust layer is worse than none. Cutting quality was not an option in a regulated business. So the timeline moved.

01
Finding Assumptions Without Bypassing Compliance
Normal user research was not possible, because advisor work is full of confidential client information. So I designed around it. When sitting with advisors I recorded how they moved through the tool and never what was on screen. Twelve sessions gave me the human side and the analytics gave me the numbers.
02
What Moved to V3
Showing confidence differently by experience level. Automatic flags when the AI's confidence drops below a set level for a client. Sources shown as charts rather than only text. All written up with owners and measures, and all already tracked, so the next sprint starts with a brief instead of a blank page.
03
How AI Tooling Compressed Design Cycles
Figma Make sped up building the card variations, since the recommendation card alone needed 14 versions for the different confidence bands. Claude Code let me prototype how the sources expand so I could test the interaction before engineering committed to it. Together that shipped a compliance-heavy tool in two releases instead of three.

"In regulated AI, the design constraint isn't time. It's how much you can change before the legal review cycle resets. Pick your changes carefully — and make every one count."

11
Reflection

What I took from this project

01
Trust Is The Product
In a regulated business, what stops people using AI is not the quality of the answer. It is having no way to check it. The most important thing on the screen is not what the AI said. It is the working behind it.
02
Slower Beats Untrusted
Speed only matters once people are actually using the thing. Paying 2.4 seconds was right, because the alternative was a faster tool nobody trusted enough to open.
03
Design The Audit Trail
Every AI screen I design now starts with a different question. Not how to display the model's answer, but how a person proves they checked it. That one change reorganises everything, from the confidence score to the edit history to how many sources you show.

Conclusion

In regulated work, shipping the model is the easy part. The job is giving a person everything they need to stand behind it.

The trust patterns shipped at Goldman Sachs — confidence visibility, citation-backed recommendations, editable AI outputs — are now adopted as the platform standard for AI-output surfaces across the broader wealth advisory product line.

Glad we could cross paths.
Out of anywhere you could be, you're here.

Next project

DreamCollege AI
All projects