Led product design for an AI tool that drafts investment advice for financial advisors. I added the things advisors needed before they would trust it: where every number came from, how sure the AI was, and the ability to edit the wording before it reached a client.
Before we start — what this is
30 seconds · no finance background needed
The part of a bank that looks after money for rich clients.
A real person at Goldman Sachs assigned to those clients. They decide what a client should do with their money, then explain it to them.
An AI that writes the first draft of that advice. It reads the client's holdings and drafts a recommendation in seconds instead of hours.
The advisor's name goes on the advice, not the AI's. Every piece of it is kept on file for seven years, and the advisor answers for it if it is wrong.
Where the advice comes from, and who answers for it
The AI showed its answer and nothing else. No sources, no reasoning, no way to check it. Advisors were being asked to stake their name on something they could not inspect — so half of them never used it.
The problem
An AI wrote investment advice for financial advisors but showed no working. Advisors had to put their name to advice they could not check, so half of them never started using it.
What I did
Attached three things to every recommendation: how sure the AI is, a link to the data behind each claim, and text the advisor can edit before a client sees it.
The result
The four decisions — jump to any sketch
The Stakes
Outcomes — two product releases
Goldman Sachs builds AI tools for the people who manage money for wealthy clients. The AI drafts investment recommendations, sums up a client's portfolio, and writes client-ready notes in seconds instead of hours. The technology did its job. Almost nobody used it.
I joined the team that owns this tool. My work covered how advisors get started with it, how the AI's advice is presented, and what sits between what the AI writes and the advisor who has to sign their name to it. In this industry a wrong recommendation is not only a bad experience. Every piece of advice is filed and kept for seven years, and the advisor answers for it.
Headquarters
New York, USA
Founded
1869
Industry
Investment Banking · Wealth Mgmt
AUM
~$2.8T (public, 2024)
Platform Users
~3,000 advisors (internal SaaS)
Compliance Surface
SEC · FINRA · internal IB controls
When I joined, only 52% of new advisors ever got going with the tool. The AI presented its advice as finished, confident text. It gave no reasoning, no source for any figure, and no way to change the wording before it went to a client. Advisors were being asked to put their name to something they could not inspect.
"I'm not going to send a client a recommendation I can't defend to compliance. The AI gives me a clean answer, but I can't see where it came from. So I just... don't use it."
— Senior wealth advisor, week-2 platform exit interviewI treated the low take-up as a research question first. The easy reading was that the screens confused people, which points at a layout fix. The evidence pointed somewhere else. Advisors did not doubt the answer. They doubted anything they could not trace back to a source.
Analytics showed where advisors dropped out, but not why any single recommendation was turned down. So I worked with the AI team to add a one-tap “this was not useful” button with an optional reason. Within three weeks we had reasons for more than 1,200 rejections. The top three were no visible source (38%), cannot edit before sending (24%), and does not say how sure it is (19%). Four out of five rejections were about trust, not about the quality of the advice.
The chosen path was the only one that let the AI, the advisor, and the compliance officer all defend the same output.
Every recommendation now takes 2.4 seconds longer to appear. That is the price of working out a confidence score, fetching the sources behind each claim, and laying out the editable fields. The product team pushed back on the delay. Here is how we defended it.
81 % of pilot rejections were trust-layer problems, not speed problems. The scale was never even close.
How sure the AI is. Every recommendation carries a score from 0 to 100 with a high, medium, or low band. Advisors can filter to high confidence only when they want to move quickly. The number comes from the model's own output rather than from a guess.
A source on every claim. Each statement links to the portfolio data, market signal, or client preference behind it. Click a link and the raw data opens beside it. This became the most-used feature in the release, with 84% of active advisors opening sources daily.
Text the advisor can change. Advisors can edit any AI output before a client sees it. Edits are saved, carry the editor's name, and survive if the AI rewrites the text. Compliance required this. It also turned out to be the part advisors valued most.
The shipped product is covered by an NDA. These are redrawn concept sketches made to explain the design thinking. They are deliberately rough, use made-up numbers, and carry none of the real visual design or client data.
01 — The change: an answer becomes a checkable answer
Everything the advisor needed in order to put their name to the advice was missing. The redesign attaches three things to the same card: how sure the AI is, where each claim came from, and the ability to change the wording. None of them alters the advice itself.
02 — How sure the AI is
In pilot sessions advisors scanned each card from the top right, so that is where the score sits. They judge whether to trust it before reading the advice rather than after. The band matters as much as the number, because a figure on its own gives nothing to compare against.
03 — Sources: the version that failed, and the one that worked
This is the detail I would have got wrong on taste alone. Pills looked tidier and tested badly, because a rounded shape already means “category” to anyone who uses software. Underlines carry twenty years of meaning that says “there is something behind this”.
04 — Getting started: teaching trust before asking for it
Nearly half of new advisors dropped out at their very first recommendation. Putting a real client on that screen asks someone to trust a system they have never seen work. Two rehearsals first, and the real thing third.
The confidence score sits top right, not at the bottom.
In pilot sessions advisors scanned each card starting at the top right. Putting the score in the first place they looked meant they judged whether to trust it before reading the advice rather than after.
Sources are underlined, not styled as tags.
The first version put sources in rounded pills. It tested badly. Advisors read pills as category labels and rarely clicked them. Plain underlines, the same convention they have seen in documents for twenty years, got about three times the clicks on first sight.
The drop-out had one clear location, the first recommendation an advisor ever saw. So we rebuilt the opening. The first screen walks through a ready-made recommendation and explains the score and the sources before any real data appears. The second gives the advisor a practice recommendation on a made-up client, where editing carries no risk. Only the third screen uses their own clients. Watch, then practise, then do.
I built and looked after the shared Figma library the recommendation card lives in. It holds reusable parts, rules for using them, and clear labelling of which elements are written by the AI and which by the advisor, a distinction that matters when the record is checked later. Handover to engineering got quicker once it was in place, and later features were faster to design and faster to review.
New advisors who got started went from 52% to 68%, a rise of 16 points, measured in the firm's own analytics across the two releases after the redesign. The gain was largest among advisors with 5 to 15 years of experience, the same group the research had flagged as hardest to convince.
AI advice accepted went from 38% to 67%, a rise of 29 points, counted per recommendation across everyone using the tool. When advisors could check the working, they accepted the advice at nearly twice the rate.
Support tickets fell by an estimated 60%. Tickets tagged as the AI being unclear dropped in the first quarter after release, based on the ticketing system's own tagging and checked against a manual read of 200 tickets. The comparison uses matching 90-day windows before and after.
The platform serves around 3,000 advisors inside the firm. A 16-point rise means roughly 480 more advisors actively using it. Each advisor looks after about $200M for clients, a range the firm publishes, so between them those advisors manage somewhere near $96B. The design did not bring that money in. It removed the thing stopping the AI from helping with money the firm already had.
"The redesign didn't change the AI model. It changed whether advisors believed the AI model. That's the entire game in regulated industries."
— Product lead, internal release retro67% acceptance is good, and the remaining third is informative. The reasons tell a split story. Advisors with 15 or more years call a single confidence number too definite and would rather see a range. Advisors with under five years turn advice down because they want more explanation, not less. One design cannot serve both.
My theory: show experienced advisors a range, such as 78 to 84%, and show newer advisors a single number with a plain-English explanation beside it. I expect that to add another 8 to 12 points of acceptance. The tracking is already in place, split by experience, along with a plan to roll it back if acceptance among senior advisors drops instead.
This is a bank, so the usual consumer app freedoms do not apply. Every decision had to pass three tests before shipping: compliance, whether anyone could explain what the AI had done, and whether advisors would trust it. Time was not the real limit. How much I could change inside two releases was.
Compliance reviews take as long as they take, and the AI model itself could not be changed. So the thing I let slip was the schedule. The work shipped across two releases instead of one, with the confidence score first and the sources and editing second. Cutting scope was risky, because half a trust layer is worse than none. Cutting quality was not an option in a regulated business. So the timeline moved.
"In regulated AI, the design constraint isn't time. It's how much you can change before the legal review cycle resets. Pick your changes carefully — and make every one count."
Conclusion
In regulated work, shipping the model is the easy part. The job is giving a person everything they need to stand behind it.
The trust patterns shipped at Goldman Sachs — confidence visibility, citation-backed recommendations, editable AI outputs — are now adopted as the platform standard for AI-output surfaces across the broader wealth advisory product line.
Glad we could cross paths.
Out of anywhere you could be, you're here.
Goldman Sachs
I can find the real reason people avoid a product
Take-up was low and the obvious answer was confusing screens. Analytics, twelve advisor sessions, and a reason-code study all said the same thing instead: advisors would not sign their name to advice they could not check.
Four out of five rejections were about trust, not answer quality
I can design AI that a regulated business will actually allow
Full automation was ruled out by law and a vague assistant was useless. I designed the middle: the AI does the work, and the advisor keeps the responsibility and the final edit.
Compliance, advisors, and the AI team all signed off
I can argue for a trade-off that looks bad on paper
The trust layer made every recommendation 2.4 seconds slower and the product team pushed back. I showed that none of the recorded rejections were about speed.
Acceptance nearly doubled, 38% to 67%
I test the interface details rather than trusting taste
Sources first appeared as rounded pills. Advisors read them as labels and ignored them, so I moved to plain underlines, the convention they already knew from documents.
About three times the clicks on first sight
I can build systems and ship them with engineers
I built the shared Figma library the recommendation card lives in, including clear labelling of what the AI wrote versus what the advisor wrote.
A compliance-heavy tool shipped in two releases instead of three
Details in this case study are covered by an NDA and have been generalised.