Notes from a conversation on The Experimentation Edge
I recently joined Ashley Stirrup on GrowthBook’s podcast, The Experimentation Edge, to talk about synthetic audiences. Not the concept. The build. What it actually took to stand one up inside a large financial services organization, and what happened when we finally pointed it at a live test.
The short version is that it worked. The longer version is more interesting, because the model was never the hard part.

The audience problem
Every experimentation program eventually runs into the same wall. You have more hypotheses than you have people to test them on. Traffic is finite. Customer patience is finite. Every live impression you spend on a mediocre variant is an impression you cannot spend on a good one.
Most organizations respond by testing less. They prioritize, they queue, they wait. The roadmap slows to the speed of available audience.
Synthetic audiences change the shape of that constraint. If you can build profiles that are grounded in your real customer data and use them to rank content by likelihood of engagement, you can do the sorting work before anything touches a live customer. The weak ideas get filtered out cheaply. The live test becomes the final exam rather than the first draft.
We built ours entirely in house. No vendor, no black box, no license renewal. The profiles are shaped by data we already had, and the ranking output is something my team can inspect, question, and correct.
The first synthetic selection we put into a live A/B test beat the control.
One result is not a validation
I want to be careful here, because this is where these stories usually get oversold.
One win is not proof that the method generalizes. It is a signal worth following, not a conclusion worth building a strategy on. Synthetic audiences carry real failure modes. They can encode the same biases as the data underneath them. They can be confidently wrong in ways that are difficult to detect, because there is no real person on the other end to contradict them. A synthetic audience that agrees with you is not evidence. It is a mirror.
The discipline is to keep validating against live behavior rather than treating the simulation as the ground truth it is standing in for. Synthetic work earns trust the same way any model does, which is slowly and with a record of calibration behind it.
The looping metric
We also talked about something less glamorous and, in day to day terms, more useful.
We did not have heat-mapping tools. We wanted to know where customers were getting stuck. So we built a looping metric in SQL that identifies where people cycle back through the same paths instead of moving forward.
Looping is a behavioral signal that shows up clearly in data you already collect. A customer who visits the same page three times in a session is telling you something. They are not browsing. They are searching for something they cannot find, or they are trying to confirm something the page did not make clear. Conversion metrics will not surface that. The customer may even convert eventually, which makes the friction invisible in the outcome data.
The broader point is that a missing tool is not always a blocker. Sometimes it is a prompt to define the behavior you care about precisely enough to write the query yourself. The definition is the hard part. The SQL is not.
The part that was actually hard
Here is what I keep coming back to. Building the audience model was tractable. Getting a large, risk averse organization to run honest experiments is the real work.
Testing at big companies has a tendency to get watered down into “let’s just try something.” The word experiment stays in the vocabulary while the rigor drains out of the practice. There is no sharp hypothesis, no defined success criterion, no willingness to be wrong. A change ships, a dashboard moves, and someone declares it a win. Nothing was learned, because nothing was ever really asked.
This is not usually a skills problem. It is an incentive problem. When teams are evaluated on wins rather than on learning, they stop designing tests that could fail. They test the safe thing. The program looks busy and produces almost nothing.
Why a center of excellence has to publish its losses
The other reflex that kills experimentation programs is institutional memory of the wrong kind. Someone proposes an idea and a voice in the room says “we tried that years ago.” No documentation, no context, no record of how it was tested or whether the conditions still hold. The idea dies to a memory that cannot be examined.
A center of excellence fixes this only if it shares losses as openly as wins. A repository of successes is a marketing artifact. A repository that includes what failed, under what conditions, and what the team concluded is an asset. It lets the next person pick up an old idea and ask the right question, which is not “did this work” but “why didn’t it, and has anything changed since.”
Failed tests are not embarrassments to be quietly retired. They are the most honest thing an experimentation program produces.
Listen to the episode
We covered all of this and more with Ashley, including who this work is actually for and where I think synthetic experimentation goes next. If you work in product, data science, marketing, or experimentation leadership inside a large organization, there should be something here you recognize.
Available now on Spotify, Apple Podcasts, and YouTube: https://share.transistor.fm/s/c63f5e2d

Leave a comment