For four years I shipped an experiment almost every week against the funnel that starts a Rocket Mortgage application, and most of them lost. That was the design. The machine that kept the losing affordable for that long is the part worth copying.
Most teams that say they run an experimentation program actually run occasional redesigns. A big swing every quarter, a launch, a wait, another swing. That can work.
What it cannot do is compound, because you only learn four times a year, and each lesson arrives attached to a change too large to trust or to trace.
At Rocket Mortgage I owned growth experiments on the questionnaire that starts a home loan. Internally it was called Launchpad: the short sequence of questions a visitor answers about their property, their finances, and their goal before we routed them into a mortgage.
At that traffic, a fraction of a percent is real money, so the job was to find those fractions on purpose, over and over, for years.
The individual wins live in the Rocket Mortgage case study. This piece is about the operating system underneath the pace.
Treat velocity like a number you report
The first rule was to make experiments per week a metric with a target, sitting on the dashboard beside leads and revenue. One a week. Fifty-odd a year.
That is a modest number beside Booking.com, which runs about 25,000 tests a year, and the mechanism is the same at either size.
When velocity is something you have to report, the backlog of ideas can never go quiet, because an empty week becomes a visible miss that someone has to explain in the room.
A quarterly-redesign team measures outcomes and hopes the ideas show up when the next big project starts. A weekly team measures the ideas themselves.
Once the count is on the board, the behavior underneath it changes: designers keep a backlog warm, analysts pre-clear metrics before a test needs them, and the build queue stays full.
Velocity is a metric. Treat experiments per week like revenue and the pipeline of ideas never gets to silt up.
A win rate you expect to miss most weeks
The benchmark most teams quietly aim for is a win rate near 25%. We set that as the floor and cleared it, landing at 42% across a quarter of twenty-four tests.
That floor tracks what Microsoft's experimentation team reported from its own platform, where only about one-third of well-designed experiments improved the key metric they were built to move.
That is the whole design. If you target a win rate you can hit most of the time, you are shipping timid changes that were going to work anyway, and learning almost nothing. A 25% goal forces bigger, stranger bets, and bigger bets lose more often.
A 42% win rate means most experiments lost, on purpose. The cadence only compounds when losing is designed to be cheap.
Make a loss cost almost nothing
Designing for cheap losses is the actual craft. A loss should cost a week of build and then nothing: no engineering debt left behind, no cleanup sprint, no meeting to assign blame.
We kept variants small enough that shipping one and killing it a month later touched barely any code. We reused instrumentation across tests, so a new experiment inherited its measurement instead of rebuilding it. And every test launched with its kill rule already written, so ending a loser was an act of bookkeeping.
When a loss is that cheap, the emotional weight comes off the whole program. Nobody defends a dying test, because nothing was staked on it and nobody spent a month of their life on it.
That is what lets a team keep firing every single week for years without flinching, and without the slow drift toward only shipping the safe bets that never move a number.
Put legal inside the loop
The screen with the most legal exposure in the entire product was the one that collects a prospect's phone number, because how and when you ask for permission to contact someone sits on top of real telemarketing law.
The Telephone Consumer Protection Act lets a consumer recover $500 in damages for each violation, and up to three times that when the violation is willful or knowing.
The default posture toward a screen like that is to freeze it: too risky to touch, so nobody tests it, so it never improves while everything around it does.
We did the opposite. We brought the legal team into the design of the experiment from the first sketch, and their constraints shaped the hypothesis itself.
The reorder experiments that grew out of this work, where the sequence of questions did the heavy lifting, get their own deep dive on the reorder wins.
The riskiest screen in the product was the one we tested most, because legal sat inside the experiment design instead of at the end of it.
See why an experiment behaved that way
Numbers tell you an experiment lost. They rarely tell you why.
Partway through the program we wired session-replay tooling into individual experiments so we could watch anonymized recordings of real people moving through a specific variant, which was a program first.
A test that lost on paper would suddenly make sense: someone hesitating at a newly worded line, a thumb hovering over the wrong field, a layout that broke on one class of device.
That turned losses into inputs. A dead experiment became the first draft of the next hypothesis, because we could see the exact moment the variant lost the visitor and change one thing about it.
Debugging experiments this way is what kept the idea supply honest: every readout, win or lose, produced a sharper question for the following week instead of a shrug.
What the machine produced
Run this way, the wins accumulated. In one representative quarter the winning experiments drove an estimated $8.9 million in incremental revenue, without adding a dollar of acquisition spend to the funnel.
Stretched across four years, the program's cumulative wins are the subject of the case study and the roughly $2 billion of incremental loan volume it describes.
The rhythm that reads the results each week is its own discipline, the weekly growth review, and this whole engine is one instance of the broader growth operating system I install.
Build your own cadence
You do not need Rocket's traffic to run this. If your experimentation program is really a redesign every quarter, the machine is the piece you are missing. The parts, in order:
- Put experiments per week on the dashboard as a target, and treat an empty week as a miss you have to explain.
- Set a win-rate goal low enough that losing is the base case, around one in four, so the program funds ambitious tests instead of timid ones.
- Engineer losses to be cheap: small variants, shared instrumentation, and a kill rule written before the test launches.
- Bring legal and compliance into the design of your riskiest tests, so the scariest screens become testable instead of frozen.
- Wire session-replay tooling into experiments so every loss tells you why it lost, and feeds the next hypothesis.
A test a week for four years takes less stamina than it sounds. The system made each test cheap enough to lose and quick enough to replace, so the losing never had to stop, and the winning had fifty chances a year to happen.
If your experimentation program is really an occasional redesign, this is the machine I build to change that. Let's talk.