Cited from real sources 6 min read Updated September 2026

A framework by Ronny Kohavi

Ronny Kohavi's OEC: Deciding What an A/B Test Optimizes For

The Overall Evaluation Criterion is the metric you judge a test against. You pick it before the test runs. Ronny Kohavi built the testing platforms at Microsoft and Amazon. He ran the same work at Airbnb. His warning: most teams skip this step. They ship a test and watch revenue. They miss that revenue rose because the product got worse.

The question nobody answers first

What are you optimizing for?

Answer "revenue" and the experiment will deliver it, by making the product worse in ways the metric cannot see.

Ronny Kohavi Lenny's Podcast Watch at 28:13

The framework

Every metric you can move, you can also game

Teams new to experimentation treat the success metric as an afterthought. Kohavi treats choosing it as the hard part, harder than the statistics.

the OEC or the overall evaluation criteria is something that I think many people that start to dabble in A/B testing miss
Kohavi, on the step teams skip Watch at 28:13

The obvious answer is the dangerous one. Optimize for revenue and you will get revenue. There are always ways to extract more of it in the short term, at the expense of the people paying you.

it's very easy to say we're going to optimize for money, revenue, but that's the wrong question
Kohavi, on the obvious metric Watch at 28:32

His example is search advertising. Put more ads on the results page and you will make more money, with no doubt and no delay. The question the revenue metric cannot answer is what that does to the person who came to search, and whether they come back next week. So the OEC is not one number. It is a number plus a guard.

there has to be some countervailing metric that tells you how do I improve revenue without hurting the user experience
Kohavi, on the shape of an OEC Watch at 28:45

How to apply it

How do you define an OEC for your test?

Six steps. The first three define the criterion, the last three keep you honest about the result.

  1. 1

    Write down the success metric before you build.

    Deciding after you see the data lets you re-score a losing test against whichever metric happened to move.

  2. 2

    Ask how you would cheat it.

    Name the change that would move the metric while making the product worse. More ads. More emails. A harder-to-find cancel button. If cheating is easy, the metric is incomplete.

  3. 3

    Add the countervailing metric that catches the cheat.

    Pair the thing you want to grow with the thing you refuse to damage. A win has to clear both, otherwise it is not a win.

  4. 4

    Expect most ideas to lose.

    Kohavi's framing is that you are trying things. Most will fail. A few will be home runs. A team that wins most of its tests is testing only safe changes.

  5. 5

    Distrust the enormous result.

    Twyman's law: a figure that looks remarkable is likely a mistake. A 30% lift from a button colour is a bug in the instrumentation until proven otherwise.

  6. 6

    Keep experimenting when things get hard.

    A crisis is the moment teams abandon measurement and go on instinct. Kohavi's counter is that your instincts were wrong most of the time in calm conditions. So a crisis is a strange moment to start trusting them.

things that are too extreme, your meter should go up and say hey I don't believe that
Kohavi, on Twyman's law Watch at 74:47

Boundary conditions

When does A/B testing fail you?

Works best when

  • You have enough traffic for a result to mean anything
  • The change is reversible and the decision is open
  • Someone owns the OEC and will not renegotiate it after the readout

Fails when

  • You are pre-product-market-fit, where the change you need is too big to A/B
  • The OEC is a single revenue number with nothing guarding it
  • The team ships a startling result instead of investigating it

The first boundary is the one founders should respect most. Experimentation is a tool for improving something that already works. Early on, the honest move is Sean Ellis's product-market fit test or Bob Moesta's switch interview. Not a 50/50 split on a landing page nobody visits.

The failure rate is also worth internalising before you build a testing culture, because it is the number that makes the OEC matter. Kohavi's point about crisis decision-making is a point about baseline humility.

if in peacetime you're wrong two-thirds to 80 of the time
Kohavi, on why a crisis is a bad time to trust instinct Watch at 49:17

Two-thirds to eighty percent of your ideas do not work when conditions are calm and you have data. That is the strongest argument for writing the OEC down in advance, and the strongest argument against deciding by seniority. It is the same discipline Annie Duke applies to kill criteria: commit to what would count as failure while you still have no stake in the answer.

The sources

Where Kohavi discusses this

Where experts disagree

Where operators disagree: define the metric, or trust the taste?

Ronny Kohavi

says write the success metric down before you build, ask how you would cheat it, and add the countervailing metric that catches the cheat. Any metric you can move you can also game, and an experiment with no OEC will faithfully optimise the wrong thing.

Archie Abrams (Shopify)

runs core product without traditional KPIs, on the view that intuition and taste guide the decisions that matter and a metric target quietly narrows what the team is willing to try.

The split maps onto what you are building. Growth and optimisation work needs Kohavi, because that is where gaming is cheap and the effects are small enough to fool you. A core product bet has no comparable baseline to measure against, which is exactly where a scorecard turns into theatre.

Useful? Send it to whoever is about to ship a test with one success metric.

Want the full playbook?

Get 128 product management frameworks.

34 frameworks 21 rules 65 heuristics & principles 49 operators

From Stewart Butterfield, Ami Vora, Codebase Guidelines, and 46 more. Drop one .md into Claude, Cursor, or ChatGPT. Your AI cites practitioners, not guesses.

See the pack

Instant .md download · One-time purchase · No subscription

New experts every week

Know when the next expert lands.

Gavel adds new operators to the database every week, each one with cited frameworks you can check and a note on where they disagree with the others. You found this page by searching. Get the next one by email instead.

53 experts 66 cited frameworks

Latest: Amjad Masad on Levels of AI Autonomy

One email a week, only when new experts shipped. Unsubscribe with one click. We never sell or share email.

Related frameworks