Blog

Research bets that end in a decision instead of fading out

How to frame exploratory research as a bet with the hardest question first and written kill criteria, so every project ends in a clear go or stop.

Every specialist group has one: the exploratory project that started eighteen months ago, has two engineers on it, produces a "promising" demo each quarter, and has never been asked to justify itself against anything. Nobody decided to keep it going. Nobody decided to stop it either. If you lead research, platform or other specialist work, your credibility depends on the opposite pattern: small bets, clearly framed, that end in a decision on a known date. This post gives you a one-page bet brief, a way to sequence the work so bad bets die early, and kill criteria that actually get used.

Why exploratory work doesn't end

Research projects rarely fail loudly. They fade, for reasons that are predictable:

  • Success was never defined. If the goal is "explore on-device inference", every result is progress. Nothing can count against the project, so nothing ends it.
  • The easy parts get done first. Teams build the pipeline, the dashboard and the demo UI because those show visible progress. The question that decides whether the idea works at all stays open.
  • Sunk cost builds up. The more you've invested, the harder it is to walk away. Barry Staw's research on escalation of commitment, starting with his 1976 paper "Knee-deep in the big muddy", showed that people who feel responsible for an earlier decision tend to put more resources into a failing course of action, not fewer.
  • Stopping looks like failure. If the people on the project expect a stopped bet to count against them in their review, they'll keep finding reasons to continue.
  • Nobody owns the stop decision. The team running the work is the worst-placed group to kill it, and nobody else is watching closely enough to.

Write a bet brief before any work starts

Treat every exploratory project as a bet with a fixed stake, and write the terms down on one page before anyone writes code:

Bet: <short name>                    Owner: <person running it>
Decider: <person who makes the go/stop call, not the owner>

Question
  The one thing we're trying to find out, as a question
  with a yes/no answer.

Hypothesis
  What we believe, specific and measurable.

Why it matters
  What we'd do differently if the answer is yes, and roughly
  what that's worth.

The hardest unknown
  The part most likely to make this impossible. We test this first.

Stake
  People, weeks and money. Nothing beyond this without a new decision.

Kill criteria
  "If by <date> we have not <measurable state>, we stop."
  One to three of these, no more.

Outcomes
  Go: what happens next, who funds it.
  Stop: what we write up, where the people go.
  Pivot: allowed only with a new brief.

The two fields people skip matter most: the hardest unknown and kill criteria.

Train the monkey first

Astro Teller, who runs X (Alphabet's "moonshot factory"), describes this with an image: suppose your goal is a monkey performing a trick while standing on a pedestal. The pedestal is easy; you know you can build it. The monkey is the real problem. If you can't train the monkey, the pedestal is wasted, yet teams often build the pedestal first because it feels like progress. Teller's rule, which Annie Duke discusses at length in her book Quit, is to tackle the monkey first. Duke adds the important point that pedestal-building doesn't just waste time; it piles up sunk cost that makes quitting harder later.

In technical R&D, pedestals are usually infrastructure and presentation, and monkeys are usually a physical, statistical or economic limit. Some hypothetical examples:

Bet Pedestal (do later) Monkey (do first)
On-device ML inference Model download service, settings UI, telemetry Can a compressed model hit the accuracy bar on the slowest supported device?
New storage engine Admin tooling, migration scripts, docs Does it hold write throughput under compaction at production data volumes?
Internal developer platform Portal, service catalog, onboarding guide Will two real product teams move a live service onto it without being ordered to?
Automated code migration CI integration, reporting dashboard Does the tool produce changes reviewers accept unedited on messy, real code?

A useful test when you review a plan: if the first month of work succeeds perfectly, have we learned whether the idea can work? If the answer is no, the plan starts with a pedestal. Reorder it.

Sometimes the monkey can't be tested without a bit of pedestal, such as a minimal harness to run the benchmark. That's fine. Build the least pedestal that lets you test the monkey, and no more.

Kill criteria that bite

Duke's advice in Quit is that good kill criteria combine a state and a date: if by a given date you are not in a given measurable state, you stop. Either half alone is weak. A date without a state becomes "let's review it in March", and in March the team explains why it's nearly there. A state without a date ("we'll stop if latency can't get below 200 ms") never triggers, because it can always get below 200 ms with a bit more time.

Compare:

Weak Strong
"We'll reassess if it isn't working." "If by March 15 the p95 latency on the reference device is above 200 ms, we stop."
"Adoption should grow over the pilot." "If by the end of week 8 fewer than two product teams run a production service on the platform, we stop."
"Accuracy needs to be competitive." "If by June 1 the compressed model scores below 92% on our evaluation set, where the current server model scores 95%, we stop."
"We'll give it a quarter." "Six weeks, two engineers. At week 6 we either meet criteria 1 and 2 or we stop; no extensions without a new brief."

Rules that keep kill criteria honest:

  1. Write them before the work starts. Criteria written after early results come in are shaped by those results.
  2. Tie them to the hardest unknown. Criteria about the pedestal ("the pipeline is built by April") measure effort, not whether the idea works.
  3. Make the measurement unambiguous. Name the device, the dataset, the metric and who runs the test. If the team can argue about whether the criterion was met, it will.
  4. Give the decision to someone outside the team. The owner presents results; the decider, usually the person who funds the work, makes the call against the written criteria.
  5. Treat a pivot as a new bet. If the results suggest a different, better question, that's valuable, but it gets a new brief with new criteria. Otherwise "pivot" becomes a way to never stop.

Run the check-ins as decisions, not updates

Put the check-in dates in the calendar when the bet is approved. Each one is a short meeting with one of four outcomes:

  • Stop. A kill criterion was hit. The team writes a closing note and moves on.
  • Continue. The criteria for this checkpoint were met; the next checkpoint stands.
  • Go. The hardest unknown is resolved; the work moves to a funded build with a normal delivery plan.
  • New bet. The original question is answered, but a different one is worth asking. Write a new brief.

"Continue, but let's give it a few more weeks" is not on the list. If the team needs more time, it's a new bet with a new stake, and the decider should ask whether they'd fund it from scratch today, knowing what they now know. That question cuts through sunk cost better than any other.

Make stopping cheap for the people involved

None of this works if stopping hurts the people who stop. Amy Edmondson, in Right Kind of Wrong, calls failures intelligent when they happen in new territory, in pursuit of a real goal, informed by available knowledge, and are as small as possible. A well-run bet that hits its kill criterion at week six is exactly that. Treat it that way:

  • Require a closing note for every stopped bet: the question, what was tried, the result, and what would have to change for the idea to be worth another look. File it where the next person with the same idea will find it.
  • Present stopped bets in the same forum as successful ones. "We spent twelve engineer-weeks and learned this won't work on our hardware, which saves the product team a quarter" is a real result.
  • Keep bet outcomes out of performance ratings. Judge people on how well they framed and ran the bet, not on whether the answer was yes.
  • Move people quickly. Line up where the team goes before the check-in, so nobody argues to continue a bet because it's the only thing keeping them busy.

Look at the portfolio, not just the bets

James March's 1991 paper on exploration and exploitation argued that organizations drift toward exploitation because its returns are quicker and more certain. Bets with kill criteria let you protect exploration without it becoming a black hole. Once a quarter, list every open bet with its stake, its next checkpoint and its status. Two numbers are worth tracking:

  • How many bets you stopped. If it's zero over a year, either every idea you pick is safe, which means you aren't exploring, or your kill criteria are too soft to trigger.
  • How long bets ran before the first real test of the hardest unknown. If it's usually more than a few weeks, your teams are still building pedestals first.

Your first step

List every exploratory project in your area that has been running for more than a quarter. For each, ask one question: what measurable result, by what date, would make us stop? If nobody can answer, that project gets a bet brief this month, with the hardest unknown first and a named decider. Framing, running and ending experiments like this is a core part of the innovation leadership area in the PTMS study guide.

Working toward PTMS Specialist?

The free study guide covers every competency area on the PTMS exam, with practical examples.

More from the blog

Practical articles for technical managers at every stage. All posts

  1. A one-page technical strategy your executives will read

    Most engineering strategies are lists of goals. How to write one that makes real choices, fits on a page and gets a decision from the people who fund it.

  2. One-on-ones that engineers don't dread

    A practical setup for one-on-ones: who owns the agenda, a running-doc template, a question bank for when things go quiet, and the habits that ruin them.

  3. Getting another team to deliver what your team depends on

    How to find cross-team dependencies early, ask for them in a way the other team can say yes to, track them weekly and escalate without burning bridges.