Testing and validation

How to set fail criteria for business experiments (before you start)

Fail criteria are the most important and most skipped step in business experimentation. Without them, every result becomes 'close enough.' Here is how to set fail criteria before you run your experiment, using three methods that work in B2B and industrial settings.

Ton van der Linden·How-to·Last updated 28 March 2026·10 min read
How to set fail criteria for business experiments (before you start)

Last month, I watched a team present the results of a six-week experiment. They had tested whether B2B buyers in the food processing industry would pay for a new predictive maintenance service. The experiment was a landing page with a “request demo” button. Their target: 30% of visitors click the button.

The result: 12%.

The team’s conclusion? “The results are promising. We need to optimize the landing page and run it again.”

No. Twelve percent against a thirty percent target is not promising. It is a fail. But because the team had not formally agreed on fail criteria for their business experiment before running it, nobody could say that out loud. And so the idea lived on, absorbing budget and attention for another quarter.

This is the most common and most costly mistake I see in testing business ideas. Not picking the wrong experiment. Not testing the wrong assumption. Skipping the step where you define what “fail” looks like before you start.

Request a Strategy Call about Testing your Business Ideas

In 30 minutes, we'll identify your riskiest assumption and the fastest experiment to test it. Or book a workshop where your team designs and launches real experiments in one day, with fail criteria set before you start.

Why fail criteria, not success criteria

You might wonder why I say “fail criteria” instead of “success criteria.” The difference is not semantic. It is psychological.

When you set success criteria, you define what good looks like. “We need 30% of visitors to click.” Sounds clear. But when the result comes in at 28%, the conversation shifts: “28% is close enough to 30%. We were almost there. Let us tweak the headline and try again.”

When you set fail criteria, you define what bad looks like. “If fewer than 20% of visitors click, this hypothesis is wrong.” When the result comes in at 12%, there is no room for “close enough.” Twelve is below twenty. The hypothesis failed.

The framing matters because of how teams process disappointing results. Success criteria invite negotiation upward. “We almost made it.” Fail criteria force a binary answer. Did we fail or not?

In 100+ sessions helping teams run experiments, I have never seen a team argue that 12% is close enough to 20%. But I have seen dozens of teams argue that 28% is close enough to 30%. The goalposts move when the criteria are framed as targets to hit rather than floors that must be cleared.

Osterwalder and Bland make this point in Testing Business Ideas: set your call-to-action criteria before you run the experiment. Not after. Not during. Before.

Three methods for setting fail criteria

The hardest part of fail criteria is not the concept. It is the number. What should the threshold actually be?

I use three methods, depending on the situation. They are not mutually exclusive. For important experiments, I recommend using at least two of them and comparing the results.

Method 1: early adopter benchmarks

Early adopters are the most forgiving customers you will ever have. They tolerate rough edges, incomplete features, and clunky processes because they want the solution that badly. If early adopters will not engage with your experiment, the broader market definitely will not.

This means your fail criterion should reflect what engaged, motivated prospects would do. Not what the average customer would do.

Example: A manufacturing company tested a new IoT-based quality monitoring service. They invited 40 prospects who had previously expressed interest in quality improvement to a webinar explaining the concept. The fail criterion: if fewer than 25% of invitees register, there is not enough pull for this idea even among warm prospects.

The logic: these 40 people had already raised their hand once. They already cared about quality improvement. If three-quarters of them could not be bothered to spend 45 minutes hearing about a potential solution, the value proposition is not strong enough.

Result: 11 out of 40 registered (27.5%). Just above the threshold. The team continued to the next experiment but with a note: demand among warm prospects is real but not overwhelming.

Method 2: industry analogs

You are rarely the first company to test something in a given channel. Other products in your industry, or similar industries, have conversion rates you can reference.

ChannelTypical B2B benchmarkSource
Cold email open rate20-25%Industry averages
Landing page conversion (cold traffic)2-5%B2B SaaS benchmarks
Landing page conversion (warm traffic)10-20%Segmented campaign data
Webinar registration (invited list)20-30%Event marketing benchmarks
Demo request (from landing page visitor)5-15%B2B sales benchmarks

Your fail criterion should sit at or below the low end of the relevant benchmark. If the industry average for landing page conversion with cold traffic is 2-5%, and your experiment hits 0.8%, you have a signal. Not about your landing page design. About the strength of the underlying value proposition.

One warning: do not cherry-pick benchmarks that make your results look good. If you are testing with warm traffic, use warm traffic benchmarks. If your experiment targets a niche segment, adjust for smaller sample effects. Intellectual honesty here separates useful experiments from expensive theater.

Method 3: back-of-napkin profitability

This is my favorite method for B2B and industrial experiments because it connects directly to business viability.

Work backwards from the business model. What does the math require?

Example: An industrial services company wanted to test a subscription model for equipment calibration. The annual subscription price was €15.000. Each customer would cost approximately €9.000 to serve. That leaves €6.000 gross margin per customer. The company needed 50 subscribers in year one to justify the investment in the service infrastructure.

The experiment: a targeted email campaign to 500 qualified prospects offering an early-access discount.

The back-of-napkin calculation:

  • Need 50 customers from a reachable market of roughly 2.000 qualified companies
  • That is a 2.5% conversion rate from the total market
  • From a targeted email campaign to 500, you need at least 10 sign-ups (2% of recipients) to be on track for the 50-customer target
  • Fail criterion: fewer than 10 expressions of serious interest (defined as booking a call to discuss terms)

This method is powerful because it removes opinion from the equation. The business model either works at these numbers or it does not. Your opinion about whether the landing page was “good enough” is irrelevant. The math is the math.

How to get team agreement before you run

Setting fail criteria alone at your desk is easy. Getting a cross-functional team to agree on them is the real challenge. Here is the process I use in workshops.

Step 1: write down the hypothesis

Be specific. Not “customers want this product.” Instead: “At least 20% of mid-size food processing companies in the Netherlands will express interest in a predictive maintenance subscription when presented with the value proposition via a targeted landing page.”

The hypothesis should include: who, what behavior, what threshold, and through which channel.

Step 2: propose the fail criterion

I ask the team: “What result would convince you that this idea is not worth pursuing further?” Not what result would make you happy. What result would make you stop.

This question changes the energy in the room. People get specific fast. “If fewer than 10 companies respond, we should kill it.” “If none of the respondents are willing to pay more than €8.000, the margin is not there.”

Step 3: challenge it from both sides

Two questions to pressure-test the proposed criterion:

“If we hit exactly this number, would you genuinely be willing to kill the idea?” If the answer is no, the criterion is too low. Raise it.

“If we miss this number by one, would you genuinely stop?” If the answer is hesitation, the criterion might be too high, or the team is not actually committed to evidence-based decisions. Both are worth discussing before you spend money on the experiment.

Step 4: document and sign off

Write the fail criterion on the experiment canvas. Not in a separate document. Not in an email. On the canvas itself, next to the hypothesis and experiment design. Everyone who will be involved in interpreting results should see it and agree.

In corporate settings, I also recommend getting sign-off from whoever controls the budget. When a VP has agreed that “fewer than 15 responses means we stop,” it is much harder for the team to argue around that number six weeks later.

For more on identifying which assumptions to test first, see the how to test business assumptions guide.

What happens when you skip this step

I have seen the pattern enough times to describe it with precision.

Week 1: The team runs an experiment without fail criteria. “We will see what the data tells us.”

Week 7: Results come in. They are mediocre. Not terrible, not great. The team discusses.

Week 8: Someone says “the results are promising, but the sample size was too small.” Another person suggests the timing was bad. A third person thinks the messaging was wrong but the underlying idea is solid.

Week 9: The team requests budget for a second round. “We learned a lot from round one. Round two will be different.”

Week 15: Round two results are slightly better but still mediocre. The same conversation happens again. More budget requested.

Week 24: The idea has consumed six months and a significant budget without ever producing strong evidence. But nobody can kill it because there was never a defined point at which “fail” meant fail.

This is pilot purgatory. And the root cause is almost always the same: no fail criteria.

The cost is not just the money spent on the experiment. It is the opportunity cost. Every month that a mediocre idea stays alive is a month that a potentially stronger idea does not get tested. In organizations that manage innovation portfolios, this is how the portfolio fills up with zombie projects, ideas that are not alive enough to succeed but not dead enough to kill.

Understanding innovation accounting helps teams track whether experiments are producing real learning or just activity.

The “close enough” problem

The most dangerous number is the one that is almost good enough.

If your fail criterion is 30% and the result is 5%, nobody argues. The idea failed. Move on. If the result is 80%, nobody argues either. Clear win.

But if the result is 27%? Or 24%? That is where teams start negotiating.

“The email went out on a Friday. Open rates are always lower on Fridays.” “Three of the non-responders were on vacation.” “Our landing page had a technical issue for the first two hours.”

Every one of these might be true. But if you allow post-hoc explanations to override pre-set fail criteria, you do not have fail criteria. You have suggestions.

My rule: if the result is below the fail criterion, the hypothesis failed. Period. The team can choose to run a new experiment with a revised hypothesis. They can choose to test a different segment or channel. But they cannot retroactively adjust the criterion to fit the result.

This sounds rigid. It is. That is the point. The entire value of fail criteria is that they cannot be negotiated after the fact. Remove that constraint and you are back to opinion-based decision making, which is exactly what testing business ideas is designed to replace.

When an experiment fails, the pivot or persevere decision framework helps teams decide what to do next.

Choosing the right experiment type from the experiment library increases the quality of evidence you collect.

A checklist before you run any experiment

Before launching an experiment, confirm these six points:

  1. The hypothesis is written down, specific, and testable
  2. The fail criterion is a number, not a feeling
  3. The fail criterion was set using at least one of the three methods (early adopter benchmarks, industry analogs, back-of-napkin profitability)
  4. The team has agreed on the fail criterion before seeing any results
  5. The person who controls the budget has signed off on the threshold
  6. The fail criterion is documented on the experiment canvas, visible to everyone involved

If any of these six points is missing, you are not ready to run the experiment. Fix it first. Running an experiment without fail criteria is like mapping your business model without talking to customers. It feels productive but produces nothing you can act on.

The time you spend agreeing on fail criteria before the experiment will save you weeks or months of debate after. In manufacturing and industrial settings where experiment budgets are significant and stakeholder alignment is hard-won, this step is not optional. It is the foundation everything else rests on.

Frequently asked questions

What are fail criteria in business experiments?

Fail criteria are pre-defined thresholds that tell you when an experiment has failed. You set them before running the experiment. For example: if fewer than 15% of prospects click the pre-order button, this hypothesis is wrong. The key word is "before." Setting the threshold after seeing results lets teams rationalize any outcome as success.

Why use fail criteria instead of success criteria?

Fail criteria work better than success criteria because of how teams interpret results. With success criteria, a team that hits 28% against a 30% target will say "close enough, let us continue." With fail criteria, the same team faces a clearer question: did we fail or not? The psychological framing makes it harder to move the goalposts. You are not asking "did we succeed enough?" You are asking "did we fail?"

How do you calculate fail criteria for a new business idea?

Three methods work well. First, use early adopter benchmarks: set criteria based on what enthusiastic early customers should do if the idea has potential. Second, use industry analogs: look at conversion rates for similar products or channels in your industry. Third, use back-of-napkin profitability: calculate what conversion rate or price point the business model needs to be viable, and set your fail criterion below that minimum.

When should you set fail criteria for an experiment?

Always before running the experiment. This is non-negotiable. The moment you see results, your judgment is compromised. Setting fail criteria after seeing data is not setting criteria at all. It is rationalizing. In practice, the best time to set them is during experiment design, when the team agrees on the hypothesis, the experiment type, the metric, and the fail threshold together.

What happens when teams skip fail criteria?

Teams that skip fail criteria almost always continue with ideas that should have been killed. The pattern is predictable: run an experiment, get mediocre results, argue that the results are "promising," request more budget, run another experiment with slightly different parameters, get mediocre results again, and repeat. This is how companies end up in pilot purgatory, spending months and significant budget on ideas that never had real evidence behind them.

Innovation Insights

Get Innovation Insights in your inbox

One email a month: the newest guides and practitioner notes, from business model design to innovation readiness, plus dates for upcoming masterclasses. Written from what actually happened in the room, at 50+ companies. No daily drip, no sales sequence.