Adasight
Playbook · Part 00

The Metric Bucket Framework

How to structure experiment metrics so a "win" doesn't quietly cannibalize another part of your product. A practical framework for primary, secondary, guardrail, and learning metrics — plus the specific technique for catching cross-vertical cannibalization before it ships.

From the Adasight × Cherto webinar Q&A Dr. Simon Jackson, Founder of Cherto
What this playbook helps you do
Structure metrics so a "win" can't hide a loss somewhere else
Catch cross-vertical cannibalization before it ships
Set a ship / no-ship rule your whole team can align on
Know what to do when a clean A/B test isn't possible
01
Part One

Most teams lump everything into "primary."

That's the trap. The more product verticals and metrics you have, the more risk there is of false positives — and of a "win" on one metric quietly stealing from another. The fix isn't a fancier statistical test. It's discipline about which bucket each metric belongs in.

"I would always have primary, secondary, guardrail, and exploring or learning metrics as a minimum set of buckets. And then you pick the metrics in those buckets for your use case."

Dr. Simon Jackson · Cherto
Primary
Definition
One metric — the decision-making metric for this test.
How to choose it
Ask: "what one number would make leadership say yes or no to this test?" Use a shared business metric if you have the traffic to move it within your test window. If not, drop to the closest metric that does have volume — and treat the real business metric as a secondary instead of forcing it to be primary.
e.g. checkout conversion rate
Secondary
Definition
Behavioral metrics close to the specific change you made.
How to choose it
Pick 1–3 metrics that should move first, before your primary metric, if the hypothesis is correct. These give you an early read, and help explain why a test won or lost even before you hit significance on the primary.
e.g. add-to-cart rate, scroll depth on the changed section
Guardrail
Definition
The metrics whose job is to protect you from cannibalization.
How to choose it
List every other team, vertical, or surface that shares your users or revenue pool. Add each one's primary metric as a guardrail on your test. If your change touches shared infrastructure, add core trust metrics too — error rate, page load time, unsubscribe rate. Part 02 below covers exactly how to apply this.
e.g. another vertical's conversion rate, page load time
Learning
Definition
Exploratory metrics you're watching out of curiosity.
How to choose it
Pick metrics tied to a future hypothesis, not this decision. Never let a learning metric block a ship call — if it starts to, that's a signal to promote it to secondary or guardrail on your next test.
e.g. scroll depth, feature discovery rate
02
Part Two

The specific technique for catching cannibalization.

This is the part most frameworks skip. It's not enough to just have a guardrail bucket — you have to know what to put in it.

The technique
Put other verticals' primary metrics into your guardrails.
If you're testing on one product vertical and the business has others, take their primary metrics and add them as your guardrails. A "win" that's actually just customers moving from Vertical B to Vertical A gets caught immediately — instead of surfacing three months later as an unexplained dip someone else has to investigate.
1
List every vertical or surface that shares users, traffic, or revenue with the one you're testing.
2
Pull each one's primary metric — the same number their own team would use to judge a win.
3
Add them as guardrails on every test that touches a shared surface, not just the obvious ones.
4
Review the guardrail list quarterly — new verticals, new surfaces, and org changes all shift what belongs here.
Worked Example

Seeing it play out on a real test.

Picture a growth team running three product lines under one roof — Ecommerce, Subscriptions, and Marketplace — all drawing from the same customer base. This quarter, they're testing a limited-time promotion banner on the Ecommerce homepage, aimed squarely at lifting checkout conversion.

Before launch, here's how the metrics got bucketed — and why:

Ecommerce conversion rateThe number leadership actually cares about for this test.
Primary
Ecommerce add-to-cart rateShould move first if the promotion is working as intended.
Secondary
Subscriptions conversion rateShares the same customer base — the promotion could pull people away from subscribing instead.
Guardrail
Marketplace conversion rateSame logic — a separate revenue line drawing on the same traffic.
Guardrail

Two weeks in, the results come back:

MetricBucketResult
Ecommerce conversion ratePrimary+4.2%
Ecommerce add-to-cart rateSecondary+2.8%
Subscriptions conversion rateGuardrail−3.1%
Marketplace conversion rateGuardrail±0.2%

What this revealed: without the guardrail bucket, this looks like a clean win — ship it, move on. With it, the picture changes: the Ecommerce lift lines up almost exactly with the Subscriptions dip. This isn't new revenue. It's very likely the same customers, redirected.

What the team did next: instead of shipping outright, they held the test at partial rollout and pulled in a lightweight follow-up — comparing the combined Ecommerce + Subscriptions revenue before and after, rather than either metric alone. That combined view is what actually determined whether the promotion created new value or just moved it around. (This is exactly the kind of triangulation Part 04 below covers.)

03
Part Three

The decision rule this enables.

Once metrics are correctly bucketed, the ship decision stops being a judgment call and becomes a rule your whole team can align on in advance.

Ship if primary is positive and no guardrails are negative.

This doesn't remove the harder statistical questions — whether to run a different test per bucket, whether to use non-inferiority testing on guardrails — but having this hygiene in place, and aligning with other parts of the business on what the guardrails should be, already puts you ahead of where most experimentation programs sit today.

04
Part Four · Bonus
When A/B isn't clean

When you can't run a clean A/B test.

Not every question has enough traffic, or a clean enough setup, for a proper randomized test. In those cases, the honest answer is to triangulate — synthesizing across several weaker signals instead of relying on one strong one.

Lightweight A/B
A smaller or shorter test than you'd normally run, treated as a signal rather than proof.
Time logs
Before/after comparisons of how long a workflow or task actually takes.
Analytics
Behavioral patterns from existing usage data, read for directional signal.
Qual analysis
Interviews, session replays, or support tickets that explain the "why" behind the numbers.

There are statistical techniques to boost power in these situations — but unless there's a serious need, they add complexity without much practical benefit. Straightforward synthesis across data points is usually the right call.

Adasight x Cherto webinar: 5 Lessons Learned from 1000s of Experiments with Dr. Simon Jackson
Watch The Full Session On YouTube

Watch 5 Lessons Learned From 1000s of Experiments

Gregor Spielmann (Adasight) × Dr. Simon Jackson (Cherto)

The full webinar covers four more lessons from thousands of experiments — including why "more tests" isn't the same as more learning, and what actually breaks when programs try to scale.

Watch on YouTube →
Next Step

Bucketing your metrics is one piece. Here's how to see the whole picture.

Adasight's Experimentation Gap Analysis is a structured look across the four places programs typically get stuck — with a clear roadmap for closing what's holding yours back.

Data
Can your team even get to the data needed to form a hypothesis?
Insights
Do results turn into a clear "why" — or end in a dashboard?
Experimentation Practice
Is there a clear success metric and process behind every test?
AI
Is your experimentation data structured enough to actually feed AI?
See the four gaps →

What the audit delivers:

A current-state assessment across process, tooling, and skills.

The specific gaps keeping tests from producing trustworthy results.

A prioritized roadmap ranked by impact and effort — not a generic checklist.

30 minutes. No pitch — just a clear view of your biggest opportunities.

Prefer To Talk It Through?
Gregor Spielmann
Gregor Spielmann
Co-Founder & COO, Adasight

Ex-Amplitude, ex-Optimizely. Helps growth and product teams build experimentation programs that compound — not just run tests.

Book a 30-min call →

Sourced from the Adasight × Cherto webinar Q&A with Dr. Simon Jackson, July 2026.