How to structure experiment metrics so a "win" doesn't quietly cannibalize another part of your product. A practical framework for primary, secondary, guardrail, and learning metrics — plus the specific technique for catching cross-vertical cannibalization before it ships.
That's the trap. The more product verticals and metrics you have, the more risk there is of false positives — and of a "win" on one metric quietly stealing from another. The fix isn't a fancier statistical test. It's discipline about which bucket each metric belongs in.
"I would always have primary, secondary, guardrail, and exploring or learning metrics as a minimum set of buckets. And then you pick the metrics in those buckets for your use case."
This is the part most frameworks skip. It's not enough to just have a guardrail bucket — you have to know what to put in it.
Picture a growth team running three product lines under one roof — Ecommerce, Subscriptions, and Marketplace — all drawing from the same customer base. This quarter, they're testing a limited-time promotion banner on the Ecommerce homepage, aimed squarely at lifting checkout conversion.
Before launch, here's how the metrics got bucketed — and why:
Two weeks in, the results come back:
| Metric | Bucket | Result |
|---|---|---|
| Ecommerce conversion rate | Primary | +4.2% |
| Ecommerce add-to-cart rate | Secondary | +2.8% |
| Subscriptions conversion rate | Guardrail | −3.1% |
| Marketplace conversion rate | Guardrail | ±0.2% |
What this revealed: without the guardrail bucket, this looks like a clean win — ship it, move on. With it, the picture changes: the Ecommerce lift lines up almost exactly with the Subscriptions dip. This isn't new revenue. It's very likely the same customers, redirected.
What the team did next: instead of shipping outright, they held the test at partial rollout and pulled in a lightweight follow-up — comparing the combined Ecommerce + Subscriptions revenue before and after, rather than either metric alone. That combined view is what actually determined whether the promotion created new value or just moved it around. (This is exactly the kind of triangulation Part 04 below covers.)
Once metrics are correctly bucketed, the ship decision stops being a judgment call and becomes a rule your whole team can align on in advance.
This doesn't remove the harder statistical questions — whether to run a different test per bucket, whether to use non-inferiority testing on guardrails — but having this hygiene in place, and aligning with other parts of the business on what the guardrails should be, already puts you ahead of where most experimentation programs sit today.
Not every question has enough traffic, or a clean enough setup, for a proper randomized test. In those cases, the honest answer is to triangulate — synthesizing across several weaker signals instead of relying on one strong one.
There are statistical techniques to boost power in these situations — but unless there's a serious need, they add complexity without much practical benefit. Straightforward synthesis across data points is usually the right call.
The full webinar covers four more lessons from thousands of experiments — including why "more tests" isn't the same as more learning, and what actually breaks when programs try to scale.
Watch on YouTube →Adasight's Experimentation Gap Analysis is a structured look across the four places programs typically get stuck — with a clear roadmap for closing what's holding yours back.
What the audit delivers:
A current-state assessment across process, tooling, and skills.
The specific gaps keeping tests from producing trustworthy results.
A prioritized roadmap ranked by impact and effort — not a generic checklist.
30 minutes. No pitch — just a clear view of your biggest opportunities.
Ex-Amplitude, ex-Optimizely. Helps growth and product teams build experimentation programs that compound — not just run tests.
Book a 30-min call →Sourced from the Adasight × Cherto webinar Q&A with Dr. Simon Jackson, July 2026.