The Art of CTO Experimentation Framework assesses experimentation maturity, helps design experiments with statistical rigor, and provides infrastructure gap analysis and sample size calculations for engineering teams.
Are we shipping features, or shipping learnings?
An experimentation maturity score and experiments designed with statistical rigour.
About 15 min · Calculator · Free
About this toolWhy it matters, common mistakes, FAQ
Are You Shipping Features — Or Shipping Learnings?
Teams that ship without validating have no way of telling which features moved a metric and which were wasted effort, so the waste never gets counted. Experimentation turns every release into evidence for the next decision.
Most teams either skip experimentation entirely ("we know what customers want") or run experiments without statistical rigor ("it went up, ship it"). Both waste resources — one by building the wrong things, the other by drawing wrong conclusions.
Questions CTOs ask
- How do you build an experimentation culture in engineering?
- Start with infrastructure: feature flags, experiment analysis tools, and guardrail metrics. Then change incentives: celebrate learnings from failed experiments, not just successful launches. Set expectations from leadership that a percentage of features should be validated through experiments. Create a low-friction process to propose and run experiments. The biggest blocker is usually not technical — it is cultural resistance to admitting uncertainty about what customers want.
- What is the minimum sample size for a meaningful A/B test?
- Sample size depends on baseline conversion rate, minimum detectable effect (MDE), significance and power. At 95% significance and 80% power, a 5% baseline needing to detect a 2 percentage-point lift takes roughly 2,200 users per variant; a 50% baseline detecting a 5 point lift takes roughly 1,600. Halving the MDE roughly quadruples the sample, which is why most startups cannot run statistically valid A/B tests on low-traffic features and are better served by qualitative methods or a staged rollout. (These figures come from the same two-proportion z-test this tool uses, so the calculator above will reproduce them — the previous version of this answer quoted numbers roughly twenty times too high and contradicted the tool on its own page.)
Related Reading
LaunchDarkly vs Flagsmith: how CTOs should choose a feature flag platform
LaunchDarkly vs Flagsmith: how CTOs should choose a feature flag platform
guidesEngineering Product Tension Framework: How CTOs Turn Conflict Into Better Shipping and Quality
Engineering Product Tension Framework: How CTOs Turn Conflict Into Better Shipping and Quality
guidesEngineering Experimentation Framework: How CTOs Build A/B Testing Rigor and a Learning Culture
Engineering experimentation framework: Are you shipping features or shipping learnings?