The first experiment isn’t the finish line.
The point of onboarding isn’t getting a platform configured. It’s getting a team to the point where they can make good decisions without needing someone else in the room every time. By 2022 I owned our packaged Optimizely Web onboarding offer and led program enablement for Web and Full Stack customers, so this is the system I actually run, not a theory of one.
Partner leadsCustomer team leads
- Align
- Implement
- Trust
- Train
- Launch
- Learn
- Self-sufficient
Co-delivery
How partner onboarding actually worked.
In 2019 Optimizely added partner-led onboarding next to its own team. Here is what one engagement looked like from the inside.
- 2019Program launchI was on the small launch team
- 27Engagements in the program’s first two yearsMostly Web Experimentation
- 20+Of those 27, named as a resourceHands-on delivery on a subset
- 50+Optimizely projects across my wider teamAll programs, not only co-delivery; a team figure
- KickoffGoals, owners, timeline
- Platform trainingThe customer’s own site, not a sandbox
- IdeationIdeas tied to a metric the business cares about
- Experiment design and roadmapHypothesis, audience, primary metric, priority
- Stats and reading resultsWhat a winner does and doesn’t mean
- Program managementIntake, cadence, who approves
The engagements were mostly Web Experimentation, with some Personalization and Full Stack. Optimizely brought customer success, training and the platform; the partner side brought scoping, discovery, metrics, ideation, governance and experiment design. That’s why kickoff starts with goals and owners, and platform training comes after: configuration is quick, deciding who owns the program isn’t.
The launch and my part in it are the partnership chapter on the Story page; what follows here is the operating sequence itself.
The system
Seven stages, each with something to show for it.
Every stage ends with an output the team keeps. If a stage can’t name its output, it isn’t finished.
01Discover
- Input
- Business goals, current reporting, who owns the program
- Action
- Agree what success means before any test is picked
- Output
- Business goals and ownership named
02Configure
- Input
- Goals, teams, measurement, capacity
- Action
- Set the primary metric and guardrails
- Output
- A KPI and measurement plan
03Integrate
- Input
- The site, the analytics tool, events and identity
- Action
- Install, connect analytics, run an A/A test and a cross-check, set audience and data requirements
- Output
- Signals the team trusts, and what the audience data needs to hold
04Launch
- Input
- A real question the business is asking
- Action
- Pick one experiment, campaign, flag or audience; set the experimentation taxonomy and the QA gate it has to pass
- Output
- The first live test, launched through a QA gate
05Enable
- Input
- The people who will run it
- Action
- Hands-on training on their own site, a RACI, test intake and prioritization, reporting, templates and office hours
- Output
- Team capability
06Govern
- Input
- A team running its first experiments
- Action
- Step back; review their decisions instead of making them, on an agreed cadence
- Output
- Decisions made without us, on a stated cadence
07Scale
- Input
- A working program
- Action
- Add teams, journeys, products, customer data and AI, one at a time
- Output
- An expansion plan; the next team starts faster
Scale sends a new team back through Discover
Discover comes first so nobody picks a first test before the team agrees what success means. A test that wins on the wrong metric teaches the team the wrong lesson.
Integrate comes before Launch on purpose. An A/A test and an analytics cross-check are cheap. Discovering six months in that the numbers don’t match costs the program.
What I leave behind
The templates outlast the engagement.
Clean, generic versions of the working documents a team keeps using after I’m gone. Recreated from my method, never from a client’s files.
Success metrics contract
The one metric that decides, and the guardrails that can veto it.
Experiment brief
Hypothesis, audience, change and expected effect on one page.
Measurement plan
Where each number comes from and how it’s cross-checked.
QA checklist
What gets checked before a test goes live, by whom.
Audience rules
When a segment has earned its own experience, and when it hasn’t.
Program intake
How ideas come in and get scored.
Learning record
What each result taught us, written for the next team.
Roadmap
The next quarter of tests and campaigns, in order.
Enablement guide
How a new team member gets from access to a first real task.
Two situations, two starting points
A new program and a maturity rescue need different first moves.
A is a team starting from purchase and setup, getting to a first result it can trust. B is a team that already owns Optimizely but isn’t getting value from it: weak decision rules, the wrong metrics, or unclear ownership of the program.
One platform. Five product teams. One set of rules.
- Team 1 to Team 5+Product teams with different roadmaps and release cycles
- Shared experimentation standardWhen to use Web or Full Stack, how legal signs off, how results are read
- Web, app phase, feature flags, measurementOne set of rules across every surface
- The problem
- Build and scale experimentation across several sites and product teams without the teams colliding.
- My role
- I led the experimentation strategy. Optimizely handled the SDK implementation.
- What changed
- Optimizely Web and Full Stack running across several sites and five-plus product teams, with the app scoped for a later phase.
- The shared standard
- The same rules for every team: when to test on the web layer, when to use Full Stack, and when legal signs off.
What I did
- Defined when to use Web versus Full Stack
- Put legal sign-off on UI changes into the test workflow
- Set shared rules across product teams
- Ran the onboarding and roadmap tracks
Not every concluded test was saying the same thing.
- About 100 concluded testsSome from an older tool, some on Optimizely
- Not all equalWinners sometimes called on a downstream metric when the primary was flat
- Primary metric, guardrails, activationMeasure the change the test makes; keep inquiries and applications as guardrails
- Decision ruleWritten down, so the next result is read the same way
- Leadership adoptionEnrollment leadership adopted the rules
- The problem
- Roughly a hundred concluded tests, but winners were sometimes called on downstream metrics when the primary metric was flat.
- My role
- I led the 2025 onboarding series. In 2026 I reviewed the program and rewrote the decision rules. The university’s own team ran the tests.
- What changed
- Enrollment leadership adopted the decision rules.
- What that set up
- Every future result gets read against the same written rule, before anyone argues about what it means.
- Same program, AI side
- The same team’s Opal cost question is on the AI and agentic systems page.
What I did
- Reviewed the concluded tests, including an older-tool archive
- Tied the primary metric to the change being tested
- Scoped activation to people who could actually see the change
- Kept inquiries and applications as guardrails
- Used the program’s own data to show how an early lift can fade
I measure onboarding by the first decision the team makes without me.
Licenses get renewed when teams can show what changed. That only happens when the program has an owner, a metric everyone agrees on, and a habit of writing down what each result taught them.
Once the team can run the system, the question becomes whether it changes the business.
Optimizely and related product names are trademarks of their owner. This section reflects my own professional experience with the platform. It is not affiliated with or endorsed by Optimizely, and client examples are described by industry, not by name.