Optimizely AI and Agentic Systems

The interesting part of Opal isn’t the agent. It’s the system around it.

I started with Opal the same way I start with any platform: what work should it actually own? That question took me from single agents to multi-agent workflows, customer data, approval gates, MCP connections, personalized pages and a Virtual Teammate.

Every build below says what it was. A proof of concept is a proof of concept.

Start with the job
  1. Inputs

    • Context
    • Tools
    • Skills
  2. Specialized agent

    One narrow job, done the same way every time

  3. Workflow

    • Ideate
    • Build
    • Review
  4. Human gateA person decides what launches, what scales and what touches a customer
    • ActionThe approved change goes out
    • MeasurementDid it change the decision?
    • LearningThe result shapes the next round

Learning becomes context for the next round

Principle

Give the agent a job.

A common first design is one giant agent asked to do everything. It costs more, drifts more, and when it’s wrong nobody can tell which part was wrong.

What I avoid

One agent, every job

  1. One general agentAsked to do everything in a single prompt
  2. High costEvery step pays for the biggest context
  3. InconsistentOutput drifts from run to run
  4. Hard to checkNobody can say which step went wrong

Big context, big bill, hard to trust.

What I build

Narrow agents and a person

  1. Agent 1: pull the dataRead only what the task needs
  2. Agent 2: check the rulesStandards, brand, accessibility
  3. Agent 3: draft the recommendationOne output, one format
  4. Person approvesThe decision stays human
  5. Agent 4: actOnly after approval

Each step can be tested, measured and approved on its own.

Experimentation

The human stays in. Just at a different step.

I don’t remove the human from experimentation. I move the human to the decision that matters.

  1. WorkflowRuns the steps in order and passes context forward
  2. IdeationTurns evidence into test ideas
    In
    Analytics and past results
    Out
    Scored ideas
  3. BuildDrafts the variation
    In
    The chosen idea
    Out
    A variation ready to review
  4. ReviewChecks setup against the plan: PASS or FAIL
    In
    Variation, audience, metrics, compliance skill
    Out
    PASS opens a task. FAIL goes back to build.
  5. CMP taskA PASS creates a task for a person
    In
    Review result
    Out
    An approval request
  6. Human approvalNothing launches until a person says yes
    In
    The task
    Out
    Go or no-go
  7. Results summaryWrites the readout in plain language
    In
    Test results
    Out
    A summary and a recommendation
  8. BacklogFeeds what we learned into the next ideas
    In
    Readouts
    Out
    An updated backlog
  9. Program overviewShows the whole program’s health
    In
    Backlog and results
    Out
    A program view for the team
Open any step to see what goes in and what comes out.
The problem
Most experimentation work is repetitive: turning evidence into ideas, building variations, checking setup, writing readouts. The decisions in between are few and they matter a lot.
What I built
In a partner co-innovation proof of concept I built a chain of narrow Opal agents for experimentation, with a CMP task as the human approval before anything launched.
Maturity
A working proof of concept inside a partner environment. It never ran for a client.
Why it’s built this way
Each agent can be tested on its own, and the review step can stop the chain with a FAIL before anything reaches a person.

What I did

  • Split the work into narrow agents: ideation, build, review, results, backlog and program overview
  • Gave review a PASS or FAIL, so the chain could stop itself
  • Made a PASS open a CMP task, so a person approves before anything launches
  • Shaped behavior with standards and compliance skills instead of longer prompts
  • Fed results back into the backlog so the next round starts smarter

Human authority

What agents do, and what people keep.

Agents can

  • Retrieve
  • Summarize
  • Compare
  • Draft
  • Check
  • Recommend
  • Route
  • Monitor

People decide

  • What launches
  • What counts as evidence
  • What promise is made to a customer
  • What becomes personalized
  • What scales
  • Anything with compliance implications
  • Anything that writes to a customer-facing system

The decision model

Not every step in an agentic system is the same kind of decision.

I sort every step into one of three buckets before I build it. Mixing them is how a team ends up trusting a guess the way it trusts a rule, or routing a compliance call through a step nobody designed to hold it.

  • Deterministic rule

    Same conditions, same answer, every time. No model output decides it. The review agent above checks a variation against the written plan and returns PASS or FAIL; nothing about that check is a guess.

  • Agent judgment

    Context has to be interpreted before there’s an answer. The ideation agent turns analytics and past results into ranked test ideas. It can be wrong, so its output is a proposal, not a launch.

  • Human authority

    The decision touches the brand, a compliance obligation, a customer or the business, and a model’s confidence doesn’t change who owns the call. The CMP task above exists so that decision has a named person attached to it.

One AI opinion is not a governance model.

  1. A page of contentDrafted by a person on the content team
  2. In parallel:
    • Brand voice reviewDoes it sound like the organization?
    • Accessibility reviewWCAG 2.1 AA checks
    • Answer-engine reviewCan AI search understand and cite it?
  3. ConsolidatorOne list of issues, in priority order
  4. Person approvesThe content owner decides what changes
The problem
Authors were publishing without quality checks, and the team had no way to check accessibility at scale.
What I built
For a children’s health system’s content team I built and tested review agents for brand voice, accessibility and answer-engine readiness, with a person approving the result.
Boundaries
No clinical guidance from AI, and no patient data in scope. The content owner decides.
Maturity
Built and tested in a sandbox. Not running for the client.
Why separate reviewers
Each check can be evaluated on its own. When one reviewer is wrong, it’s clear which one to fix.

Strategy

The problem wasn’t another dashboard. It was turning the dashboards into a decision.

Audience data, content, search readiness, test history and journey evidence all lived in different places. Every strategy conversation started from zero.

OSA: a strategy assistant made of seven Opal agents

  1. Audience suggesterProposes audiences from customer data
    In
    Customer data profiles
    Out
    Suggested segments, confirmed by a person
  2. Content reviewChecks existing content against the goals
    In
    CMS content
    Out
    Content gaps
  3. GEO auditLooks at readiness for AI search
    In
    Site content
    Out
    Search findings
  4. Experiment blueprinterTurns findings into test plans
    In
    Gaps and findings
    Out
    Test plans, parameters confirmed by a person
  5. Personalization ideasMatches audiences to experiences
    In
    Audiences and content
    Out
    Personalization ideas
  6. Customer journeyPlaces each idea on the journey
    In
    Ideas and journey stages
    Out
    A journey view
  7. RoadmapSequences everything into a plan
    In
    All of the above
    Out
    A prioritized roadmap
What I built
I built a working prototype that chained seven Opal agents into one strategy workflow. I wrote the instructions it runs on (Optimizely now calls these skills): brand, KPIs, personas and data governance. I directed the companion app with an AI coding assistant.
Human checks
A person confirms audiences before they’re created and confirms test parameters before a plan is set. A person also chooses where results go.
Maturity
Built as a working prototype to explore how several Optimizely capabilities could feed one strategy workflow. Some connections used sample data, so it wasn’t run end to end in production.
The problem it solves
Customers have useful evidence spread across several systems and still struggle to turn it into the next decision. This put it in one sequence: audiences, content, search, tests, personalization, journey, then a prioritized roadmap.

Use Opal for what only Opal needs to do.

  1. Experimentation MCPPull results straight from the platform
  2. Data-pull agentFetch once
  3. Cache the resultsNever pay twice for the same numbers
  4. Significance and metric checkAgainst a written statistical standard
  5. SummaryPlain-language readout
  6. OpalOnly where an Optimizely-specific action is needed
The problem
One day of post-test analysis used roughly 500 to 600 Opal credits.
What I recommended
Move the repeatable steps out of Opal. Pull the data once, cache it, check significance against a written standard, and only call Opal where an Optimizely-specific action is needed.
What happened
The team moved its repeatable analysis off Opal.
Why it matters
Using less of the platform was the better answer here. Opal stayed on the steps that need it, and the cost became predictable.

The safest automation is the one that knows what it isn’t allowed to do.

  1. My applicationAssessment answers and recommendations
  2. Read-only MCP gatewayFive tools, read access only
  3. OpalCalls the tools like any other
  4. Concierge and Virtual TeammateAnswers and guidance from what it read
  5. Write actions: lockedHeld behind consent, approval, audit and a kill switch. Not switched on.
What I built
I built a read-only MCP gateway with five tools and verified it inside Opal. I configured a Virtual Teammate in a partner environment, connected it to a read-only system, and kept write actions switched off.
The locked lane
I designed a write path that would need consent, approval, an audit trail and a kill switch. I left it off on purpose.
What building it showed
The read path shipped. The write path stays designed and switched off until consent, approval and an audit trail are real.

One-to-one pages

I learned the autonomous pattern, then built the decision boundary myself.

As a partner I was walked through Optimizely’s own method for generating one-to-one pages. That product is theirs. Separately, I built my own version with rules instead of a model, so every page could be traced back to a decision.

Optimizely’s agentic pattern

An agent builds the page

  1. Customer data and audience
  2. Research
  3. Page-generation agent
  4. CMS
  5. Governance and human review
  6. Distribution

Optimizely describes Limitless 1:1 Personalization as starting with human review and allowing more autonomy as quality is proven.

My prototype

Rules build the page

  1. Assessment answers
  2. Rule-based recommendationNo model writes the copy
  3. CMS (SaaS) content API
  4. Read back through GraphVerify nothing was lost
  5. A page for one person

I built a prototype that wrote a personalized page for each visitor into Optimizely CMS (SaaS) from their assessment answers, then verified it through Graph. It was verified in production, then retired as the live path.

What building it showed: generating the page was the easy part. The work was everything around it:

  • Source data
  • Decision authority
  • Content limits
  • Approval
  • Measurement
  • Persistence
  • Update rules

Virtual Teammates

A teammate is a role, not a chat window.

Optimizely launched Virtual Teammates in 2026. I configured one in a partner environment, connected it to a read-only system, and designed the jobs I’d hand it next. That’s exploration, not a program.

Chat

One request, one answer

  1. You ask
  2. It answers
  3. It forgets the job

Useful, but you are the one remembering the work.

Virtual Teammate

A standing role

  1. Its own identity and accessScoped permissions
  2. MemoryPicks up where it left off
  3. Tools and jobsWorks on a schedule or a trigger
  4. Proposes, then waitsA person approves by default
  5. FeedbackGets better at the role

Optimizely describes these as coworkers with their own login, memory and scoped access.

The question isn’t what the teammate can do. It’s what recurring responsibility I can safely hand to it.

What I’d build next

Worth building. Not yet built.

These are my ideas, not Optimizely’s roadmap. Each one keeps a person on the decisions that matter.

  • CRO manager teammate

    Watches the experiment program so the team doesn’t have to chase it.

    Monitors
    Backlog, analytics, behavior evidence, results, program health
    Does
    Flags issues, prepares readouts, proposes the next decision

    A person owns launch, interpretation and scale.

  • Customer value architect

    Finds the value a customer already paid for and isn’t using.

    Reads
    Licensed products, connected tools, usage, goals, available data
    Outputs
    Next use case, missing integration, enablement gaps

    A person decides what to recommend to the customer.

  • Experiment learning network

    Turns past tests into better next tests.

    Flow
    Past results, patterns, new hypothesis, pre-analysis, QA, post-analysis
    Keeps
    A reusable record of what worked and why

    A person decides which learning becomes a rule.

  • Personalization readiness agent

    Says no until a segment has evidence for a different experience.

    Asks
    Do we know who this is? Is there a signal? Does it predict a different need? Do we have a better experience to show? Can we measure it?
    If not
    Don’t personalize yet

    A person decides when the evidence is good enough.

  • Content governance workflow

    Moves every page through the same checks before and after publishing.

    Steps
    Create, accessibility, brand, search and answer-engine readiness, compliance
    After
    Publish, then keep monitoring

    A person approves before anything goes live.

Why this matters now

Three kinds of experience now meet in one platform.

The way I see it, the platform now serves three audiences at once: the marketers who run it, the customers who see it, and the AI agents that read and act on it. Here is the work I’ve done for each.

  • Marketer experience

    Enablement, workflows and approval steps, so a team can run the platform on their own.

  • Customer experience

    Experimentation and personalization, where the business decision actually shows up on a page.

  • Agent experience

    Opal agents, MCP connections and a Virtual Teammate, with clear limits on what each one can touch.

I’ve spent most of my career asking what makes someone take the next step. Optimizely gave me better ways to test the answer. Opal is giving us better ways to act on what we’ve learned.

The part I still care about hasn’t changed: the system should help the team make a better next decision.

If any of these stories are ones you’d want to talk through, that’s the conversation I enjoy most.

Optimizely and related product names are trademarks of their owner. This section reflects my own professional experience with the platform. It is not affiliated with or endorsed by Optimizely, and client examples are described by industry, not by name.