Skip to content
Experimento A / B   growth lab

GrowthBook 5.1 Adds Contextual Bandits and an MCP Server

By the Experimento team | Updated 2026 | method-checked
figure_01 News
GrowthBook 5.1 Adds Contextual Bandits and an MCP Server
Graphic by Experimento
Share this chartFacebookWhatsAppX

Every experimentation vendor shipped the same idea this month, in three different shapes: let an AI agent drive the platform. Here is what actually arrived, and which parts are worth switching on.

GrowthBook 5.1: contextual bandits, a hosted MCP server, a Slack app

GrowthBook released version 5.1 on 22 September. Three headline additions.

Contextual bandits choose which variation a user sees at the individual level, learning from that user’s context and shifting traffic automatically. GrowthBook positions them for time-sensitive campaigns and tests with many variants. That framing is right and worth repeating: a bandit optimises reward during the test, an A/B test measures effect afterwards. If you need a clean read on what a change did, run the test. If you have a three-week promotion and no interest in the causal estimate, run the bandit.

A hosted MCP server with OAuth lets agents publish flags, create experiments and work through what GrowthBook describes as 30-plus skills. A Slack app, in beta, lets you tag @growthbook to query data and get daily or weekly digests.

Alongside those: funnel metrics for multi-step experiences, a Learnings feature for documenting what an experiment taught you, User Journeys analytics, AI dashboards and a better SQL Explorer. The Learnings feature is the quietly useful one. Most programmes lose more value to forgotten results than to bad statistics, and a searchable record of what has already been tried is the cheapest thing you can add to a CRO programme. Release notes at GrowthBook.

Amplitude’s argument: your agent is not broken, it is guessing

Amplitude published a piece on 23 September under the title “Your agent isn’t broken. It’s guessing,” about configuring context for AI agents. The point generalises past Amplitude’s product: an agent given access to your analytics without a defined metric layer, event dictionary and set of guardrails will produce fluent, confident, wrong answers, and the failure looks like a model problem when it is a data definition problem.

This is the same discipline that makes experimentation work at all. If nobody can say which metric is the primary one before a test starts, no amount of tooling saves the analysis. Our notes on hypothesis frameworks and what makes a good conversion rate cover the definitions worth settling first. The post is on the Amplitude blog.

Optimizely put two agents in Opal

Earlier this month, on 2 September, Optimizely added two Feature Experimentation agents to the Agent Directory in Opal. A Governance agent reads a Feature Experimentation project and produces a read-only HTML governance report. A Feature Flag Implementation agent generates a portable SKILL.md file so a coding agent can implement a flag in your codebase.

The governance one is the more interesting of the pair, because flag debt is a real and boring problem: teams ship flags, the experiment concludes, nobody removes the flag, and eighteen months later the codebase has a hundred permanent branches nobody dares delete. An automated report that lists them is a small thing that saves an argument. Release notes at Optimizely.

What this run of releases means if you are choosing a tool

Three vendors, one direction. Agent access is becoming table stakes rather than a differentiator, which means it is a poor reason to pick a platform. The things that still differ, and still decide whether a programme works, are the statistics engine, whether the tool reads from your warehouse, and how much traffic you actually have. Our comparison of A/B testing tools and the sample size calculator are the better place to start than any changelog.

// the readout

Get the Experimento newsletter

Independent guides and reviews, straight to your inbox. No spam.

9,400+ growth folks no spam, ever

Confidence 95%. Opt out anytime.