GrowthBook 5.1 Adds Contextual Bandits and an MCP Server
Every experimentation vendor shipped the same idea this month, in three different shapes: let an AI agent drive the platform. Here is what actually arrived, and which parts are worth switching on.
GrowthBook 5.1: contextual bandits, a hosted MCP server, a Slack app
GrowthBook released version 5.1 on 22 September. Three headline additions.
Contextual bandits choose which variation a user sees at the individual level, learning from that user’s context and shifting traffic automatically. GrowthBook positions them for time-sensitive campaigns and tests with many variants. That framing is right and worth repeating: a bandit optimises reward during the test, an A/B test measures effect afterwards. If you need a clean read on what a change did, run the test. If you have a three-week promotion and no interest in the causal estimate, run the bandit.
A hosted MCP server with OAuth lets agents publish flags, create experiments and work through what GrowthBook describes as 30-plus skills. A Slack app, in beta, lets you tag @growthbook to query data and get daily or weekly digests.
Alongside those: funnel metrics for multi-step experiences, a Learnings feature for documenting what an experiment taught you, User Journeys analytics, AI dashboards and a better SQL Explorer. The Learnings feature is the quietly useful one. Most programmes lose more value to forgotten results than to bad statistics, and a searchable record of what has already been tried is the cheapest thing you can add to a CRO programme. Release notes at GrowthBook.
Amplitude’s argument: your agent is not broken, it is guessing
Amplitude published a piece on 23 September under the title “Your agent isn’t broken. It’s guessing,” about configuring context for AI agents. The point generalises past Amplitude’s product: an agent given access to your analytics without a defined metric layer, event dictionary and set of guardrails will produce fluent, confident, wrong answers, and the failure looks like a model problem when it is a data definition problem.
This is the same discipline that makes experimentation work at all. If nobody can say which metric is the primary one before a test starts, no amount of tooling saves the analysis. Our notes on hypothesis frameworks and what makes a good conversion rate cover the definitions worth settling first. The post is on the Amplitude blog.
Optimizely put two agents in Opal
Earlier this month, on 2 September, Optimizely added two Feature Experimentation agents to the Agent Directory in Opal. A Governance agent reads a Feature Experimentation project and produces a read-only HTML governance report. A Feature Flag Implementation agent generates a portable SKILL.md file so a coding agent can implement a flag in your codebase.
The governance one is the more interesting of the pair, because flag debt is a real and boring problem: teams ship flags, the experiment concludes, nobody removes the flag, and eighteen months later the codebase has a hundred permanent branches nobody dares delete. An automated report that lists them is a small thing that saves an argument. Release notes at Optimizely.
What this run of releases means if you are choosing a tool
Three vendors, one direction. Agent access is becoming table stakes rather than a differentiator, which means it is a poor reason to pick a platform. The things that still differ, and still decide whether a programme works, are the statistics engine, whether the tool reads from your warehouse, and how much traffic you actually have. Our comparison of A/B testing tools and the sample size calculator are the better place to start than any changelog.
More from Experimento
related resultsBest A/B Testing Tools in 2026, Compared by What They Actually Do
A practical 2026 comparison of the best A/B testing tools by what each one is actually good at, from Optimizely and VWO to GrowthBook and PostHog.
read result →A/B Testing Tools Compared: Which Platform Fits Your Team
A/B testing tools compared for 2026: VWO, Optimizely, AB Tasty, Convert, GrowthBook, Statsig and PostHog, matched to marketing, CRO and engineering teams.
read result →Data-Driven Design: Let Research Guide Your Product Decisions
A practical guide to data-driven design: how to pair analytics with user research, run honest tests, and make product decisions on evidence, not opinion.
read result →