Skip to content
Experimento A / B   growth lab

What Is A/B Testing? And When Concept Testing Beats It

By the Experimento team | Updated 2026 | method-checked

A/B testing splits live traffic between two versions of something and measures which one produces more of the outcome you care about. Concept testing asks a sample of your target audience to react to an idea before it exists. Both are ways of replacing an argument with evidence, and teams routinely reach for the wrong one, usually because A/B testing is the famous one and concept testing is not.

The choice is not a matter of taste. It is decided almost entirely by two things: whether the thing exists yet, and whether you have enough traffic for a split test to say anything at all. Most teams who want to A/B test cannot, and do not find out until they have burned six weeks on an inconclusive result.

What A/B testing actually is

You take a page, an email, a flow or a feature. You create a variant with one thing changed. You send visitors randomly to one or the other, at the same time, and you count how many complete the goal in each group.

Three things make it a genuine experiment rather than a comparison:

The split is random and simultaneous. Running version A in March and version B in April is not an A/B test. It is a before-and-after, and any difference could be seasonality, a campaign, a press mention or the weather.

Only one thing changes. If the variant has a new headline, a new image and a new button colour and it wins, you have learned that the bundle won. You have not learned which part did the work, and you cannot carry that lesson anywhere else.

The result is judged against a stated threshold before you start. Otherwise you will stop the test the moment it is winning, which is the most common way a team ends up rolling out a change that does nothing.

The output is a behavioural fact: more people did the thing. That is genuinely valuable and genuinely narrow. As Nielsen Norman Group puts it, A/B testing “will not provide any insights into why these changes occur,” which is why they recommend combining it with qualitative research rather than treating a win as an explanation.

The traffic maths that decides whether you can run one

This is the part most introductions skip, and it is the part that determines whether A/B testing is available to you at all.

The sample size you need depends on your baseline conversion rate and the size of the effect you want to be able to detect. Smaller baseline, or smaller effect, means dramatically more traffic. The relationship is not linear, and it punishes low-converting funnels hard.

A concrete example. A checkout converting at 2%, looking to detect a 10% relative improvement (2% to 2.2%), needs on the order of 200,000 users per variant. That is 400,000 users in total, or roughly 80 days at 5,000 visitors a day. If your page does not see 5,000 visitors a day, that test simply is not runnable in a useful timeframe.

Three levers, and only three:

  • Raise the baseline. Test high up the funnel where more people convert. A 30% click-through rate needs a fraction of the sample a 2% checkout rate needs.
  • Accept a larger minimum detectable effect. You can find a 40% lift much faster than a 5% one. You just cannot detect the 5% one at all, so stop pretending the test failed to find it.
  • Test fewer variants. Splitting traffic four ways quarters the sample in each arm. On constrained traffic, two arms or nothing.

Run your own numbers before you write a single line of variant copy. Our A/B test sample size calculator, minimum detectable effect calculator and test duration calculator between them tell you in about a minute whether the test you have in mind is possible.

If the answer is no, that is not a failure. It is the point at which concept testing becomes the correct tool rather than the consolation prize.

What concept testing is, and what it answers

Concept testing puts an idea in front of a sample of your target audience and measures their reaction, before you have built anything. It is a survey method, not a traffic method, so it works with hundreds of respondents rather than hundreds of thousands of visitors, and it works for products, pricing, propositions, packaging and messaging that do not exist yet.

It answers a different class of question. An A/B test tells you which of two live things performed better with people who already found you. Concept testing tells you whether an idea appeals, to whom, how it compares with alternatives, and what people say is wrong with it. That last part matters: respondents can be asked why, which a split test can never do.

The three designs worth knowing

Monadic. Each respondent sees one concept only, and evaluates it in isolation. This is the cleanest read, because nobody is comparing, so the scores reflect how the concept lands on its own. It needs a separate group of respondents for every concept, which makes it the most expensive design. Use it when the concepts are very different, or complex, or when you need an absolute rather than relative read.

Sequential monadic. Each respondent sees several concepts, one after another, and rates each. Far fewer respondents needed, and it gives you a direct comparison. The cost is bias: order effects favour whatever came first, and people get tired and less discriminating by the fourth concept. Rotate the order and keep the list short. Use it when concepts are similar and budget is finite.

MaxDiff. Respondents repeatedly pick the most and least appealing option from small sets. It is the best method for ranking a long list of features, benefits or messages, because forcing a choice avoids the problem where everything scores 4 out of 5 and nothing is separated. Use it when you have twelve benefit statements and need to know the three worth building the launch around.

Whichever design you use, the metrics that carry weight are purchase or use intent, relevance, uniqueness and believability. Overall “liking” is the least predictive thing you can measure, and the easiest to get a flattering number on.

Which one to use

The honest decision rule, in order:

  1. Does the thing exist and is it live? No, then you cannot A/B test it. Concept test.
  2. Do you have the traffic? Run the sample size numbers on your real baseline. If the test needs more than about six weeks, it is not a test, it is a delay. Concept test the direction instead, ship the better-supported version, and measure it.
  3. Is the question “which wins” or “why”? A/B testing answers the first. Concept testing, user research and session analysis answer the second. If you cannot articulate why the variant should win, you are not ready to test it, you are guessing with extra steps. Our CRO hypothesis framework covers writing one that is falsifiable.
  4. Is the decision reversible and small, or expensive and one-way? Button copy is cheap to change and worth testing live. A pricing model, a brand name or a product line is expensive and slow to reverse, which is exactly the territory concept testing exists for.
Choosing between A/B testing and concept testing Decision panel. If the thing is not built or not live, concept test. If it is live but the required sample size would take more than about six weeks, concept test the direction and ship. If it is live and the traffic supports the sample size, A/B test one change at a time. If the question is why rather than which, use qualitative research alongside either method. Which method, in the order the questions actually arrive Work down. The first "no" decides it. 1. Does it exist and is it live? No: nothing to split traffic between. Concept test the idea 2. Does your traffic reach the sample size? No, or it would take over ~6 weeks: that is a delay, not a test. Concept test, then ship and measure 3. Both yes, and you have a stated hypothesis? Change one thing. Fix the duration in advance. Run whole weeks. A/B test it Asking "why" rather than "which"? Neither method answers it. A split test reports behaviour, not reasons. Pair either one with qualitative research. Sources: Nielsen Norman Group on A/B testing limitations; standard sample size requirements. Chart by Experimento.
Chart by Experimento. The decision is settled by whether the thing is live and whether your traffic reaches the required sample size, in that order.

The two methods work best in sequence rather than in competition. Concept test to narrow six ideas to two. Build the two. A/B test them if the traffic supports it. Then use qualitative research to work out why the winner won, so the lesson transfers to the next thing you build.

The practical landscape since Google Optimize closed

Worth knowing if you are picking a tool now: Google Optimize and Optimize 360 shut down on 30 September 2023. It had been the free on-ramp that most small teams started on, and Google did not build a replacement inside GA4, choosing instead to open its APIs so third-party testing tools could integrate with Google Analytics.

The practical effect is that experimentation now generally means picking a dedicated platform. Commercial options include VWO, Optimizely, AB Tasty and Kameleoon. Product analytics tools such as PostHog now include experiments alongside their analytics, and GrowthBook is an open-source option for teams that would rather host it themselves. We compare the trade-offs in A/B testing tools compared.

Concept testing runs on a different stack entirely, since it is survey research: panel providers and insight platforms rather than a JavaScript snippet on your site. That difference is a feature, because it means it does not depend on your traffic at all.

Where A/B tests go wrong

Even when the traffic supports a test, these four kill more results than bad ideas do.

Peeking and stopping early. Checking daily and stopping the moment significance appears inflates your false positive rate substantially. Set the duration in advance from the sample size calculation and leave it alone.

Stopping mid-week. Behaviour differs by day. Always run whole weeks, and at least two of them, so weekday and weekend traffic are represented in both arms.

Sample ratio mismatch. If your 50/50 split arrives as 53/47, something is broken in the assignment or the tracking, and the result is not trustworthy no matter how good it looks. Check it every time with a sample ratio mismatch calculator.

Treating a flat result as a failure. “No detectable difference” is a real finding. It tells you that lever does not move this metric at this size, which should stop you spending three more sprints on it. Most tests do not win, and a programme where everything wins is a programme that is measuring itself wrong.

Our guide to conversion rate optimisation puts all of this into a programme rather than a series of one-off tests.

Frequently asked questions

What is the difference between A/B testing and concept testing? A/B testing splits live traffic between two existing versions and measures actual behaviour. Concept testing surveys a sample of your target audience about an idea that does not exist yet, and can ask why. A/B testing needs traffic and a live product; concept testing needs neither.

How much traffic do I need for an A/B test? It depends on your baseline conversion rate and the smallest effect you want to detect. A funnel converting at 2% chasing a 10% relative lift needs roughly 200,000 users per variant. A higher-converting step, or a willingness to detect only larger effects, cuts that sharply. Calculate it before you build the variant.

Can I run a concept test instead of an A/B test if my site has low traffic? Yes, and for most low-traffic sites it is the better answer. Concept testing works with hundreds of respondents rather than hundreds of thousands of visitors, so it gives you a defensible read on direction in days rather than an inconclusive split test in months.

What is monadic concept testing? A design where each respondent evaluates one concept only, with no comparison. It gives the cleanest, least biased read on how a concept performs on its own, at the cost of needing a separate respondent group per concept. Sequential monadic shows each respondent several concepts in turn, which is cheaper but introduces order and fatigue bias.

How long should an A/B test run? Long enough to reach the sample size the calculation demands, and always in whole weeks, with a minimum of two. Never stop early because it looks like it is winning: checking repeatedly and stopping at the first significant reading is the fastest way to ship a change that does nothing.

Is A/B testing still worth it now Google Optimize has gone? Yes, for teams with the traffic. Optimize closed in September 2023 and Google did not replace it inside GA4, so you now pick a dedicated platform such as VWO, Optimizely, AB Tasty, PostHog or the open-source GrowthBook. The method did not change, only the free on-ramp.

// the readout

Get the Experimento newsletter

Independent guides and reviews, straight to your inbox. No spam.

9,400+ growth folks no spam, ever

Confidence 95%. Opt out anytime.