A/B testing, also called split testing, compares two versions of a webpage, email, ad, CTA, or other marketing asset to see which performs better. By changing one variable and keeping others constant, businesses base decisions on performance data rather than opinion.
Trusting a test result is harder than setting it up because early leads often fade, and a click win may not translate into revenue. Testing is therefore a decision discipline for marketing and sales teams, not a website feature.
What Is A/B Testing?
A/B testing is a controlled experiment with two versions of the same asset. The control is the version already in use, and the variation is identical except for one change, such as a new subject line or button label.
Both versions are evaluated against a single primary metric, such as click-through rate. Changing just one variable ensures clear results, making it easy to attribute any performance lift directly to that change. Statistical significance and sample size, two core A/B testing principles, then confirm whether that lift is real or just random fluctuation.
How Does A/B Testing Work?
The testing tool randomly splits traffic or a contact list into two groups. Group A receives the control, Group B receives the variation, and the tool records each group against the primary metric.
Random assignment prevents bias. If only loyal customers saw a variation, it might perform better because of who saw it, not what changed. Randomly splitting users ensures both groups share the same mix of visitors, so any performance difference comes directly from the test variation.
A significance check confirms whether a performance gap is real or driven by temporary data variation. As UCLA’s statistical consulting group notes, power, effect size, sample size, and alpha are mutually dependent variables.
Establishing your target effect size, significance level, and statistical power before launch calculates the exact sample size needed for reliable decision-making.
Why Is A/B Testing Important?
A/B testing eliminates internal bias and protects revenue:
- Prevents costly mistakes: It isolates failing updates early, stopping underperforming changes before they reach your full audience.
- Accelerates revenue growth: Startups using A/B testing achieve 30% to 100% higher performance gains within a year (Management Science, 2022).
What Can You A/B Test?
- Website and landing pages
Headlines, copy, images, form length, CTA wording, and layout are the usual candidates. Test form length early, because every extra field trades lead volume for lead detail.
- Email marketing
You can compare subject lines, preview text, body copy, CTAs, send times, and formats. Email is well-suited to testing because the list is already known, so random splitting is simple.
- Sales outreach
Sales teams can test cold email subject lines, opening messages, call scripts, presentations, and follow-up timing. A rep might compare a question-led opener against a result-led one across matched prospect lists. Sales A/B testing on documents and demos judges the winner on meeting rate, because a meeting is the outcome those assets exist to win.
Across all three channels, effective A/B tests share three defining traits:
- High volume – The asset reaches enough people each week to produce a readable result in a reasonable time.
- One clear metric – The change maps to a single action, such as a click or reply, rather than revenue months later.
- A bold change- The variation differs enough from the control to produce a gap the sample can detect.
How to Run an A/B Test
- Identify the goal
Define what the business wants to improve and pick one primary KPI. For a newsletter, that may be click-through rate; for a sales sequence, it may be meetings booked.
- Form a hypothesis
State what will change, what should happen, and why.
For example: “Changing the CTA from Learn More to Get Pricing will raise click-through rate, because visitors on this page are comparing costs.”
- Choose one variable
Change only the element under test. Keep the audience, timing, offer, and design constant so the results reflect that one change.
- Create the control and variation.
Duplicate the control, then apply only the proposed change to the duplicate. Check that both versions render correctly on mobile and desktop, so you don’t mistake a layout bug for a result.
- Audience and sample size
Sample size for comparing two conversion rates depends on the baseline rate, minimum detectable effect, significance level, and power. Evan Miller’s sample size calculator takes these four inputs and returns the visitors needed per variation.
The required sample size therefore determines whether a test is practical for a given audience. For web experiments, Kohavi and Thomke wrote in Harvard Business Review that “any company that has at least a few thousand daily active users can conduct these tests.”
- Run the test
Run both versions at the same time and fix the test period before launch. In his essay How Not To Run an A/B Test, Miller advises fixing the sample size in advance and trusting the numbers only once the experiment ends.
- Analyze the results
Compare the primary KPI and check significance. Then rule out external factors such as a holiday, and compare new and returning visitors to check that the result holds for both.
A/B Testing Metrics to Track
Evaluate each test using an isolated primary metric directly tied to the change, while tracking a secondary control metric to protect broader business performance.
Primary Metrics (Direct Impact):
- Conversion & Click-Through Rates: Validate immediate user action on landing-page elements and call-to-action buttons.
- Click & Engagement Rates: Best for email copy and send times. According to Litmus, Apple Mail Privacy Protection (MPP) inflates open rates, so click rate is the only reliable measure of true engagement.
- Pipeline Metrics (Leads & Meetings Booked): Connects B2B demo pages and sales emails directly to qualified pipeline activity.
Secondary Control Metrics (Funnel Protection):
- Bounce Rate: Signals messaging misalignment when a new headline or layout fails to meet visitor expectations.
- Revenue & Sales Conversions: Ensures top-of-funnel conversion gains translate into genuine revenue without compromising deal quality.
A/B Testing Examples
| Test | Control | Variation | Primary metric |
| Email subject line | Your March product update | Three changes that save you time | Open and click rate |
| Landing page CTA | Learn more | Get pricing | Click-through rate |
| Sales email | Feature-led opening | Customer result opening | Replies, meetings booked |
| Website headline | All-in-one platform | Close deals from one inbox | Conversion rate |
A/B Testing vs Multivariate Testing
A/B Testing (Single-Variable):
- How it works: Changes one element (such as a headline or CTA button) to compare a variation against the original.
- Traffic needed: Moderate. Splits traffic 50/50, generating statistically valid results quickly.
- Best for: Email sequences, high-friction landing pages, and lower-volume campaigns.
Multivariate Testing (Multi-Variable):
- How it works: Changes multiple elements simultaneously (such as headline, hero image, and CTA together) to identify the best combination.
- Traffic needed: High to extremely high. Every added combination splits traffic further, requiring massive sample sizes.
- Best for: High-traffic homepages and core website conversion funnels.
Common A/B Testing Mistakes to Avoid
The costliest A/B testing mistakes are lapses in discipline, because genuine wins are rare.
Ron Kohavi, a former Airbnb vice president, co-wrote Trustworthy Online Controlled Experiments. He and Harvard Business School professor Stefan Thomke reported in Harvard Business Review that only about 10% to 20% of experiments at two large search engines were positive. With wins that rare, a flawed test can easily mistake noise for a win.
- Bundling several changes in one variation –The test shows that the package won, but not which change drove the lift. Building both versions from shared CRM email templates keeps the untested elements identical.
- Starting without a hypothesis –Without a stated expectation, you can present any random difference afterward as a meaningful finding.
- Using too small a sample – Small samples miss modest real effects, producing false negatives and discarding good ideas.
- Ending the test early- Stopping when one version looks ahead inflates false positives, because early leads often fade as data accumulates.
- Skipping the sample ratio check-If a 50/50 split delivers noticeably uneven group sizes, Kohavi and Thomke warn that this “sample ratio mismatch” often voids the results.
- Ignoring the guardrail metric – A subject line that wins on opens but loses on clicks doesn’t improve the campaign.
- Generalizing across audiences- A result from enterprise buyers may not hold for small businesses, so confirm it per segment.
- Ignore external factors, seasonal peaks, outages, or a rival’s sale can distort behavior during the test window, so check the calendar before declaring a winner.
Kohavi and Thomke also cite Twyman’s law, which holds that “any figure that looks interesting or different is usually wrong.” Check a surprisingly large win again before the team acts on it.
A/B Testing Tools
Tool selection depends on the specific marketing channel and technical requirements of the test:
- Dedicated Experimentation Platforms: Power website and web app testing using visual drag-and-drop editors, split-URL testing, and advanced audience targeting.
- Email & Automation Platforms: Run native A/B tests on subject lines, email copy, and send times directly in campaign workflows.
- CRM & Sales Systems: Test sales sequences and outreach messaging while mapping performance directly to lead conversion, deal stages, and closed revenue.
How Vtiger One Supports A/B Testing
Vtiger One lets marketing teams send different email campaign variations to two subscriber groups. Campaign reports then show unique opens, clicks, bounces, and unsubscribes for each send.
Segments built from past campaign behavior define who receives each test, and email sequences let sales compare outreach approaches. You can then add winning variants to nurture flows in marketing automation software, so a tested message stays in use after the test ends.
Frequently Asked Questions
What is A/B testing in simple terms?
A/B testing shows two versions of something, such as an email or a web page, to two random groups of people. The version with more clicks or replies wins. Teams use it to settle choices such as subject lines, CTA wording, and page layouts with evidence instead of debate.
What is an example of A/B testing?
A company sends one subject line to half its newsletter list and a different one to the other half. If the second produces a meaningfully higher open and click rate across a large enough list, the team adopts that style and tests the next idea.
What is the difference between A/B testing and split testing?
In everyday use, there is no difference, because both names describe dividing an audience between two versions. Some practitioners reserve split testing for tests that send traffic to two separate URLs, while A/B testing changes an element on the same page.
How long should an A/B test run?
Run a test until it reaches the sample size calculated before launch, and for at least one full business cycle. For many websites, that means one or two full weeks, so you capture both weekday and weekend behavior. Email tests often finish sooner, because the whole list receives the message at once.
How much traffic is needed for A/B testing?
The required traffic depends on the current conversion rate, the smallest improvement you want to detect, the confidence level, and the statistical power. A low-converting page testing a small change needs far more visitors than a bold email test. A sample size calculator turns those inputs into the number each version needs.
What is statistical significance in A/B testing?
Statistical significance asks how likely a difference at least this large would be if both versions performed the same. Many teams test at a 5% significance level, often called 95% confidence. Significance makes chance an unlikely explanation, but it cannot show whether the lift matters commercially.
What is the difference between A/B testing and multivariate testing?
A/B testing compares two versions that differ by one change, so results come faster and are easier to interpret. Multivariate testing changes several elements at once and tests their combinations. It shows how elements interact, but it needs much more traffic to reach a reliable answer.
