Testing your pricing pages has the potential for the biggest impact on a business’s top-line.
Every percentage increase in conversion results in a direct increase in the profit.
The challenge comes when you try to implement an A/B test in a slow-moving market with limited traffic coming to your pricing pages (usually measured by purchases or sign-ups).
The standard marketing advice regarding conversion rate optimization programs (CROs) is ineffective in this situation.
Most companies provide tools designed for mass monetization where tens of thousands of purchases are expected.
Yet when you have only a few hundred visitors to your SaaS/B2B sites’ pricing pages every month, it’s almost impossible to use traditional statistical significance to calculate results.
Waiting for statistical significance could lead to long wait times to receive test results.
Seasonal fluctuations in your traffic may negatively impact the value of your test data.
As competitors react quickly to changes occurring within your test environment, you will remain at a disadvantage while waiting for the turnaround time of a software dashboard's green lights.
You need to develop a wholly different operating process to run successful experiments on low-traffic pricing pages.
It is imperative that you do not waste funds on testing minor cosmetic changes to your landing pages.
You do not have the volume of data necessary to evaluate results accurately.
Additionally, poor random assignment of traffic will increase your chances of running a false-positive A/B test.
Balancing scientific reliability with the realities of slow sales processes will be challenging.
You will need to isolate test variations strictly, establish realistic minimum detectable effects, and be absolutely committed to running properly powered A/B experiments.
How to conduct A/B experiments in low-traffic environments
If you are operating in a low-traffic environment, immediately consider this when establishing your test protocols.

Do not attempt to implement traditional testing processes without considering the traffic volume constraint that exists.
Traffic volumes dictate the type of methodology to use; therefore, methodology must be established with this assumption in mind.
Testing by standards requires large amounts of traffic to compare two versions of a web page.
Without enough traffic, testing should be focused on major differences in the structure and function of the web page rather than just minor adjustments.
It is important to isolate one variable to measure its effect.
Low visitor volume limits your ability to determine what effect other changes made at the same time had on your visitors.
If you implemented a price change, tier name change and billing toggle change all at once, you would not have enough data to properly identify the cause of any movement of conversions in one direction or another.
Stop obsessing over significance based on frequentist statistics.
In low-volume testing, there is no mathematical basis for establishing a 95% confidence interval based on 400 monthly visitors.
You will need to leverage Bayesian approaches, or sequential testing methodologies that allow you to make risk-based decisions to determine what changes worked based on data collected incrementally.
Qualitative research is the tool that bridges the gap between low-volume web pages and the minimum detectable effect (MDE).
When volume is too low for statistical conclusions, user testing, exit surveys and sales team evaluation of pricing changes should be used to validate pricing changes before making them available to potential customers.
The math behind low traffic testing
The basic math associated with low-traffic testing is related to the conversion rate, the sample size and the MDE.
It takes an excessive amount of traffic to detect a small change in traffic behaviour.
Many times, teams under-estimate the amount of traffic required to run a valid experiment.
They too quickly identify the incorrect conversion rate of a variant, when watching their dashboard over a two week period, and declare a test winner.
This is not experimentation; this is gambling.
The problem with standard experimentation calculators
When you use a standard experimentation calculator to plug in your numbers, the calculator will provide you with the required sample size based on static assumptions.
They typically assume that you have an 80% chance of finding an effect (statistical power) and that you are willing to pay 5% (alpha level) for a false positive.
These baseline assumptions generate sample sizes in a low-traffic environment, where you would need a year to collect your samples.
In addition, most calculators of statistical power also assume the presence of a perfect test environment in which there are no cookie deletions, no multi-device use, and a 60-day (or so) B2B purchase cycle.
If your test runs for four months to reach the required sample size, then by that time, the traffic mix of the users of both samples may have changed.
If you conduct your test over a long period of time, the original results will be invalidated due to the influence of other external factors on the two sample groups (e.g., seasonal holidays, new products or services).
Reality check for Minimum Detectable Effect (MDE)
The MDE is the smallest change in conversion rate you want to test.
If you want to test for an increase in conversion of 2% on the pricing page, you will need hundreds of thousands of visitors to your site.
If you only have 500 visitors a month, you will only be able to test changes with the expectation that they will generate a huge increase in conversion (e.g., 20-30%).
Minor changes will not produce a 30% increase in conversion.
If you change a button from green to blue, the chances that you will achieve your MDE in a low-traffic environment are slim.
To achieve the required MDE, the variant must create a fundamental change in the way the user assesses value.
You should be testing high-risk, high-reward variables.
If the change is minor enough that it won't directly change the baseline metrics, you shouldn't spend your time testing it on a low-traffic page.
A/B test triage - Evaluating whether or not A/B test will work before you start
Before you start designing your experiment, you need to determine whether or not you should even use an A/B Test (A/B Tests are often incorrectly assigned).
In some cases, the use of an A/B Test will not provide usable data because the conditions required for accurate data are not met.
Many groups attempt to create A/B Tests for situations where, mathematically, the chances of success are extremely low or impossible.
Accurate results require a stable baseline.
If your website receives fluctuating traffic due to a single LinkedIn post or a sporadic email, the baseline of a test will be distorted.
Volume threshold
You should establish your maximum duration of a test before running it.
General consensus in the industry is that A/B Tests should generally not last more than four to six weeks; beyond that time limit, issues arise with sample pollution, which degrades the reliability of the sample data collected.
To establish the maximum duration for your A/B Test, you must determine how many visitors you will receive during the six-week period.
Multiply the average number of visitors during that time by the baseline conversion rate.
The result provides the expected volume of test participants during the six-week period.
If the statistical analysis indicates that you can detect a 40% increase in conversion rate due to your proposed change, you need to evaluate whether your proposed change will actually generate a 40% increase in conversion rate.
If your answer is no, then standard-controlled A/B Testing is not applicable.
Qualitative validation - When statistical testing will not occur
The only means to validate your concept is to conduct qualitative validation since statistics cannot be relied upon for testing and evaluating your proposed concept.
Instead of conducting a blind A/B Test where you will never achieve statistically valid results, you should conduct deep usability studies with a sample of your target audience to collect qualitative feedback regarding both the existing and new pricing page designs.
During the Usability Study, you must encourage the Qualified Buyer to vocalize their reasoning behind their evaluation of your pricing page design.
Identify their ideal plan selection and the reason for their choice.
Determine areas of confusion and ask for clarification.
For example, if your five (5) competitors clearly do not understand the value metric of the new page, you do not need an A/B test to prove that the new page is ineffective.
Qualitative research provides context for quantitative data, thus enabling you to analyze the reasoning behind the numbers while taking risks if you do not have the quantity of traffic to support a conclusive "what".
Implementation of direct ship process
In some cases, the impact that delaying a crucial price update could create could cause more damage than the risk associated with launching it without performing an A/B test beforehand.
If your organization is transitioning from a user-based to a usage-based pricing model due to a significant redirecting of business focus, conducting an A/B test on this occasion would be counter-productive.
Documentation of baseline metrics should be carried out, followed by the implementation of the new pricing model across 100% of traffic and the monitoring of performance in the subsequent cohorts for 90 days.
This is a strategic rollout and it should be regarded as such.
The distinction between an experiment and a strategic rollout is critical to avoid misallocation of time, energy and cost in analyzing results.
How to execute a successful A/B pricing page test when traffic is minimal
If your idea passes the initial evaluation process and the idea is suitable for testing, execution will fall predominantly upon your ability to effectively conduct the A/B test and minimize the amount of data contamination in addition max performance of your test process.

A professional pricing page redesign will entail changing the monthly price, changing the names of the plans, adding checkmarks for any new features, and re-designing how the information is shown on the page.
And they will do all of this in one variant.
If you see a 15% drop in conversion, you cannot know why this drop happened.
Was the price point too high?
Were the names of the pricing plans confusing?
Did the new layout place the primary call-to-action too low on the page?
To get accurate results from your testing, you must control for all but the one specific item you are testing.
For instance, if you want to see what effect the "Recommended" plan has, don't change the price.
Don't change the copy.
Change nothing but how the single column (the "Recommended" plan) is visually highlighted.
Only by keeping everything else constant can you make a reliable conclusion about causality.
Traffic segmentation and sticky assignment
For a high-stakes pricing test, you must ensure that user assignment is perfectly sticky; once a user is assigned to a specific variant, they must be assigned that same variant whenever they return, regardless of how they access the pricing page.
If a user visits the pricing page on Tuesday and sees Variant A (which is $49/month), then returns on Thursday, via mobile phone, and sees Variant B (which is $79/month), then no trust exists for that user.
Consequently, the high-stakes pricing test has failed.
To achieve sticky assignment, you must configure any experimentation platform to assign variants at the account (or authenticated user) level whenever possible, instead of relying solely on browser cookies that can easily be deleted.
If a large percentage of your traffic is anonymous, you must accept more error margin regarding cross-device contamination, but you should actively exclude known sources of bad traffic (like mobile social media clicks with very low intent) from the audience for the test.
Establish downstream guardrails
The purpose of conducting a clean pricing test is to assess revenue, not just click-through rates.
Optimizing a pricing page's initial "Click to Start Trial" button can be hazardous.
If you were to cut your prices in half, it would likely double your click-through rate, making your conversion rate look great on the testing dashboard.
However, if the result is that those users churn at a doubled rate or return only half of their lifetime value, then that would indicate a catastrophic failure for the business.
Therefore, it is crucial that your test plan includes "hard guardrails" around metrics such as lead quality, sales pipeline velocity, and eventual activation rates.
If Variant B increases your initial signups by ten percent but also decreases ARPU by fifteen percent, then it has failed as a variant.
The tradeoff thresholds should be clearly defined in a clean test before applying any visitor.
Statistical methodology selection
If the frequency of testing necessitates a substantially larger volume of traffic than your business can produce, you will have to modify your mathematical approach.
The statistical methodology that you choose will determine how you evaluate the findings of your tests.
The peril of traditional frequentist methodologies
The frequentist testing methodology is founded on the principle of a null hypothesis, which demands that your sample sizes be large enough to allow you to reject the null hypothesis with a certain degree of confidence that there is no difference between the control and the variant.
Given that most frequentist testing will produce data with a frequency of low volume, the result is typically flatlined tests, in which you run a test for six weeks to have an insignificant 60% result generated by the tool at the end of the six-week period.
The strict frequentist testing rules would indicate that you have generated no new knowledge, and you have wasted six weeks' worth of traffic.
Implementing bayesian decision-making approaches
Bayesian Decision-Making is much more user-friendly as a tool for quantifying low sample sizes.
A Bayesian approach does not ask the obvious question: "Is there absolute proof that Variant B is different from the Control?"
It asks the more appropriate question: "Based on the information I have collected thus far, what is the probability that Variant B will outperform Control?"
By applying Bayesian decision making, a business is able to confirm risk-weighted strategies.
For example, if a test lasts four weeks and the model shows 85% likelihood of success compared to its predecessor, a business can confidently move forward with it—the risk of making this decision is very low (though not zero).
In addition, most B2B companies that operate in low-traffic segments will benefit from having the ability to make decisions based on an 80% to 85% confidence ratio instead of being held hostage to a 95% confidence level that will likely never be achieved.
Sequential test design concepts
Sequential Testing allows for interim data access while still preserving the integrity of the experiment.
Normal Test Design prohibits early termination for any variant exhibiting substantial strength, as opposed to Sequential Testing which supports previously defined "cutoff" values for statistical performance results.
These cutoffs indicate significant performance differential that enables termination of a test once the endpoint has been detected.
Accessing interim data during a Sequential Test Design is crucial for reducing risk associated with the pricing page testing process.
It would be counterintuitive for marketers to spend four weeks confirming that their new pricing strategy has dropped conversion rates by 40%.
You have a mathematical reason for terminating a test early if it is producing poor results from sequential testing of multiple versions.
What you will be testing with limited traffic
If you can only run as many as two to three tests with traffic that is sufficient to produce definitive results, your backlog of tests will need to be very selective in nature.

You will want to emphasise structural changes over surface design changes.
Structural architecture change testing
Testing for a change in the structure of your proposed solution is a significant upgrade in value.
For example; when testing a two-tier approach vs. a three-tiered approach, or condensing four complex layouts down to two streamlined options can potentially change how your target market evaluates your product.
As a result of such changes your prospects will mentally shift from purchasing a product to what option is the best fit for their needs.
These types of structural modifications can have a considerable impact on your minimum detectable effect.
Billing defaults and commitment anchors
The default setting of your billing toggle is a common and significant pricing experiment you can conduct.
If your current billing page is set to default to monthly billing, changing the default to an annual billing option can provide you with significant cash flow benefits.
Your prospects may experience "visual shock" at the large annual numbers presented to them but regardless of whether your conversion is lower, the customers that do convert will provide you with substantial upfront capital.
There are many factors you can test structurally, including the placement of your social proof, removing extraneous navigation links, or moving your highest price tier to the left side of the page (to serve as a price anchor) which all can greatly impact your prospect's psychology.
Things you should never test
Button colours, font sizes or minor edits to text (below the fold).
Vanity tests are a waste of your precious traffic and rarely yield reliable statistical outcomes.
If there is anything wrong with or intentionally misleading in your pricing page copy, only minor changes to your headline position will impact your revenue positively enough to warrant the length of time you spent running such a test.
Therefore, use your limited traffic to run tests that will change your business' unit economics.
Conclusion
The testing of prices in a low-traffic environment is a matter of strict operational discipline.
Copying the testing workflows of enterprise companies who have millions of visitors each day will not be successful.
You need to determine your Minimum Detectable Effect realistically and only run large, significant structural tests.
You need to stop relying on rigid frequentist significance to help you make decisions about risk-adjusted measures, and instead begin to use Bayesian probabilities to help you.
Above all, isolate your variables during your testing process and ensure that the assignment logic you used is set up correctly, as this will help protect your data integrity.
If you are not able to adhere to these standards, stop using the A/B testing software and use qualitative user research combined with direct observation to monitor individual tests.
A poorly executed test can destroy your revenue pipeline and give you a false sense of security about the data you have collected.
Frequently Asked Questions (FAQs)
What is the maximum duration for A/B testing that is low traffic?
As a "baseline" benchmark, a standard test should be limited to six to eight weeks of duration.
The following issue can arise when performing a sample A/B test on a pricing page for low-traffic websites is the error or "noise" received from sample contamination due to the user contacting multiple devices or other seasonal behaviors as described above.
If a test requires a baseline of three months of analytic data from a specific website's pricing page to achieve a valid sample size, the level of base traffic is much too low and you cannot expect valid results.
Can microconversions be used as an alternative to revenue?
Microconversion events (eg click on the pricing tier) can replace the long turnaround of receiving substantial sampling sizes for pricing tests to allow rapid conclusions to be reached; however there is a great deal of risk involved in using microconversion events.
For example, if version A results in a 20% increase in click-through rates to the checkout page versus version B, that could actually result in version A having fewer purchases completed than version B due to a pricing shock when the billing occurs.
Therefore when using microconversion events to "call" a test early, it is critical that you maintain tracking on those items to validate whether your revenue decreased significantly after the test's conclusion.
Is geo-splits a valid alternative to price points for SaaS offerings?
Geo-split testing (where you show variant A to people in the United Kingdom versus variant B to people in Canada) may work well for major consumer brands, but there are extreme flaws in using this methodology for low-traffic SaaS brands.
In addition, buying power, market maturity, and the number of competing alternatives are significantly different by region, and therefore the baseline conversion rates of these regions are also likely to be different.
For low-traffic instances, the degree of noise generated by the geographical differences between these two regions completely overshadows the signal generated by your price changes.
What happens if you stop a test before reaching a statistically significant level of results?
Stopping a frequentist experiment or A/B price testing before the specified statistical significance level has been reached means that the results cannot be definitively attributed to changes made during the experiment, and that any changes in results will be more likely caused by random variance.
If you must stop an incomplete experiment to determine your next steps, we recommend switching to a Bayesian analysis to know how likely you are to have a successful outcome from the changes you tested, or assuming that the control was the successful variant (assuming the control is the same as the previous baseline) until you achieve additional data to determine accuracy.
