Skip to main content
Create A/B tests in the Algolia dashboard. Choose a test type, configure the variants and traffic allocation, then review and launch the test.

Before you begin

Check that your control index receives search traffic and that your app sends the events needed for the metrics you want to measure. Test each configuration with your search UI before you launch an A/B test. For example, changes to a can affect filtering, and ranking changes can reduce search relevance.

Create and launch a test

1

Choose a test type

Select your Algolia application and open the A/B Testing page. Click Create test, then select what you want to compare:
  • A feature: test a supported feature or compare its configurations. Available features depend on your app and permissions.
  • Index configurations: compare search-time settings on a single index, such as typo tolerance, synonyms, facets, geo search, and distinct.
  • Index vs. Index test: compare separate indices or replicas. Use this option for settings that require separate indices, such as searchable attributes or custom ranking. Screenshot of the "Choose test type" step, showing feature, index configuration, and multiple-index options.
2

Build the test

Select the index that receives your production search traffic as the control.For feature and index configuration tests, the control and variants share a baseline index. Configure each variant with the change you want to measure. For an Index vs. Index test, select a separate index for each variant.The dashboard checks every selected index:
  • An index that’s already part of an active A/B test blocks the next step.
  • An index without recent search traffic or events shows a warning. You can continue, but the test might not collect enough data for useful results.
When testing with replica indices, use the main index, which receives search traffic, as the control. Use the replica as a test variant. A replica doesn’t receive traffic when used as the control because Algolia assigns variants during the .
Screenshot of the "Build test" step, showing a baseline index with control and variant configurations.
3

Configure the test

Choose how much traffic to send to each variant. Start with an equal split unless you want to limit exposure to an untested configuration. The traffic percentages must total 100%.For supported test types, select a target metric and the smallest relative change you want to detect. The dashboard uses your historical traffic and events to estimate the test duration. If duration estimation isn’t available for the selected test type, set the duration manually. For more information, see Understand sample size and test duration.Select a statistical method. Frequentist methods use statistical significance and a p-value to evaluate results. Bayesian methods use probability to be better and a credible interval, which represents the range of likely uplift. Bayesian is the default, and the dashboard reuses the method from the last test you created successfully. To learn more, see Bayesian experimentation.Screenshot of the "Configure" step, showing traffic allocation, target metric, and custom test duration.
4

Review and launch the test

Enter a descriptive test name and an optional hypothesis. Review the test type, variants, traffic allocation, target metric, and duration. To change a section, click its edit button.If everything looks correct, click Launch test.Screenshot of the "Review" step, showing the test name, hypothesis, and configuration summary.
You can stop a test before its planned duration or wait until it finishes. You can’t restart a stopped test because interrupted testing can make its results inaccurate.

Understand sample size and test duration

Accurate sample size determination is essential for effective A/B testing and reliable results. The sample size estimation tool simplifies this process by estimating the number of searches needed to confidently detect meaningful differences between experiment variations.

Sample size estimator

The estimator uses historical traffic and a baseline metric, such as conversion rate or click-through rate. Select the metric and the smallest relative change you want to detect. The calculation uses 80% statistical power and a 5% significance level, adjusted for multi-variant testing. The estimator rounds the duration up to whole weeks to account for traffic patterns. Key features:
  • Historical baseline rate. The tool uses your existing data to determine the current performance metric, such as conversion or click-through rate, against the measured changes.
  • Fixed statistical parameters. A statistical power of 80% and a significance level of 5% (adjusted for multi-variant testing) provide a balance between detecting real effects and minimizing false positives.
  • Customizable metrics and effect size. Select a metric, such as conversions or clicks, and set the minimum effect size you want to detect.

Choose an appropriate effect size

The effect size is the smallest relative change in a metric you consider significant enough to act on. Choosing an appropriate effect size is essential for accurate and efficient A/B tests. For example, if you would adopt a change only if it increases your conversion rate by 5%, set your effect size to 5%. With a baseline conversion rate of 10% for the control variant, a test variant would need to show a conversion rate of 10.5% (a 5% relative gain) to count as an improvement. The sample-size estimator shows the number of searches required to detect that improvement with 80% power at a 5% significance level (adjusted in multi-variant testing). When choosing an effect size, consider:
  • Impact. Consider the smallest change that would make a meaningful improvement in your goals. For example, a 2% relative increase in conversion rate might be significant for you, while another organization might aim for a 5% relative change.
  • Historical data. Review past experiments to understand typical variations and set an effect size that’s realistic and achievable based on historical performance. For example, if other changes you have tested result in a 3% relative increase in conversion rate, you might set your effect size to 3%.
  • Balance between sensitivity and practicality. Smaller effect sizes, such as 1% to 2%, require larger sample sizes but let you detect subtle changes. Larger effect sizes (for example, 5% to 10%) require smaller samples and are easier to detect but may overlook smaller, yet important, changes.

View the results

Open a test from the overview page to review its status, variants, and metrics.

Overview page

The overview page lists the A/B tests in your , including each test’s ID, name, start date, planned duration, and status. The list shows active tests first, followed by other tests from newest to oldest. To change the sort order, click the chevron next to Start date. Screenshot of the "A/B Testing" page listing test IDs, names, start dates, durations, and statuses. Select a test to open its details page.

A/B test details page

Screenshot of A/B test details page showing 'Personalization Impact Test' with Control and Variant B metrics and CTR breakdown. At the top you see a card for each variant (Control, B, C, …) showing its traffic share, tracked users and searches. The Metric breakdown section shows a table for every metric: Metric breakdown tables with CTR, ATCR, CVR and more The Metrics breakdown section compares the metrics for each tested variants with the following columns:
  • Variant. The tested variant.
  • Tracked searches. Searches with clickAnalytics set to true. Metrics that don’t require this parameter use Searches instead.
  • Metric value. The observed value of the metric (for example, CTR, or ATCR).
  • Difference. The relative change from the control, expressed as a percentage.
  • Confidence interval. A range that expresses uncertainty in the estimated relative change from the control.
  • Confidence rating. Shows how reliable the observed difference is (for example, trending confident, or unconfident). For details, see Confidence. Hover over any rating to see the exact p-value.
You can also open the full Analytics dashboard for the test’s variants by clicking View analytics in the top right corner. For frequentist results, use the Confidence over time chart in Additional insights to inspect a variant and metric. The chart shows how the estimate and confidence interval change during the test. Screenshot of "Confidence over time" showing Variant B's conversion-rate difference and confidence intervals from day 3 to day 30. If possible, wait until the A/B test has finished before interpreting the results.
A/B tests show results in the dashboard within the first hour after you created them, but metric comparisons won’t show until at least 1 week or 20% of the test duration has elapsed. This mitigates the risk of drawing inaccurate conclusions due to insufficient data. Test results are updated every day.

Test statuses

The potential test statuses are:
  • In progress - Not enough data: the test has started, and there is insufficient data to draw reliable conclusions. Wait for at least one week, or 20% of the test duration, before metrics begin to show. Hover over the badges to see data, but avoid drawing conclusions during this stage.
  • In progress: the test is collecting data and comparing metrics.
  • Failed: Algolia couldn’t create the test. Check the index and test configuration, then try creating the test again or contact Algolia support.
  • Stopped: you stopped the test. Test traffic returns to the control. Algolia retains the test’s metadata and metrics, and the test remains visible in the dashboard. Its results might be inconclusive.
  • Completed: the test has finished. Your app is back to normal: index A performs as usual, receiving 100% of search requests.
Deleting a test permanently deletes its metadata and metrics and removes it from the Algolia dashboard.

How to interpret results

Compare the size of the improvement with the effort, cost, and risk of adopting the change. For example, a 4% relative increase in click-through rate or conversion rate might not justify restructuring your . A low implementation cost doesn’t make an uncertain result reliable. Consider the evidence, the expected benefit, and any negative effects on other important metrics before adopting a variant. For more information, see How A/B test scores are calculated.

Confidence

A/B tests use either frequentist or [Bayesian statistical methods]](/doc/guides/ab-testing/what-is-ab-testing/in-depth/bayesian-experimentation). The confidence ratings in this section apply to frequentist tests and depend on the test status and p-value. A p-value is the probability of observing results at least as extreme as those measured, assuming no difference between variants and that the test’s statistical assumptions hold. It isn’t the probability that the result is a false positive, the false-positive rate, or the probability of reproducing the result. A small p-value provides evidence against the assumption that the variants perform the same. A large p-value doesn’t show that the variants perform the same. Frequentist tests use these confidence labels:
  • Not enough data to interpret: the test has started, and there is insufficient data to draw reliable conclusions on the performance of the variants. Wait for at least a week, or 20% of the test duration, before metrics begin to show. Hover over the badges to see data, but avoid drawing conclusions during this stage.
  • No data: the test doesn’t have data to compare. Check that your app sends the events required for the selected metric.
  • Unconfident: the available data doesn’t provide enough evidence for a difference. The rating can change as the test collects data.
  • Trending confident: the data suggests a difference, but the test is still running. The rating can change.
  • Inconclusive: the test has finished, but the confidence is too low to determine whether the observed change is due to chance. This might be because of insufficient data or high similarity between the variants.
  • Confident: the test is complete, and the observed change probably reflects the true impact.
A Confident or Trending confident result doesn’t mean the change is beneficial. The label describes the evidence for a difference, not its direction or value. For example:
  • A Confident decrease in conversion rate is evidence against adopting the change when your goal is to increase conversions.
  • A Confident increase in conversion rate supports considering the change. Also assess its size, cost, and effect on other metrics.
  • A Trending confident increase is preliminary. For a frequentist test, wait until the planned end before making a decision.
  • An Inconclusive result doesn’t show whether the change has an effect. Consider a new test with a longer planned duration. More data can improve precision, but it doesn’t guarantee a conclusive result.

Collect enough data

Don’t base decisions on insufficient data or low confidence results. A small sample can make it hard to determine the effect. Use the statistical indicators for your selected method when assessing the results. For frequentist tests, follow the planned duration rather than stopping as soon as the result looks favorable.

Recommendations

  • Test before going live. Be wary of breaking anything. For example, ensure both your test indices work with your UI. Small changes can break your interface or strongly affect the user experience. For example:
    • Making a change that affects a facet can cause the facet’s UI logic to fail.
    • Changing a simple ranking on index B can make the search results so bad that users of this index have terrible results. This isn’t the purpose of A/B testing. Index B should theoretically be better or at least as good as index A.
  • Don’t change your A or B indices during a test. Don’t adjust settings during testing. This pollutes your test results, making them unreliable. If you must update your data, do so synchronously for both indices and, preferably, restart your test. Changing data or settings during a test can break your search experience and undermine your test conclusions.
  • Don’t use the same index for several A/B tests. You can’t use the same index in more than one test at the same time. You’ll get an error.
  • To test different settings on the same index, select that same index for both variants. No replica needed.
  • Make only small changes. The more features you test simultaneously, the harder it is to determine causality.

API clients

Use the Algolia dashboard for most A/B tests. Use an Algolia API client when you need to automate test management, for example:
  • Create similar tests for several websites that use separate indices.
  • Start tests from your backend after data changes or in response to analytics. Monitor the results and be ready to revert changes that reduce important metrics.
Your needs the permissions required by each operation:
  • Create or delete A/B tests: editSettings.
  • Retrieve A/B test results: analytics.
Last modified on September 23, 2026