I search through your content to help you find answers to your questions, fast.
Re-evaluating LLM encoders for semantic search
In this study, we explore the transferability of MTEB and publicly available ecommerce dataset benchmark performance to real-world retail search applications.
The rapid advancements in Large Language Model (LLM) encoders have greatly improved semantic search across various domains, especially in scenarios where traditional token-level embeddings are inadequate. Text embeddings are essential for enabling machines to represent human language, empowering them for tasks such as semantic search. By converting human language into vectors, embeddings allow LLMs to grasp nuanced linguistic features and relationships between textual elements, such as contextual meanings by going beyond the token overlapping nature of keyword search. Words are inherently high-dimensional and sparse, and embedding them into lower dimensional dense vectors enable LLMs to process information a lot faster. At Algolia, fine-tuned LLMs (Masked Large Language Models) empower an end-to-end search experience (Figure 1). Benchmarks like the Massive Text Embedding Benchmark (MTEB) have been instrumental in evaluating these models on tasks such as retrieval, clustering, and ranking using public datasets. However, a critical question remains: do high performances on these benchmarks translate effectively to specialized domains like retail semantic search?
Figure 1: NeuralSearch (NS) at Algolia is an end-to-end search experience empowered by LLMs
In this study, we explore the transferability of MTEB and publicly available ecommerce dataset benchmark performance to real-world retail search applications. Our findings indicate that models excelling on MTEB do not necessarily maintain their superiority when applied to retail-specific datasets. To investigate this, we curated a private retail dataset mirroring the relevance scoring system of publicly available ecommerce datasets. We leveraged the existing MTEB codebase in our evaluation pipelines to ensure consistency in our evaluation methodology.
By evaluating a range of both open-source and commercial LLM encoders on our dataset, we observed discrepancies in model rankings compared to those on MTEB specifically. Models that performed exceptionally well on MTEB did not always deliver the best results in our retail context. These findings highlight the limitations of relying solely on general benchmarks for model selection in specialized domains.
Our article underscores the necessity for domain-specific evaluations and benchmarks. Without them, organizations risk deploying models that are ill-suited for their specific needs, potentially compromising the effectiveness of their semantic search systems. We advocate for the development of more specialized benchmarks and the inclusion of domain-specific datasets in model evaluation processes to ensure optimal performance in real-world applications.
The evaluation of text embeddings plays a crucial role in understanding their effectiveness across various natural language processing tasks, including retrieval, clustering, and classification. Open benchmarks such as MTEB, BEIR, STS-B, and TREC/MS MARCO offer standardized frameworks to assess the quality of embeddings in diverse scenarios. BEIR (Benchmarking Information Retrieval) provides a comprehensive suite of retrieval tasks covering multiple domains. STS-B (Semantic Textual Similarity Benchmark) focuses on sentence-level similarity scores, measuring embeddings' ability to capture semantic equivalence. TREC/MS MARCO benchmarks are widely used for large-scale passage ranking and retrieval, emphasizing real-world information retrieval systems. Among these, MTEB stands out as the most versatile and comprehensive choice for retrieval purposes. It evaluates embeddings across a broad range of languages, domains, and tasks, making it uniquely suited for retrieval scenarios that require robust performance in diverse and multilingual environments. With its expansive coverage and detailed task evaluation, MTEB ensures embeddings are tested for generalizability and efficacy, making it the ideal benchmark for retrieval applications.
MTEB is introduced due to challenges associated with many existing benchmarks leveraging a small set of datasets from a single task not covering their possible applications, and various models are constantly being proposed without proper evaluation. Although there are both open and closed sourced LLM models, which makes MTEB the most comprehensive benchmark for LLM embedding models, many evaluations are based on models that are often dated, focusing on models and datasets from a time when the field was less mature. In addition, the rapid advancement of LLM encoders in semantic search made MTEB the primary standard for developers to pick the best performing models. However, generic LLM encoders evaluated on public datasets may not capture the nuances of specialized domains. Retail semantic search involves unique challenges, including diverse product descriptions, varying customer search behaviors, and industry-specific terminology, therefore, MTEB does not necessarily excel in the retail context, highlighting a significant gap between benchmark performances and real-world applicability of LLM models.
In this article, we used three different datasets:
The list of datasets used for this study is shared below:
| MTEB dataset (En) | Public ecommerce dataset (En) | Private ecommerce dataset (En-De-Es-Ru-Fr) |
|---|---|---|
| hotpotqa | esci | Fashion and apparel |
| fever | wands | Electronics and technology |
| arguana | crowdflower | Home and living |
| climate-fever | marqo | Beauty and personal care |
| dbpedia | homedepot | Toys and hobbies |
| fiqa | Health and wellness | |
| msmarco | ||
| nq | ||
| cqadupstack-english | ||
| nfcorpus |
Table 1: Algolia internal benchmarks include the Open Retrieval Benchmark, a public ecommerce dataset benchmark, and a private ecommerce benchmark curated from customer performance data. Note that the private benchmark is developed in collaboration with customers to evaluate how models perform on their specific data distributions
This study uses only the corpus associated with the relevant documents of the split of concern which is test split. This is different from how MTEB code base computes the metrics. This approach is favored for better compatibility with ecommerce datasets. Also, it reduces the chances that relevant pair matches are ignored in the evaluations. All evaluations are conducted on test splits (that represent 20% of all datasets) for public and private ecommerce datasets. Private datasets are multilingual, and both English and multilingual models are selected for evaluation. Also, private datasets have only relevant pairs which is similar to MTEB retrieval datasets. For open-source LLMs, top performing models on MTEB that are 500 million and fewer parameters are selected. Closed-source models selected are OpenAI, Google, and Cohere models that are accessed at the time of analysis.
LLMs are evaluated on various datasets using standardized metrics such as nDCG@10, F1@10, and MAP@10. The metrics such as nDCG (normalized Discounted Cumulative Gain) and F1 are for evaluation of retrieval performances of LLM models, whereas MAP are primarily for evaluation of their ranking performances. Each metric is explained in detail below.
nDCG is the normalized DCG (Discounted Cumulative Gain) and is calculated by summing the relevance gains from each product, discounted by the position at which it appears. The numerator shows the relevance score of the product at position i and the denominator shows the discounting factor that reduces the contribution of products that appear in lower ranks. To evaluate the effectiveness of DCG, it is compared to the highest possible DCG that could be achieved if the products were ranked in the perfect order of relevance, i.e., IDCG. This normalization makes it straightforward to compare between different sets of results.
F1: The harmonic mean of precision and recall, providing a single metric that balances both. Recall is the metric that determines how many of all that are actually positive did the model correctly identify, whereas Precision is the metric that determines how many of all the items predicted to be positive were actually correct. The F1 score is high only when both precision and recall are high, which is useful specifically when the dataset is imbalanced.
MAP: Average Precision is the average of the Precision scores computed after each relevant product is retrieved. It is called Mean Average Precision (MAP) when averaged over many queries. MAP is an indicator of whether the majority of relevant products are shown at the top positions. MRR focuses on the most relevant products to be at the top where MAP focuses the majority of relevant products ranked at top positions.
Algolia builds embedding models that are approximately 500M parameters to ensure desired latency. Each LLM is architecturally optimised and quantised to ensure lower latency (Table 2). The Algolia AI team continuously assesses state-of-the-art LLMs, selecting those with top performance and permissive licenses. Fine-tuning these models for ecommerce contexts ensures superior performance tailored to industry-specific needs. The fine-tuning methodology combines the best practices from cutting-edge research with the AI team's expertise. Leveraging automated AI training and evaluation pipelines, the process optimizes model performance by simultaneously exploring numerous hyperparameters on the same training dataset, resulting in the best possible models. Some of the techniques inspired by the latest research, without delving into exact details, are outlined in the table below:
| Stage | Technique | Comments |
|---|---|---|
| Fine-tune (infoNCE loss) | Stratified public ecommerce datasets in batches | To ensure the best possible outcome is achieved from infoNCE loss, stratified datasets are curated in the same batch. |
| Hard-negative fine-tune | Synthetic hard negatives for further fine-tuning | A combination of GenAI labelling and tuned hard negative mining is leveraged to ensure the resultant model can separate the decision boundary between vague samples. |
Table 2: Algolia leverages two-stage training approach: 1. fine-tuning using infoNCE loss 2. hard-negative fine-tuning leveraging synthetic datasets
Algolia embedding LLMs (as of December 2024) with their specifications are provided in Table 3. All Algolia LLMs are trained on publicly available ecommerce datasets, and no Algolia customer datasets are used for training purposes. GenAI labeling and hard negative mining are combined to create synthetic datasets for further fine-tuning. Algolia v2410 models are open-sourced under MIT license, and they can be accessed at Hugging Face. Note that latency is computed on a local machine with an i9 CPU.
| Model | License | Base | Datasets | Dimension | Latency |
|---|---|---|---|---|---|
| Algolia-Large-EN-Generic-v2410 | MIT | gte-large | Public ecommerce (+Syn.) | 1024 | 90 ms / 40 ms (opt.) |
| Algolia-Large-Multilang-Generic-v2410 | MIT | solon-embeddings-large-0.1 | Public ecommerce (+Syn.) | 1024 | 90 ms / 40 ms (opt.) |
| Algolia-Large-All-Generic-v2412 | MIT | snowflake-arctic-embed-l-v2.0 | Public ecommerce (+Syn.) | 1024 | 90 ms / 35 ms (opt.) |
Table 3: Algolia LLMs are state-of-the-art embedding models fine-tuned and optimised for ecommerce search
All three Algolia LLMs ranked in the top 20 for multilingual ecommerce datasets including English (Figure 2). Algolia v2412 is at the top of the leaderboard with its NDCG@10 performance. Algolia Multilingual v2410, which is available on Hugging Face with permissive license as of December 2014, ranks at 6th place above Cohere multilingual LLM. The performance difference between our v2410 and v2412 multilingual models is approximately 3% in NDCG@10.
Figure 2: Algolia LLM performance on ecommerce benchmark that includes Algolia private and public ecommerce datasets
Algolia ecommerce benchmark using just Algolia private datasets is curated from Algolia customers that collaborated to evaluate the model performance for their specific data distribution. Algolia private ecommerce datasets include query and document sentences as pair and an associated relevancy score, and dominant languages in the benchmark are English, French, German, and Spanish. The relevancy scores are derived from the analytics based on the past user interactions with the products for the associated queries. Removing public ecommerce datasets doesn’t impact the leaderboard significantly, yet Algolia v2412 becomes the second in the leaderboard leading the best Google embedding model on the leaderboard. All three Algolia LLMs are still ranked in the top 20 (Figure 3).
Figure 3: Algolia LLM performance on ecommerce benchmark that includes Algolia private ecommerce datasets
This benchmark keeps all datasets that are non-English, so dominant languages in this benchmark are French, German, and Spanish. Note that there are some other non-European languages, yet they are minority in numbers in this benchmark. Removing private English ecommerce datasets doesn’t impact the leaderboard significantly, and the Algolia English v2410 model drops from 12th place to 13th place in the leaderboard as expected (Figure 4). Also, there are some new models that climb in the top 20, namely Alibaba gte-large-en-1.5.
Figure 4: Algolia LLM performance on ecommerce benchmark that includes Algolia multilingual private ecommerce datasets
Algolia ecommerce benchmark includes ecommerce datasets in English only. Since the majority of Algolia customer datasets are in English, it is essential to benchmark Algolia LMMs against other open-source and commercially available counterparts. Dominant verticals in this benchmark are Fashion and Technology. Algolia v2412 model retains its place in the leaderboard (Figure 5), whereas v2410 English model climbs to the 9th place and v2410 multilingual falls from 9th place to 13th place. The Snowflake arctic-embed-m-v1.5 model jumps to the 5th place from 24th place. Lajavaness bilingual models end up in 32nd place for large and 35th place for large-8k models. Such drastic changes in the leaderboard shows how some LLM embedding models are sensitive to English. It is essential to evaluate sensitivity of models on any specific language to make sure they match the expectations of the deployment data distributions.
Figure 5: Algolia LLM performance on ecommerce benchmark that includes Algolia English private ecommerce datasets
Public ecommerce benchmark includes datasets such as ESCI, WANDS, Home Depot, Crowdflower, and Marqo. All these datasets with relevance scores are in English and publicly available. Due to their availability, it is no surprise many open-source models are trained on them. In the public English ecommerce benchmark, Algolia v2412 model is at 3rd place, whereas Algolia v2410 English model is in 6th and Algolia v2410 Multilingual model in 10th place (Figure 6). Algolia models perform consistently well across all ecommerce benchmarks. It is noticeable that Snowflakes models are gathered at the top of the leaderboard, indicating that they are trained strongly on public ecommerce datasets. OpenAI and Google models are no longer within the top 20, and they are clustered at around 30th place, which shows that these commercial models are not trained heavily on public ecommerce datasets. Cohere multilingual and English models are within top 20 in the leaderboard.
Figure 6: Algolia LLM performance on ecommerce benchmark that includes Public English ecommerce datasets
MTEB includes English retrieval datasets only. Algolia keeps MTEB to make sure Algolia models aren’t diverging much from their foundational language understanding given that Algolia LLMs are fine-tuned on top of state-of-the-art open-source LLM embedding models. Compared to the Private ecommerce benchmark, there are seven models displaced in MTEB, the most notable of which are Algolia v2410 Multilingual, Google gecko@003, and JinaAI embedding-v3 (Figure 7). Although these three models are highly ranked in private ecommerce datasets, they aren’t within the top 20 on the MTEB leaderboard. This is an indication that MTEB and any other open-source benchmarks aren’t enough to tell which LLM embedding models can be used for a specific context. It is essential to build an internal benchmark that includes datasets with deployment data distributions to ensure the model preferred is indeed the right one.
Figure 7: Algolia LLM performance on the MTEB benchmark that includes English retrieval datasets
It is essential to evaluate LLMs in the context of a specific problem. The performance difference of Google, JinaAI, and Algolia models on MTEB and Algolia ecommerce benchmarks shows how important it is to make sure internal benchmarks with relevant datasets are available for evaluation purposes. At Algolia, new LLMs are added to our benchmarks typically within 24 hours. This enables our AI team to select the state-of-the-art LLM embedding model, and fine-tune our next generation models based on foundational open-source models with permissive license using our fully automated training and evaluation pipelines. Algolia v2410 models are state-of-the-art for their size and use cases, and are now available under an MIT licence. Please use them and provide us with feedback at ai-research@algolia.com.
Powered by Algolia AI Recommendations