Search is the one feature every user touches and almost nobody thanks you for. Someone types a few keywords into a box, and either the right product is at the top of the list or it is not. For a long time at Fullscript it often was not, and our users were quick to tell us. Earlier this year, Dave Currie and I had the honour of standing on a stage at RubyConf in Las Vegas to talk about how we fixed that with a bit of classical machine learning, written entirely in Ruby. You can watch the full talk here: Teaching Ruby to Rank.
This post covers the same ground as the talk, but with more of the technical detail we had to leave out of a 30 minute slot, and a walkthrough of the gem we built along the way, which we are getting ready to release.
The Static Boosting Problem
Before 2024, search was probably the most complained about thing on our platform. At the time, our search was a single Elasticsearch query (We later migrated to Opensearch). The user's query was tokenized, matched against the name, description and ingredients fields, and scored with BM25. If you have not worked with a search engine before, BM25 is the default relevance formula in Lucene-based engines like OpenSearch and Elasticsearch. It is built on term frequency (how many times a token appears in a document) and inverse document frequency (how rare that token is across the whole index), so rare tokens that match are weighted much more heavily than common ones.
A simplified version of what that query looked like:
1{2 "query": {3 "bool": {4 "should": [5 { "match": { "name": { "query": "magnesium", "boost": 3.0 } } },6 { "match": { "description": { "query": "magnesium", "boost": 1.0 } } },7 { "match": { "ingredients": { "query": "magnesium", "boost": 1.5 } } },8 { "rank_feature": { "field": "popularity", "boost": 2.0 } },9 { "knn": { "embedding": { "vector": [0.12, -0.44, "..."], "k": 50, "boost": 1.0 } } }10 ]11 }12 }13}
Each of those subqueries has a static boost value that encodes how much we think that part of the query matters. The name probably matters more than the description, so it gets a bigger boost, and so on. The important word there is static. Those boosts are the same for every query anyone ever types.
Tuning it went like this. A bug report lands for a query like "magnesium". We adjust the name boost, maybe the popularity boost, maybe a token filter, and we make sure "magnesium" looks good. Then we vibe test a handful of other queries, because we just changed the weights for everything. However, since we are only human and we cannot test everything, other queries would get worse. Then the next bug report lands, and the loop restarts.
- magnesiumNDCG@50.55
- 1Calcium + Magnesium + Drelevance 1 of 4
- 2Magnesium Oxide 500relevance 2 of 4
- 3Daily Multivitaminrelevance 1 of 4
- 4Bisglycinate Calm 200 mgrelevance 4 of 4
- 5Herbal Sleep Tearelevance 0 of 4
Best match is buried at #4
Extremely unfriendly, inefficient search engine.
- vitamin cNDCG@50.96
- 1Vitamin C 1000 mgrelevance 4 of 4
- 2Vitamin C Gummiesrelevance 2 of 4
- 3Immune Defenserelevance 2 of 4
- 4Buffered Ester-Crelevance 3 of 4
- 5Citrus Bioflavonoidsrelevance 1 of 4
Best match is #1
- sleep supportNDCG@50.84
- 1Sleep Support Blendrelevance 3 of 4
- 2Melatonin 3 mgrelevance 4 of 4
- 3Valerian Root Capsulesrelevance 3 of 4
- 4Nighttime Relax Gummiesrelevance 2 of 4
- 5Daily Multivitaminrelevance 0 of 4
Best match is buried at #2
- omega 3NDCG@50.98
- 1Omega-3 Fish Oilrelevance 4 of 4
- 2Algae Omega Veganrelevance 2 of 4
- 3Krill Oil Softgelsrelevance 3 of 4
- 4Flaxseed Oilrelevance 2 of 4
- 5Cod Liver Oilrelevance 2 of 4
Best match is #1
Measuring "Good" Before Fixing It
The real problem with vibe testing is that we had no shared definition of good. Two engineers could look at the same result set and disagree about whether the change helped. So before touching the ranking again, we picked a metric: normalized discounted cumulative gain (NDCG), which is more or less the industry standard for search relevance.
NDCG has three parts, and the name spells them out backwards.
Gain is a grade for how relevant a single result is to the query. Take the query "vitamin c". A bottle of Vitamin C 1000mg is a perfect match. Ascorbic acid is the same thing under a different name, so it is also a perfect match. A zinc and vitamin C lozenge is relevant but not what you asked for. Vitamin D is not vitamin C at all, even if it matched on the token "vitamin". Where those grades come from is its own problem that the next section will cover.
Discount penalizes a result by its position. A great result at rank one is worth a lot more than a great result at rank seven, because people would rather not scroll down the page to find their ideal result. The real formula is DCG = 1 / log₂(rank + 1), but for a worked example a simple 1 / rank makes the point:
Rank | Product | Grade | Discount | Discounted gain |
|---|---|---|---|---|
1 | Vitamin C 1000mg | 4 | 1.00 | 4.00 |
2 | Vitamin D3 5000 IU | 0 | 0.50 | 0.00 |
3 | Ascorbic Acid Powder | 4 | 0.33 | 1.33 |
4 | Zinc + Vitamin C Lozenges | 3 | 0.25 | 0.75 |
5 | Omega-3 Fish Oil | 0 | 0.20 | 0.00 |
6 | Citrus Cold Remedy | 2 | 0.17 | 0.33 |
7 | Daily Multivitamin | 1 | 0.14 | 0.14 |
Cumulative means we add the discounted gains up. This result set has a DCG of 6.56.
That number is unbounded and depends on how many relevant products exist for a query, so it is hard to compare across queries. The normalized part fixes that. Sort the same seven products by grade, best first, and compute the DCG of that ideal ordering: 4 + 2 + 1 + 0.5 + 0.2 + 0 + 0 = 7.7. Divide the actual by the ideal and you get an NDCG of 6.56 / 7.7 = 0.85. Perfect ranking is 1.0. Now every query is on the same scale, and "did this change help" becomes a question with a measurable answer.
- 1Vitamin C 1000mg41.004.00
- 2Vitamin D3 5000 IU00.500.00
- 3Ascorbic Acid Powder40.331.33
- 4Zinc + Vitamin C Lozenges30.250.75
- 5Omega-3 Fish Oil00.200.00
- 6Citrus Cold Remedy20.170.33
- 7Daily Multivitamin10.140.14
Judgment Lists From Clicks
A graded result set for a query is called a judgment list, and to compute NDCG you need enough of them to be a set representative of your real world searches. There are two ways to get grades. Explicit judgments come from a human rater, or these days an LLM, looking at a query and a product and assigning a score. We did this for a while when we were starting out, and it is a fine way to bootstrap. Implicit judgments come from user behaviour, and that is where we ended up.
Our users run tens of thousands of queries a day, and a lot of them are the same queries. If we log the query, the ordered result set, and every view, add to cart and purchase that came from it, then we can turn that behaviour into grades for thousands of queries without anyone rating anything by hand. We did not have to invent a format for this. OpenSearch's User Behavior Insights (UBI) plugin defines a standard schema for exactly these two things, a queries index and an events index joined by a query id, and we adopted it wholesale. It meant our logging matched what a wider search community was already building tooling around, and it gave us one less thing to argue about.
Going from raw events to grades takes a few steps:
- Join events to queries. Every event carries the id of the query that produced it, so we know which result set a click came from and what position the product was in.
- Decide what the user actually looked at. Within one search session we assume the user scanned the list top to bottom and stopped somewhere. The deepest position with any action is that stopping point, and every product above it counts as an impression. Products below it were probably never seen, so they are dropped rather than counted as ignored. This is called a cascade examination model, and it is the main thing standing between you and a training set that says everything at rank 30 is terrible.
- Aggregate across sessions. Within one session a product gets at most one of each action, so a user who opens the same product three times counts as one view. Then sum impressions, views, add to carts and purchases per query and product across every session.
- Turn counts into a rate, and weight it. A purchase has a higher intent signal than a view, so the grade is a weighted sum of conversion rates: click-through rate, add to cart rate and purchase rate. We want these grades to correlate with business outcomes, not just engagement.
- Smooth it. A product with one impression and one purchase is not a 100% converter, it is a product we know nothing about. Each rate is smoothed with a Beta-Binomial prior that pulls low-impression products toward the dataset-wide average and only lets them move away as the evidence accumulates.
One session: search “magnesium”
Click a row to cycle its action. The deepest action sets the stopping point, and a purchase also counts as an add to cart and a view.
Smoothing: cvr with a Beta-Binomial prior
(1 + 6) / (1 + 6 + 144) = 0.0464
α = 0.04 × 150, β = 0.96 × 150
label = Σ weight × smoothed rate: ctr ×1, atcr ×2, cvr ×3
In the gem, that whole label definition is easily configurable in the gem:
1dataset_builder:2 conversion_rate_targets:3 - name: ctr4 numerator: view_count5 weight: 16 nu: 507 - name: atcr8 numerator: add_to_cart_count9 weight: 210 nu: 10011 - name: cvr12 numerator: purchase_count13 weight: 314 nu: 150
Each rate is (numerator + α) / (impression_count + α + β), where α and β are derived from the global rate for that action and nu, the prior strength in pseudo-observations. The label for a row is Σ weight × smoothed_rate. By default we take the square root of that sum before training, but that transformation is configurable. These judgment lists are now two things at once: the way we measure our search engine, and the training data for a model that can rank better than our static boosts ever could.
Learning to Rank
Learning to rank (LTR) is the technique of training a model to order search results for a query, using judgment lists as the ground truth. Our OpenSearch query retrieves the initial result set, which we pass to our LTR model to improve the ordering. The whole loop:
Queries, result sets and user events, stored in the UBI schema. The one part nobody can package up for you.
Ltr::Ubi.query(...)Ltr::Ubi.event(...)Features
Features are what the model actually sees for each query and product pair. We split them into two groups by when they are computed.
Search-time features depend on the query. The BM25 score of the name subquery, the score of the ingredients subquery, the overall relevancy score, the length of the query. These can only be computed when the query runs, so we ask OpenSearch to return them alongside the results.
Feature store features do not depend on the query. Price, popularity, reorder rate, how many times a product has been purchased. These are the same whether you searched "magnesium" or "sleep support", so we compute them at deployment time and store them in a feature store keyed by product id. At inference we look them up instead of recomputing them, which keeps reranking fast.
There is one more wrinkle. BM25 scores are unbounded. A score of 12 might be excellent for one query and mediocre for another, which makes it hard for a model to learn what a "bad" score looks like. So for every bm25_* feature, and for the overall relevancy score, we automatically derive two more per query: the product's rank by that score within the result set, and its score relative to the top score. bm25_name becomes bm25_name, bm25_name_rank and bm25_name_relative_score. A product that did not match on that field at all has no score to rank, so it gets a rank of -1 rather than a NaN, which lets the model treat "did not match the name" as its own signal. Those derived features are usually among the most important ones in the final model.
Query “magnesium”
Query “sleep support”
Raw BM25 is unbounded, so 12.4 for “magnesium” and 4.2 for “sleep support” are both the best score for their query, and not comparable.
The model
The model is XGBoost, trained with the rank:ndcg objective. If you have not run into it, XGBoost builds an ensemble of decision trees. Each node in a tree splits on one feature at one threshold, a product's features walk it down to a leaf, and the leaves across all the trees sum to a score. The ranking objective means it is not trying to predict the grade of any single product, it is trying to get the order of products within a query right, and it is optimizing NDCG directly while it does so. The objective is a config, so rank:pairwise or rank:map are a one-line change, but the pipeline assumes a ranking objective throughout: the data is grouped by query, the metrics are ranking metrics, and the lambdarank_* parameters only mean something to the lambdarank family.
Every feature and hyperparameter for the model lives in the gem config:
1model:2 type: xgboost3 name: fs-ltr-us-patient4 hyperparameters:5 objective: rank:ndcg6 eval_metric: ndcg@127 lambdarank_unbiased: true8 n_estimators: 5009 learning_rate: 0.110 max_depth: 611 min_child_weight: 112 subsample: 0.513 colsample_bytree: 0.614 early_stopping_rounds: 201516search_time_features:17 - relevancy18 - bm25_name19 - bm25_description20 - bm25_ingredients2122feature_engineering:23 - rank24 - relative_score2526feature_store:27 path: "tmp/ltr/feature_store.json"28 include_in_deployment: true29 features:30 - price31 - purchase_count32 - purchase_conversion
Optimization
Two things determine how good the model is: which features it gets, and the hyperparameters it is trained with. We tune both automatically before training.
For features, we run greedy backward feature selection. Start with every candidate feature, train a model, then try dropping each feature one at a time and keep the drop that helps most. Repeat until no drop helps or you hit a configured minimum. A small complexity penalty, subtracted per feature, means a feature has to earn its place: a drop that costs less NDCG than the penalty saves still counts as an improvement, since every feature is one more thing to compute at query time. There are two cheaper strategies in the gem, rfe and importance_ranked, which rank features by importance and sweep a cut point instead of trying every drop. They cost one training per feature instead of one per feature per step, which matters once you have many features.
For hyperparameters, we use Bayesian optimization. Rather than grid searching every combination of tree depth, learning rate and subsample ratio, etc., it fits a Gaussian process to your initial sample of hyperparameter values, then uses expected improvement to pick the next most promising point in the search space. This process lets you optimize your hyperparameter values much faster than grid search. The Gaussian process, the Cholesky decomposition under it, and the acquisition function are all implemented in Ruby in the gem.
1optimization:2 metric_to_optimize: ndcg@123 min_features: 44 complexity_penalty: 0.015 feature_selection:6 strategy: greedy_backward7 n_init: 58 iterations: 509 hyperparameter_search_space:10 max_depth: { min: 3, max: 10 }11 learning_rate: { min: 0.1, max: 0.3 }12 min_child_weight: { min: 1, max: 10 }13 subsample: { min: 0.3, max: 0.9 }14 colsample_bytree: { min: 0.5, max: 0.9 }
Splits and the deploy gate
The dataset is split three ways at the query level, never at the product level, so a query's products are always in the same split. Training fits on train. The optimizer scores its candidates on val. The final reported metrics come from test. One honest footnote: XGBoost's early stopping watches training.early_stopping_split, which defaults to test, so the number of trees is the one thing test does get a say in. Features and hyperparameters, the choices with real room to overfit, never see it. We did not always have this separation. For a while a single test split was doing all three jobs, which meant the NDGC score we were gating deploys on was from the same data set we optimized our features and hyperparameters against. It looked great, but it was optimistic.
1dataset_builder:2 split_ratios:3 train: 0.74 val: 0.155 test: 0.1567deploy:8 enabled: true9 metric: ndcg10 minimum_metric_value: 0.5
After training, the model is evaluated on the test split. If its NDCG clears the threshold, it gets deployed. If not, nobody gets paged, it just does not ship, and we go investigate. The pipeline also scores the validation split and logs the gap against test, because a wide gap is a good sign that the tuning captured noise instead of ranking quality. The metrics that need a yes-or-no notion of relevance, like map@10 and precision@5, use a percentile cut on the label rather than a fixed grade, because a smoothed conversion rate almost never reaches a round number you could name. evaluation.min_relevance_percentile: 60 means at most the top 40% of labels in the split count as relevant.
Reranking at search time
Putting it together for the query "magnesium": OpenSearch runs the initial retrieval and returns the top 50 products along with each product's search-time features. We look up the feature store features for those 50 product ids, derive the rank and relative score features within this result set, and hand the model 50 rows. It returns 50 scores. We sort by score and return the list. The model does not know what magnesium is. It knows that for this query, this product had the top BM25 name score, a relative ingredients score of 0.9, a strong purchase conversion rate, and a mid-range price, and that products with those properties tended to get bought.
OpenSearch returns the top 50 by BM25. Showing 8 here, in BM25 order.
- 1Magnesium Lotion13.2$182.1%10.600.62–
- 2Magnesium Citrate Powder12.6$227.4%21.001.52–
- 3Magnesium Glycinate 120ct11.9$2611.8%30.952.31–
- 4Magnesium L-Threonate10.7$386.1%40.901.74–
- 5Magnesium + B6 Sleep Blend9.4$218.9%50.851.97–
- 6Cal-Mag Liquid7.8$174.2%60.550.88–
- 7Multimineral Complex5.1$293.3%70.400.41–
- 8Electrolyte Drink Mix3.6$155.0%80.350.57–
Illustrative data
The Gem
Everything above started life as a library inside our Rails monolith, wired to our OpenSearch cluster and our AWS account. We are extracting it into a gem, ltr, that covers the full lifecycle: collecting behavioural data, aggregating it into a dataset, optimizing, training, deploying and reranking. It requires Ruby 3.3+ and is built on xgb, rover-df, numo-narray and ActiveSupport. The AWS SDKs for S3, SSM and SageMaker, the OpenSearch client and SQLite are all dependencies (for now), so bundle install pulls them in, but no client is created until you use the feature that needs it. You can build a dataset from a local OpenSearch instance, then train, and rerank against a model on disk without an AWS account.
We are putting the finishing touches on the gem now and will link it here as soon as it is released. In the meantime, here is what using it looks like.
Installation and host configuration
1gem "ltr"
Ltr.configure wires the gem to your app. Inside a Rails app the defaults pick up Rails.logger and build an S3 client from whatever AWS region and credentials are in the environment. What they do not pick up is Rails.root: root defaults to the working directory, and every relative path in your config (the override file, the dataset, the local model) resolves against it. A rake task run by cron does not always start where you think it does, so we set it. error_reporter is an optional callable that gets handed errors the pipelines recover from, so they land in your error tracker instead of only in the log:
1Ltr.configure do |config|2 config.logger = Rails.logger3 config.s3_client = Aws::S3::Client.new(region: "us-east-1")4 config.root = Rails.root.to_s5 config.error_reporter = ->(error, context) { Sentry.capture_exception(error, extra: context) }6end7
Pipeline settings live in YAML. The gem ships a complete default config, and you layer an override file on top of it per model. Config files are ERB-evaluated, so secrets can come from the environment. The defaults are our defaults, though, all the way down to our S3 bucket and our SageMaker execution role in the deploy block, so there is a handful of keys you will always override: the model name, where the dataset lives, how to reach OpenSearch, which features exist in your index, and, if you deploy to SageMaker, the bucket, the execution role and the region. Here is the smallest override we would actually run:
1model:2 name: fs-ltr-us-patient3 hyperparameters: # the optimizer writes here4 objective: rank:ndcg5 eval_metric: ndcg@126 max_depth: 67 learning_rate: 0.189ubi_data_loader:10 max_days_ago: 3011 n_events: 50000012 opensearch:13 url: <%= ENV["OPENSEARCH_URL"] %>14 username: <%= ENV["OPENSEARCH_USERNAME"] %>15 password: <%= ENV["OPENSEARCH_PASSWORD"] %>16 index:17 queries: ubi_queries18 events: ubi_events1920search_time_features: # and here21 - relevancy22 - bm25_name23 - bm25_description24 - bm25_ingredients2526feature_store:27 features: # and here28 - price29 - purchase_count30 - purchase_conversion3132optimization:33 search_time_features: # the pool feature selection draws from34 - relevancy35 - bm25_name36 - bm25_description37 - bm25_ingredients38 feature_store_features:39 - price40 - reorder_rate41 - purchase_count42 - purchase_conversion4344training:45 modelling_dataset_path: "s3://my-bucket/ltr/us_patient/dataset.csv"46 s3_region: us-east-14748dataset_enricher:49 class: "Search::Reranking::DatasetEnricher"5051deploy:52 enabled: true53 sagemaker_bucket: my-bucket54 execution_role: <%= ENV["SAGEMAKER_EXECUTION_ROLE"] %>55 instance_type: ml.m5.large56 metric: ndcg57 minimum_metric_value: 0.5
1config = Ltr::ConfigLoader.new(override_config_path: "config/ltr/us_patient.yml")
Everything else falls back to the gem defaults. Loading validates the merged config and raises Ltr::ConfigurationError naming the key when something is off: split ratios that do not sum to one, a deploy.metric that is not in evaluation.metrics, local: true without a local_path. It does not check feature names against your dataset, metric names against the ones it knows, or whether you remembered to replace our bucket with yours, so those surface later, and we will come back to how. We run anywhere from one to four models per search engine, split by things like geography and user type, and each one is just another override file.
Step 1: Collect the data
This is the part the gem cannot do for you, and honestly it is one of the hardest parts of the whole project. You need to log every query with its ordered results and scores, and every user action against those results, with the query id that produced them. We use the OpenSearch User Behavior Insights (UBI) schema for this, and the gem's dataset pipeline reads that schema directly.
The two indices need three things from their mappings: timestamp as a date, because the loader pages through events by it; query_id and action_name as keyword, because they are filtered exactly; and query_attributes stored but not indexed, because its scores object is read straight out of _source and its key order is the displayed rank, which is not something you want an analyzer anywhere near:
1PUT ubi_queries2{3 "mappings": {4 "properties": {5 "query_id": { "type": "keyword" },6 "user_query": { "type": "text" },7 "timestamp": { "type": "date" },8 "client_id": { "type": "keyword" },9 "application": { "type": "keyword" },10 "query_attributes": { "type": "object", "enabled": false }11 }12 }13}1415PUT ubi_events16{17 "mappings": {18 "properties": {19 "query_id": { "type": "keyword" },20 "action_name": { "type": "keyword" },21 "timestamp": { "type": "date" },22 "event_attributes": {23 "properties": { "product_id": { "type": "keyword" } }24 }25 }26 }27}
Ltr::Ubi builds documents in the right shape. It does not write anything, you index the hash with whatever client you already have. It reads the id field names from the gem's default config rather than from your override, and a couple of places in the pipeline assume the defaults too, so leave dataset_builder.query_id_field and doc_id_field at query_id and product_id:
1# After running the initial retrieval query2scores = hits.to_h do |hit|3 [hit["_id"], hit["fields"].slice("relevancy", "bm25_name", "bm25_description", "bm25_ingredients")]4end56query_doc = Ltr::Ubi.query(7 query_id: query_id,8 user_query: params[:q],9 timestamp: Time.current,10 scores: scores, # insertion order is the displayed rank11 application: "web",12 client_id: current_user.id.to_s # ids are strings, the gem will tell you if not13)14opensearch.index(index: "ubi_queries", body: query_doc)
1# When the user does something with a result2event_doc = Ltr::Ubi.event(3 query_id: query_id,4 action_name: "add_to_cart", # or "view", "purchase", ...5 doc_id: product.id.to_s,6 timestamp: Time.current7)8opensearch.index(index: "ubi_events", body: event_doc)
The scores hash values are the search-time features the model will train on, and the position each product was displayed at is what the cascade examination model needs. The action names you log are the ones you reference in conversion_rate_targets, and only those. An action you log but never reference, say add_to_favorites, does not move the cascade cutoff, and a query whose only events are unreferenced actions produces no training rows at all. If you want an action to count, give it a target, even a low-weight one.
One more thing about identity. Judgment lists are grouped by the normalized query text, stripped and downcased, so "Magnesium " and "magnesium" are one query. If the same text means different things in different contexts, we have separate catalogs per country, name the query document fields that should split them in dataset_builder.additional_qid_fields.
Step 2: Build the dataset
1result = Ltr::DatasetAggregation::DatasetPipeline.call(config)2result[:rows] # => 1843023result[:dataset_path] # => "s3://my-bucket/ltr/us_patient/dataset.csv.gz"
This pulls the newest n_events events from the last max_days_ago days, fetches the queries those events reference, joins them, runs the cascade examination and aggregation described above, computes labels, derives the engineered features, assigns splits, and writes a gzipped CSV to training.modelling_dataset_path, local or S3. n_events is the knob that sizes your dataset, not the day count, and it defaults to a modest 10,000, so set it. And the path you configure ends in .csv; the file that gets written is that path plus .gz. The pipeline raises on anything it cannot recover from, an empty window, an enricher class that does not exist, so the result you get back is always a success. It carries rows, columns, enriched and dataset_path.
If your products have features that do not live in the search index, like price or a reorder rate from your orders table, you add them here with an enricher. Subclass DatasetEnricher, return a dataframe keyed by product_id, and name the class in config. The join key is cast to an integer on both sides, so your product ids need to be numeric today; string keys like UUIDs are on the list:
1module Search2 module Reranking3 class DatasetEnricher < ::Ltr::DatasetAggregation::DatasetEnricher4 def fetch_product_features(unique_product_ids)5 rows = Product.where(id: unique_product_ids).pluck(:id, :price, :reorder_rate)6 Rover::DataFrame.new(7 "product_id" => rows.map { |id, _, _| id },8 "price" => rows.map { |_, price, _| price.to_f },9 "reorder_rate" => rows.map { |_, _, rate| rate.to_f }10 )11 end12 end13 end14end
Products your frame leaves out get nil in those columns, which XGBoost reads as missing.
A few things to know before you open the CSV. Queries with fewer than dataset_builder.min_documents_per_query products (three by default) are dropped, because there is nothing to rank. The view_count, add_to_cart_count and purchase_count columns are per-product totals across every query, because they are features the model will see again at inference time; the per-query counts the label was computed from are not in the file, and neither is impression_count. If you need to explain one row's label, dataset_builder.debug_mode: true keeps those, and you must never train on the result. Splits are assigned with a fixed seed, so rebuilding from the same events gives the same split.
The default aggregation engine does all of this in memory, which is fine for a few hundred thousand events. Once a pull no longer fits in RAM, one line of config switches it to an out-of-core engine that stages the download in a scratch SQLite file and aggregates there, bounding peak memory by the OpenSearch page size instead of the size of the pull:
1dataset_builder:2 aggregation_engine: sqlite
Both engines produce the same labels and features for the same events. The SQLite engine filters events down to the cascade actions on the OpenSearch side, so n_events counts only those, and unreferenced actions never reach the dataset. Its scratch database goes under dataset_aggregation.staging_root, the system temp dir by default, needs room for the whole uncompressed pull, and is deleted whether the run succeeds or fails. We wrote that one because our training servers were, to quote the talk, beefy, and Ruby is not shy about memory.
Step 3: Optimize
1result = Ltr::Optimize::OptimizePipeline.call(config)2result[:features] # => { search_time_features: [...], feature_store_features: [...] }3result[:hyperparameters] # => { objective: "rank:ndcg", max_depth: 7, learning_rate: 0.18, ... }
This loads the dataset, runs feature selection over optimization.search_time_features and optimization.feature_store_features, runs Bayesian optimization over hyperparameter_search_space, and writes the winners back. It rewrites your override file in place, replacing the search_time_features, feature_store.features and model.hyperparameters blocks. It replaces, it does not insert: a block that is not in the file is skipped without a word, and the run's result then lives only in the log. That is why the override above carries all three, even though two of them would have fallen back to the defaults anyway. Comments outside those blocks survive; comments inside them do not, since the lines they annotated are gone. With ssm.enabled: true it also writes them to AWS SSM Parameter Store so the training pipeline can pick them up from there. Every trial fits on train, early-stops on training.early_stopping_split, and is scored on optimization.split, which is val.
A few things about the search space. A parameter is treated as an integer when both bounds are integers, so max_depth: { min: 3, max: 10 } is rounded and clamped, and min: 3.0 would quietly make it a float. Ranges whose upper bound is at least twenty times the lower are searched on a log scale. There are no categorical parameters. n_init is how many Latin hypercube points seed the Gaussian process before expected improvement takes over, and the loop stops early if ten iterations in a row fail to improve. A feature in your pools that is not in the dataset gets a warning and is dropped from the search.
This is the expensive step. It trains a lot of models. We run it daily, but weekly would be fine for most, and you can skip it entirely and train with the hand-set features and hyperparameters in your config.
Step 4: Train and deploy
1result = Ltr::Train::ModellingPipeline.call(config)23result[:metrics]4# => { "map@10" => 0.71, "mrr@10" => 0.79, "ndcg" => 0.87, "ndcg@5" => 0.83, "ndcg@12" => 0.85,5# "precision" => 0.34, "precision@5" => 0.52, "recall@5" => 0.61 }67result[:optimization_split_metrics] # the same metrics on val, for the gap89result[:deployed] # => true, the SageMaker gate10result[:deployment] # => { deployed: true, reason: :threshold_met, metric: "ndcg", score: 0.87, threshold: 0.5, endpoint: {...} }
Training reads the dataset, trains the XGBoost ranker with the optimized features and hyperparameters, evaluates it on the test split with every metric in evaluation.metrics, and then checks the gate. The metric names are ndcg, map, mrr, precision and recall, with or without an @k.
The failure mode to watch for here is quiet. A feature named in config but missing from the dataset, because of a typo or because the enricher did not produce it, is logged and dropped, and training carries on without it. It only raises if no features are left. Check the training log for dropped features the first time you run a new config.
If deploy.enabled is on and deploy.metric clears deploy.minimum_metric_value, the model ships: the booster, its feature list, the feature store JSON and the Python inference handler are tarred up and uploaded to deploy.sagemaker_bucket, and the endpoint <model.name>-endpoint is created or updated to serve them. The pipeline blocks until the endpoint is InService, a few minutes for an update, so budget for that in whatever schedules it. On the first run, the container image is looked up by training.s3_region and only the four US regions are mapped, and the tarball is built inside the gem's install directory, so that path has to be writable. The feature store is built from the dataset you just trained on, so its values are as fresh as your last dataset build.
The credentials running the pipeline need to write to the bucket, create and update SageMaker models, endpoint configs and endpoints, and pass the execution role; whatever calls the reranker at query time needs sagemaker:InvokeEndpoint. The endpoint is a real instance billing around the clock, so instance_type is a cost decision as much as a latency one.
For local development, or if you would rather not run SageMaker, there is a local deploy mode that saves the booster to disk:
1deploy:2 enabled: false3 local: true4 local_path: "tmp/ltr/models/us_patient.ubj"
Local deploy is independent of the SageMaker gate and does not have one of its own: if local is on, the booster is written, whatever the score. It reports under its own key, and result[:deployed] stays false because that one is the SageMaker answer:
1result[:deployed] # => false2result[:local_deployment] # => { deployed: true, reason: :local_saved, path: "tmp/ltr/models/us_patient.ubj" }
Step 5: Rerank
The inference payload is a list of items, each with an id and the fields the model needs:
1items = hits.map do |hit|2 {3 "id" => hit["_id"],4 "fields" => [5 { "name" => "relevancy", "value" => hit["_score"] },6 { "name" => "bm25_name", "value" => hit["fields"]["bm25_name"] },7 { "name" => "bm25_description", "value" => hit["fields"]["bm25_description"] },8 { "name" => "bm25_ingredients", "value" => hit["fields"]["bm25_ingredients"] },9 ]10 }11end
The SageMaker client takes that list wrapped in an items key, because that is the JSON the endpoint expects. The local reranker takes the bare list. Ltr::Inference::Base picks the backend for you, local when deploy.local is on and SageMaker otherwise, but it passes your payload through untouched, so hand it the shape of the backend it is going to pick:
1response = Ltr::Inference::SagemakerClient.call({ "items" => items }, config)2response = Ltr::Inference::LocalReranker.call(items, config)34# => {5# "items" => [6# { "item" => "prod_812", "score" => 2.31 },7# { "item" => "prod_104", "score" => 1.97 },8# { "item" => "prod_559", "score" => 0.42 },9# ...10# ],11# "took_seconds" => 0.00612# }1314reranked_ids = response["items"].map { |item| item["item"] }
Items come back sorted by score, ids as strings. Either way, the _rank and _relative_score features are derived on the fly from the fields you send, using the model's own feature list to figure out which ones it needs. A field you leave out reaches XGBoost as NaN, which it handles as missing rather than failing. A field you send with a nil value does not: locally it becomes 0.0, so drop the field rather than sending nil. On SageMaker the handler first tries the feature store shipped in the tarball for any field you did not send, so you only have to send the search-time features, and a value you do send always wins. Locally there is no feature store lookup, so merge the feature store features into the fields yourself.
The local reranker caches the booster per path for the life of the process. If you redeploy a new .ubj under a running app, call Ltr::Inference::LocalReranker.reset! or restart.
The SageMaker client has aggressive default timeouts, half a second to connect and half a second to read, configurable under inference.sagemaker. A reranker that times out should fall back to the original order, not hold up the page, and that fallback is your job, not the gem's: a timeout raises, and the caller catches it. Ours looks like this:
1def rerank(hits)2 response = Ltr::Inference::SagemakerClient.call({ "items" => items_for(hits) }, config)3 order = response["items"].map { |item| item["item"] }4 hits.sort_by { |hit| order.index(hit["_id"]) || order.size }5rescue Seahorse::Client::NetworkingError, Aws::SageMakerRuntime::Errors::ServiceError => e6 Ltr.config.report_error(e, query: params[:q])7 hits8end
Note that the AWS SDK's default retry policy still applies to the runtime client, so a bad half second can turn into a couple of them before that rescue fires. Budget for that or turn retries down on the client.
Running it on a schedule
The gem is a library, not a daemon. There is no scheduler in it. What it does ship is a rake task per pipeline. Each one takes the override file either as a task argument or from the environment, and the argument wins if you pass both:
1bin/rails "ltr:aggregate[config/ltr/us_patient.yml]"2bin/rails "ltr:optimize[config/ltr/us_patient.yml]"3bin/rails "ltr:train[config/ltr/us_patient.yml]"45# or, which reads better in a crontab6LTR_CONFIG=config/ltr/us_patient.yml bin/rails ltr:train
aggregate logs the row count and where the dataset landed, optimize logs the features and hyperparameters it chose, and train logs the metrics and the reason on each deploy gate. In our app, cron runs the three nightly, in that order, once per model, and spinning up a new model for a new search engine is a new override file and three more cron entries.
The first time through, the things worth checking are the row count from aggregate (zero or tiny means the mapping or the action names are off), the dropped-feature lines from train, and the deploy reasons, which will be one of :disabled, :below_threshold or :threshold_met for SageMaker and :not_configured or :local_saved for local.
Did It Work?
Our users have not sent many messages saying "good job on search". What we have instead is a Slack channel dedicated to negative search feedback that has gone almost silent. We get maybe one message a month, and our NDCG dashboards have moved in the direction we want them to. In search, silence is the compliment.
The first release was the biggest single jump, but we have been iterating since 2024 and the gains have kept coming. Having judgment lists changed how we work as much as it changed the results. Every change to retrieval, whether it is a new synonym filter or a tweak to query understanding, gets run through the judgment lists offline for an NDCG score before it goes near an A/B test. We go into production informed instead of hopeful.
Conclusion
Learning to rank is not new, but doing it end to end in Ruby without a Python sidecar felt new to us, and it worked. If I had to boil the whole project down: a shared metric beats gut feel, your best training data is probably something you already collect, and you do not have to leave Ruby to do machine learning.
Thank you for reading this post! If you want the version with all the diagrams and a couple of bad jokes, watch the talk. Dave and Lucas have also written about the other stages of our retrieval pipeline, query understanding and personalization, here on Builder's Corner.

-2.png%3F2026-02-25T16%3A15%3A02.533Z&w=3840&q=75)