Search is the one feature every user touches and almost nobody thanks you for. Someone types a few keywords into a box, and either the right product is at the top of the list or it is not. For a long time at Fullscript it often was not, and our users were quick to tell us. Earlier this year, Dave Currie and I had the honour of standing on a stage at RubyConf in Las Vegas to talk about how we fixed that with a bit of classical machine learning, written entirely in Ruby. You can watch the full talk here: Teaching Ruby to Rank.



This post covers the same ground as the talk, but with more of the technical detail we had to leave out of a 30 minute slot, and a walkthrough of the gem we built along the way, which we are getting ready to release.

The Static Boosting Problem

Before 2024, search was probably the most complained about thing on our platform. At the time, our search was a single Elasticsearch query (We later migrated to Opensearch). The user's query was tokenized, matched against the name, description and ingredients fields, and scored with BM25. If you have not worked with a search engine before, BM25 is the default relevance formula in Lucene-based engines like OpenSearch and Elasticsearch. It is built on term frequency (how many times a token appears in a document) and inverse document frequency (how rare that token is across the whole index), so rare tokens that match are weighted much more heavily than common ones.

A simplified version of what that query looked like:

1{
2 "query": {
3 "bool": {
4 "should": [
5 { "match": { "name": { "query": "magnesium", "boost": 3.0 } } },
6 { "match": { "description": { "query": "magnesium", "boost": 1.0 } } },
7 { "match": { "ingredients": { "query": "magnesium", "boost": 1.5 } } },
8 { "rank_feature": { "field": "popularity", "boost": 2.0 } },
9 { "knn": { "embedding": { "vector": [0.12, -0.44, "..."], "k": 50, "boost": 1.0 } } }
10 ]
11 }
12 }
13}


Each of those subqueries has a static boost value that encodes how much we think that part of the query matters. The name probably matters more than the description, so it gets a bigger boost, and so on. The important word there is static. Those boosts are the same for every query anyone ever types.

Tuning it went like this. A bug report lands for a query like "magnesium". We adjust the name boost, maybe the popularity boost, maybe a token filter, and we make sure "magnesium" looks good. Then we vibe test a handful of other queries, because we just changed the weights for everything. However, since we are only human and we cannot test everything, other queries would get worse. Then the next bug report lands, and the loop restarts.

Whack-a-Query
1open bug report
  • magnesium
    NDCG@50.55
    1. 1Calcium + Magnesium + Drelevance 1 of 4
    2. 2Magnesium Oxide 500relevance 2 of 4
    3. 3Daily Multivitaminrelevance 1 of 4
    4. 4Bisglycinate Calm 200 mgrelevance 4 of 4
    5. 5Herbal Sleep Tearelevance 0 of 4

    Best match is buried at #4

    Extremely unfriendly, inefficient search engine.
  • vitamin c
    NDCG@50.96
    1. 1Vitamin C 1000 mgrelevance 4 of 4
    2. 2Vitamin C Gummiesrelevance 2 of 4
    3. 3Immune Defenserelevance 2 of 4
    4. 4Buffered Ester-Crelevance 3 of 4
    5. 5Citrus Bioflavonoidsrelevance 1 of 4

    Best match is #1

  • sleep support
    NDCG@50.84
    1. 1Sleep Support Blendrelevance 3 of 4
    2. 2Melatonin 3 mgrelevance 4 of 4
    3. 3Valerian Root Capsulesrelevance 3 of 4
    4. 4Nighttime Relax Gummiesrelevance 2 of 4
    5. 5Daily Multivitaminrelevance 0 of 4

    Best match is buried at #2

  • omega 3
    NDCG@50.98
    1. 1Omega-3 Fish Oilrelevance 4 of 4
    2. 2Algae Omega Veganrelevance 2 of 4
    3. 3Krill Oil Softgelsrelevance 3 of 4
    4. 4Flaxseed Oilrelevance 2 of 4
    5. 5Cod Liver Oilrelevance 2 of 4

    Best match is #1

Every boost is shared by every query. Make "magnesium" happy and watch what happens to everyone else.


Measuring "Good" Before Fixing It

The real problem with vibe testing is that we had no shared definition of good. Two engineers could look at the same result set and disagree about whether the change helped. So before touching the ranking again, we picked a metric: normalized discounted cumulative gain (NDCG), which is more or less the industry standard for search relevance.

NDCG has three parts, and the name spells them out backwards.

Gain is a grade for how relevant a single result is to the query. Take the query "vitamin c". A bottle of Vitamin C 1000mg is a perfect match. Ascorbic acid is the same thing under a different name, so it is also a perfect match. A zinc and vitamin C lozenge is relevant but not what you asked for. Vitamin D is not vitamin C at all, even if it matched on the token "vitamin". Where those grades come from is its own problem that the next section will cover.

Discount penalizes a result by its position. A great result at rank one is worth a lot more than a great result at rank seven, because people would rather not scroll down the page to find their ideal result. The real formula is DCG = 1 / log₂(rank + 1), but for a worked example a simple 1 / rank makes the point:

Rank

Product

Grade

Discount

Discounted gain

1

Vitamin C 1000mg

4

1.00

4.00

2

Vitamin D3 5000 IU

0

0.50

0.00

3

Ascorbic Acid Powder

4

0.33

1.33

4

Zinc + Vitamin C Lozenges

3

0.25

0.75

5

Omega-3 Fish Oil

0

0.20

0.00

6

Citrus Cold Remedy

2

0.17

0.33

7

Daily Multivitamin

1

0.14

0.14

Cumulative means we add the discounted gains up. This result set has a DCG of 6.56.

That number is unbounded and depends on how many relevant products exist for a query, so it is hard to compare across queries. The normalized part fixes that. Sort the same seven products by grade, best first, and compute the DCG of that ideal ordering: 4 + 2 + 1 + 0.5 + 0.2 + 0 + 0 = 7.7. Divide the actual by the ideal and you get an NDCG of 6.56 / 7.7 = 0.85. Perfect ranking is 1.0. Now every query is on the same scale, and "did this change help" becomes a question with a measurable answer.

Rank "vitamin c" yourself
  1. 1Vitamin C 1000mg41.004.00
  2. 2Vitamin D3 5000 IU00.500.00
  3. 3Ascorbic Acid Powder40.331.33
  4. 4Zinc + Vitamin C Lozenges30.250.75
  5. 5Omega-3 Fish Oil00.200.00
  6. 6Citrus Cold Remedy20.170.33
  7. 7Daily Multivitamin10.140.14
Grades are fixed. Only the order changes. Drag or use the arrows to move products and watch DCG and NDCG respond.


Judgment Lists From Clicks

A graded result set for a query is called a judgment list, and to compute NDCG you need enough of them to be a set representative of your real world searches. There are two ways to get grades. Explicit judgments come from a human rater, or these days an LLM, looking at a query and a product and assigning a score. We did this for a while when we were starting out, and it is a fine way to bootstrap. Implicit judgments come from user behaviour, and that is where we ended up.

Our users run tens of thousands of queries a day, and a lot of them are the same queries. If we log the query, the ordered result set, and every view, add to cart and purchase that came from it, then we can turn that behaviour into grades for thousands of queries without anyone rating anything by hand. We did not have to invent a format for this. OpenSearch's User Behavior Insights (UBI) plugin defines a standard schema for exactly these two things, a queries index and an events index joined by a query id, and we adopted it wholesale. It meant our logging matched what a wider search community was already building tooling around, and it gave us one less thing to argue about.

Going from raw events to grades takes a few steps:

  1. Join events to queries. Every event carries the id of the query that produced it, so we know which result set a click came from and what position the product was in.
  2. Decide what the user actually looked at. Within one search session we assume the user scanned the list top to bottom and stopped somewhere. The deepest position with any action is that stopping point, and every product above it counts as an impression. Products below it were probably never seen, so they are dropped rather than counted as ignored. This is called a cascade examination model, and it is the main thing standing between you and a training set that says everything at rank 30 is terrible.
  3. Aggregate across sessions. Within one session a product gets at most one of each action, so a user who opens the same product three times counts as one view. Then sum impressions, views, add to carts and purchases per query and product across every session.
  4. Turn counts into a rate, and weight it. A purchase has a higher intent signal than a view, so the grade is a weighted sum of conversion rates: click-through rate, add to cart rate and purchase rate. We want these grades to correlate with business outcomes, not just engagement.
  5. Smooth it. A product with one impression and one purchase is not a 100% converter, it is a product we know nothing about. Each rate is smoothed with a Beta-Binomial prior that pulls low-impression products toward the dataset-wide average and only lets them move away as the evidence accumulates.
From clicks to grades

One session: search “magnesium”

Click a row to cycle its action. The deepest action sets the stopping point, and a purchase also counts as an add to cart and a view.

4 impressions2 views1 add to carts0 purchases

Smoothing: cvr with a Beta-Binomial prior

Raw rate100.0%
Smoothed rate4.6%

(1 + 6) / (1 + 6 + 144) = 0.0464
α = 0.04 × 150, β = 0.96 × 150

label = Σ weight × smoothed rate: ctr ×1, atcr ×2, cvr ×3

Everything above the deepest action counts as seen. Everything below is dropped, not punished. Low-volume products are pulled toward the average until evidence accumulates.


In the gem, that whole label definition is easily configurable in the gem:

1dataset_builder:
2 conversion_rate_targets:
3 - name: ctr
4 numerator: view_count
5 weight: 1
6 nu: 50
7 - name: atcr
8 numerator: add_to_cart_count
9 weight: 2
10 nu: 100
11 - name: cvr
12 numerator: purchase_count
13 weight: 3
14 nu: 150


Each rate is (numerator + α) / (impression_count + α + β), where α and β are derived from the global rate for that action and nu, the prior strength in pseudo-observations. The label for a row is Σ weight × smoothed_rate. By default we take the square root of that sum before training, but that transformation is configurable. These judgment lists are now two things at once: the way we measure our search engine, and the training data for a model that can rank better than our static boosts ever could.

Learning to Rank

Learning to rank (LTR) is the technique of training a model to order search results for a query, using judgment lists as the ground truth. Our OpenSearch query retrieves the initial result set, which we pass to our LTR model to improve the ordering. The whole loop:

Six steps, every night

Step 1 of 6

Collect

Queries, result sets and user events, stored in the UBI schema. The one part nobody can package up for you.

Ltr::Ubi.query(...)Ltr::Ubi.event(...)
Nightly via cron: aggregate → optimize → train
Collect, aggregate, featurize, optimize, train and deploy, rerank. Then repeat.

Features

Features are what the model actually sees for each query and product pair. We split them into two groups by when they are computed.

Search-time features depend on the query. The BM25 score of the name subquery, the score of the ingredients subquery, the overall relevancy score, the length of the query. These can only be computed when the query runs, so we ask OpenSearch to return them alongside the results.

Feature store features do not depend on the query. Price, popularity, reorder rate, how many times a product has been purchased. These are the same whether you searched "magnesium" or "sleep support", so we compute them at deployment time and store them in a feature store keyed by product id. At inference we look them up instead of recomputing them, which keeps reranking fast.

There is one more wrinkle. BM25 scores are unbounded. A score of 12 might be excellent for one query and mediocre for another, which makes it hard for a model to learn what a "bad" score looks like. So for every bm25_* feature, and for the overall relevancy score, we automatically derive two more per query: the product's rank by that score within the result set, and its score relative to the top score. bm25_name becomes bm25_name, bm25_name_rank and bm25_name_relative_score. A product that did not match on that field at all has no score to rank, so it gets a rank of -1 rather than a NaN, which lets the model treat "did not match the name" as its own signal. Those derived features are usually among the most important ones in the final model.

Making unbounded scores comparable

Query “magnesium”

productbm25_namerankrelative
Magnesium Glycinate 120ct12.4
Magnesium Citrate Powder9.8
Cal-Mag Liquid7.1
Multimineral Complex—
Magnesium Lotion3.3

Query “sleep support”

productbm25_namerankrelative
Sleep Support Blend4.2
Melatonin Gummies3.9
Night Calm Tea2.0
Valerian Root Capsules1.1
Magnesium Glycinate 120ct—

Raw BM25 is unbounded, so 12.4 for “magnesium” and 4.2 for “sleep support” are both the best score for their query, and not comparable.

Each bm25_* score gains a per-query rank and a score relative to the top hit. No match gets a rank of -1.


The model

The model is XGBoost, trained with the rank:ndcg objective. If you have not run into it, XGBoost builds an ensemble of decision trees. Each node in a tree splits on one feature at one threshold, a product's features walk it down to a leaf, and the leaves across all the trees sum to a score. The ranking objective means it is not trying to predict the grade of any single product, it is trying to get the order of products within a query right, and it is optimizing NDCG directly while it does so. The objective is a config, so rank:pairwise or rank:map are a one-line change, but the pipeline assumes a ranking objective throughout: the data is grouped by query, the metrics are ranking metrics, and the lambdarank_* parameters only mean something to the lambdarank family.

Every feature and hyperparameter for the model lives in the gem config:

1model:
2 type: xgboost
3 name: fs-ltr-us-patient
4 hyperparameters:
5 objective: rank:ndcg
6 eval_metric: ndcg@12
7 lambdarank_unbiased: true
8 n_estimators: 500
9 learning_rate: 0.1
10 max_depth: 6
11 min_child_weight: 1
12 subsample: 0.5
13 colsample_bytree: 0.6
14 early_stopping_rounds: 20
15
16search_time_features:
17 - relevancy
18 - bm25_name
19 - bm25_description
20 - bm25_ingredients
21
22feature_engineering:
23 - rank
24 - relative_score
25
26feature_store:
27 path: "tmp/ltr/feature_store.json"
28 include_in_deployment: true
29 features:
30 - price
31 - purchase_count
32 - purchase_conversion


Optimization

Two things determine how good the model is: which features it gets, and the hyperparameters it is trained with. We tune both automatically before training.

For features, we run greedy backward feature selection. Start with every candidate feature, train a model, then try dropping each feature one at a time and keep the drop that helps most. Repeat until no drop helps or you hit a configured minimum. A small complexity penalty, subtracted per feature, means a feature has to earn its place: a drop that costs less NDCG than the penalty saves still counts as an improvement, since every feature is one more thing to compute at query time. There are two cheaper strategies in the gem, rfe and importance_ranked, which rank features by importance and sweep a cut point instead of trying every drop. They cost one training per feature instead of one per feature per step, which matters once you have many features.

For hyperparameters, we use Bayesian optimization. Rather than grid searching every combination of tree depth, learning rate and subsample ratio, etc., it fits a Gaussian process to your initial sample of hyperparameter values, then uses expected improvement to pick the next most promising point in the search space. This process lets you optimize your hyperparameter values much faster than grid search. The Gaussian process, the Cholesky decomposition under it, and the acquisition function are all implemented in Ruby in the gem.

1optimization:
2 metric_to_optimize: ndcg@12
3 min_features: 4
4 complexity_penalty: 0.01
5 feature_selection:
6 strategy: greedy_backward
7 n_init: 5
8 iterations: 50
9 hyperparameter_search_space:
10 max_depth: { min: 3, max: 10 }
11 learning_rate: { min: 0.1, max: 0.3 }
12 min_child_weight: { min: 1, max: 10 }
13 subsample: { min: 0.3, max: 0.9 }
14 colsample_bytree: { min: 0.5, max: 0.9 }


Splits and the deploy gate

The dataset is split three ways at the query level, never at the product level, so a query's products are always in the same split. Training fits on train. The optimizer scores its candidates on val. The final reported metrics come from test. One honest footnote: XGBoost's early stopping watches training.early_stopping_split, which defaults to test, so the number of trees is the one thing test does get a say in. Features and hyperparameters, the choices with real room to overfit, never see it. We did not always have this separation. For a while a single test split was doing all three jobs, which meant the NDGC score we were gating deploys on was from the same data set we optimized our features and hyperparameters against. It looked great, but it was optimistic.

1dataset_builder:
2 split_ratios:
3 train: 0.7
4 val: 0.15
5 test: 0.15
6
7deploy:
8 enabled: true
9 metric: ndcg
10 minimum_metric_value: 0.5


After training, the model is evaluated on the test split. If its NDCG clears the threshold, it gets deployed. If not, nobody gets paged, it just does not ship, and we go investigate. The pipeline also scores the validation split and logs the gap against test, because a wide gap is a good sign that the tuning captured noise instead of ranking quality. The metrics that need a yes-or-no notion of relevance, like map@10 and precision@5, use a percentile cut on the label rather than a fixed grade, because a smoothed conversion rate almost never reaches a round number you could name. evaluation.min_relevance_percentile: 60 means at most the top 40% of labels in the split count as relevant.

Reranking at search time

Putting it together for the query "magnesium": OpenSearch runs the initial retrieval and returns the top 50 products along with each product's search-time features. We look up the feature store features for those 50 product ids, derive the rank and relative score features within this result set, and hand the model 50 rows. It returns 50 scores. We sort by score and return the list. The model does not know what magnesium is. It knows that for this query, this product had the top BM25 name score, a relative ingredients score of 0.9, a strong purchase conversion rate, and a mid-range price, and that products with those properties tended to get bought.

Reranking "magnesium"

OpenSearch returns the top 50 by BM25. Showing 8 here, in BM25 order.

ProductBM25 namePriceConversionName rankIngredientsScoreMove
  1. 1Magnesium Lotion13.2$182.1%10.600.62–
  2. 2Magnesium Citrate Powder12.6$227.4%21.001.52–
  3. 3Magnesium Glycinate 120ct11.9$2611.8%30.952.31–
  4. 4Magnesium L-Threonate10.7$386.1%40.901.74–
  5. 5Magnesium + B6 Sleep Blend9.4$218.9%50.851.97–
  6. 6Cal-Mag Liquid7.8$174.2%60.550.88–
  7. 7Multimineral Complex5.1$293.3%70.400.41–
  8. 8Electrolyte Drink Mix3.6$155.0%80.350.57–

Illustrative data

OpenSearch retrieves, the model reorders. Illustrative scores, real mechanics.


The Gem

Everything above started life as a library inside our Rails monolith, wired to our OpenSearch cluster and our AWS account. We are extracting it into a gem, ltr, that covers the full lifecycle: collecting behavioural data, aggregating it into a dataset, optimizing, training, deploying and reranking. It requires Ruby 3.3+ and is built on xgb, rover-df, numo-narray and ActiveSupport. The AWS SDKs for S3, SSM and SageMaker, the OpenSearch client and SQLite are all dependencies (for now), so bundle install pulls them in, but no client is created until you use the feature that needs it. You can build a dataset from a local OpenSearch instance, then train, and rerank against a model on disk without an AWS account.

We are putting the finishing touches on the gem now and will link it here as soon as it is released. In the meantime, here is what using it looks like.

Installation and host configuration

Gemfileruby
1gem "ltr"


Ltr.configure wires the gem to your app. Inside a Rails app the defaults pick up Rails.logger and build an S3 client from whatever AWS region and credentials are in the environment. What they do not pick up is Rails.root: root defaults to the working directory, and every relative path in your config (the override file, the dataset, the local model) resolves against it. A rake task run by cron does not always start where you think it does, so we set it. error_reporter is an optional callable that gets handed errors the pipelines recover from, so they land in your error tracker instead of only in the log:

config/initializers/ltr.rbruby
1Ltr.configure do |config|
2 config.logger = Rails.logger
3 config.s3_client = Aws::S3::Client.new(region: "us-east-1")
4 config.root = Rails.root.to_s
5 config.error_reporter = ->(error, context) { Sentry.capture_exception(error, extra: context) }
6end
7


Pipeline settings live in YAML. The gem ships a complete default config, and you layer an override file on top of it per model. Config files are ERB-evaluated, so secrets can come from the environment. The defaults are our defaults, though, all the way down to our S3 bucket and our SageMaker execution role in the deploy block, so there is a handful of keys you will always override: the model name, where the dataset lives, how to reach OpenSearch, which features exist in your index, and, if you deploy to SageMaker, the bucket, the execution role and the region. Here is the smallest override we would actually run:

config/ltr/us_patient.ymlyaml
1model:
2 name: fs-ltr-us-patient
3 hyperparameters: # the optimizer writes here
4 objective: rank:ndcg
5 eval_metric: ndcg@12
6 max_depth: 6
7 learning_rate: 0.1
8
9ubi_data_loader:
10 max_days_ago: 30
11 n_events: 500000
12 opensearch:
13 url: <%= ENV["OPENSEARCH_URL"] %>
14 username: <%= ENV["OPENSEARCH_USERNAME"] %>
15 password: <%= ENV["OPENSEARCH_PASSWORD"] %>
16 index:
17 queries: ubi_queries
18 events: ubi_events
19
20search_time_features: # and here
21 - relevancy
22 - bm25_name
23 - bm25_description
24 - bm25_ingredients
25
26feature_store:
27 features: # and here
28 - price
29 - purchase_count
30 - purchase_conversion
31
32optimization:
33 search_time_features: # the pool feature selection draws from
34 - relevancy
35 - bm25_name
36 - bm25_description
37 - bm25_ingredients
38 feature_store_features:
39 - price
40 - reorder_rate
41 - purchase_count
42 - purchase_conversion
43
44training:
45 modelling_dataset_path: "s3://my-bucket/ltr/us_patient/dataset.csv"
46 s3_region: us-east-1
47
48dataset_enricher:
49 class: "Search::Reranking::DatasetEnricher"
50
51deploy:
52 enabled: true
53 sagemaker_bucket: my-bucket
54 execution_role: <%= ENV["SAGEMAKER_EXECUTION_ROLE"] %>
55 instance_type: ml.m5.large
56 metric: ndcg
57 minimum_metric_value: 0.5


1config = Ltr::ConfigLoader.new(override_config_path: "config/ltr/us_patient.yml")


Everything else falls back to the gem defaults. Loading validates the merged config and raises Ltr::ConfigurationError naming the key when something is off: split ratios that do not sum to one, a deploy.metric that is not in evaluation.metrics, local: true without a local_path. It does not check feature names against your dataset, metric names against the ones it knows, or whether you remembered to replace our bucket with yours, so those surface later, and we will come back to how. We run anywhere from one to four models per search engine, split by things like geography and user type, and each one is just another override file.

Step 1: Collect the data

This is the part the gem cannot do for you, and honestly it is one of the hardest parts of the whole project. You need to log every query with its ordered results and scores, and every user action against those results, with the query id that produced them. We use the OpenSearch User Behavior Insights (UBI) schema for this, and the gem's dataset pipeline reads that schema directly.

The two indices need three things from their mappings: timestamp as a date, because the loader pages through events by it; query_id and action_name as keyword, because they are filtered exactly; and query_attributes stored but not indexed, because its scores object is read straight out of _source and its key order is the displayed rank, which is not something you want an analyzer anywhere near:

1PUT ubi_queries
2{
3 "mappings": {
4 "properties": {
5 "query_id": { "type": "keyword" },
6 "user_query": { "type": "text" },
7 "timestamp": { "type": "date" },
8 "client_id": { "type": "keyword" },
9 "application": { "type": "keyword" },
10 "query_attributes": { "type": "object", "enabled": false }
11 }
12 }
13}
14
15PUT ubi_events
16{
17 "mappings": {
18 "properties": {
19 "query_id": { "type": "keyword" },
20 "action_name": { "type": "keyword" },
21 "timestamp": { "type": "date" },
22 "event_attributes": {
23 "properties": { "product_id": { "type": "keyword" } }
24 }
25 }
26 }
27}


Ltr::Ubi builds documents in the right shape. It does not write anything, you index the hash with whatever client you already have. It reads the id field names from the gem's default config rather than from your override, and a couple of places in the pipeline assume the defaults too, so leave dataset_builder.query_id_field and doc_id_field at query_id and product_id:

1# After running the initial retrieval query
2scores = hits.to_h do |hit|
3 [hit["_id"], hit["fields"].slice("relevancy", "bm25_name", "bm25_description", "bm25_ingredients")]
4end
5
6query_doc = Ltr::Ubi.query(
7 query_id: query_id,
8 user_query: params[:q],
9 timestamp: Time.current,
10 scores: scores, # insertion order is the displayed rank
11 application: "web",
12 client_id: current_user.id.to_s # ids are strings, the gem will tell you if not
13)
14opensearch.index(index: "ubi_queries", body: query_doc)


1# When the user does something with a result
2event_doc = Ltr::Ubi.event(
3 query_id: query_id,
4 action_name: "add_to_cart", # or "view", "purchase", ...
5 doc_id: product.id.to_s,
6 timestamp: Time.current
7)
8opensearch.index(index: "ubi_events", body: event_doc)


The scores hash values are the search-time features the model will train on, and the position each product was displayed at is what the cascade examination model needs. The action names you log are the ones you reference in conversion_rate_targets, and only those. An action you log but never reference, say add_to_favorites, does not move the cascade cutoff, and a query whose only events are unreferenced actions produces no training rows at all. If you want an action to count, give it a target, even a low-weight one.

One more thing about identity. Judgment lists are grouped by the normalized query text, stripped and downcased, so "Magnesium " and "magnesium" are one query. If the same text means different things in different contexts, we have separate catalogs per country, name the query document fields that should split them in dataset_builder.additional_qid_fields.

Step 2: Build the dataset

1result = Ltr::DatasetAggregation::DatasetPipeline.call(config)
2result[:rows] # => 184302
3result[:dataset_path] # => "s3://my-bucket/ltr/us_patient/dataset.csv.gz"


This pulls the newest n_events events from the last max_days_ago days, fetches the queries those events reference, joins them, runs the cascade examination and aggregation described above, computes labels, derives the engineered features, assigns splits, and writes a gzipped CSV to training.modelling_dataset_path, local or S3. n_events is the knob that sizes your dataset, not the day count, and it defaults to a modest 10,000, so set it. And the path you configure ends in .csv; the file that gets written is that path plus .gz. The pipeline raises on anything it cannot recover from, an empty window, an enricher class that does not exist, so the result you get back is always a success. It carries rows, columns, enriched and dataset_path.

If your products have features that do not live in the search index, like price or a reorder rate from your orders table, you add them here with an enricher. Subclass DatasetEnricher, return a dataframe keyed by product_id, and name the class in config. The join key is cast to an integer on both sides, so your product ids need to be numeric today; string keys like UUIDs are on the list:

1module Search
2 module Reranking
3 class DatasetEnricher < ::Ltr::DatasetAggregation::DatasetEnricher
4 def fetch_product_features(unique_product_ids)
5 rows = Product.where(id: unique_product_ids).pluck(:id, :price, :reorder_rate)
6 Rover::DataFrame.new(
7 "product_id" => rows.map { |id, _, _| id },
8 "price" => rows.map { |_, price, _| price.to_f },
9 "reorder_rate" => rows.map { |_, _, rate| rate.to_f }
10 )
11 end
12 end
13 end
14end


Products your frame leaves out get nil in those columns, which XGBoost reads as missing.

A few things to know before you open the CSV. Queries with fewer than dataset_builder.min_documents_per_query products (three by default) are dropped, because there is nothing to rank. The view_count, add_to_cart_count and purchase_count columns are per-product totals across every query, because they are features the model will see again at inference time; the per-query counts the label was computed from are not in the file, and neither is impression_count. If you need to explain one row's label, dataset_builder.debug_mode: true keeps those, and you must never train on the result. Splits are assigned with a fixed seed, so rebuilding from the same events gives the same split.

The default aggregation engine does all of this in memory, which is fine for a few hundred thousand events. Once a pull no longer fits in RAM, one line of config switches it to an out-of-core engine that stages the download in a scratch SQLite file and aggregates there, bounding peak memory by the OpenSearch page size instead of the size of the pull:

1dataset_builder:
2 aggregation_engine: sqlite


Both engines produce the same labels and features for the same events. The SQLite engine filters events down to the cascade actions on the OpenSearch side, so n_events counts only those, and unreferenced actions never reach the dataset. Its scratch database goes under dataset_aggregation.staging_root, the system temp dir by default, needs room for the whole uncompressed pull, and is deleted whether the run succeeds or fails. We wrote that one because our training servers were, to quote the talk, beefy, and Ruby is not shy about memory.

Step 3: Optimize

1result = Ltr::Optimize::OptimizePipeline.call(config)
2result[:features] # => { search_time_features: [...], feature_store_features: [...] }
3result[:hyperparameters] # => { objective: "rank:ndcg", max_depth: 7, learning_rate: 0.18, ... }


This loads the dataset, runs feature selection over optimization.search_time_features and optimization.feature_store_features, runs Bayesian optimization over hyperparameter_search_space, and writes the winners back. It rewrites your override file in place, replacing the search_time_features, feature_store.features and model.hyperparameters blocks. It replaces, it does not insert: a block that is not in the file is skipped without a word, and the run's result then lives only in the log. That is why the override above carries all three, even though two of them would have fallen back to the defaults anyway. Comments outside those blocks survive; comments inside them do not, since the lines they annotated are gone. With ssm.enabled: true it also writes them to AWS SSM Parameter Store so the training pipeline can pick them up from there. Every trial fits on train, early-stops on training.early_stopping_split, and is scored on optimization.split, which is val.

A few things about the search space. A parameter is treated as an integer when both bounds are integers, so max_depth: { min: 3, max: 10 } is rounded and clamped, and min: 3.0 would quietly make it a float. Ranges whose upper bound is at least twenty times the lower are searched on a log scale. There are no categorical parameters. n_init is how many Latin hypercube points seed the Gaussian process before expected improvement takes over, and the loop stops early if ten iterations in a row fail to improve. A feature in your pools that is not in the dataset gets a warning and is dropped from the search.

This is the expensive step. It trains a lot of models. We run it daily, but weekly would be fine for most, and you can skip it entirely and train with the hand-set features and hyperparameters in your config.

Step 4: Train and deploy

1result = Ltr::Train::ModellingPipeline.call(config)
2
3result[:metrics]
4# => { "map@10" => 0.71, "mrr@10" => 0.79, "ndcg" => 0.87, "ndcg@5" => 0.83, "ndcg@12" => 0.85,
5# "precision" => 0.34, "precision@5" => 0.52, "recall@5" => 0.61 }
6
7result[:optimization_split_metrics] # the same metrics on val, for the gap
8
9result[:deployed] # => true, the SageMaker gate
10result[:deployment] # => { deployed: true, reason: :threshold_met, metric: "ndcg", score: 0.87, threshold: 0.5, endpoint: {...} }


Training reads the dataset, trains the XGBoost ranker with the optimized features and hyperparameters, evaluates it on the test split with every metric in evaluation.metrics, and then checks the gate. The metric names are ndcg, map, mrr, precision and recall, with or without an @k.

The failure mode to watch for here is quiet. A feature named in config but missing from the dataset, because of a typo or because the enricher did not produce it, is logged and dropped, and training carries on without it. It only raises if no features are left. Check the training log for dropped features the first time you run a new config.

If deploy.enabled is on and deploy.metric clears deploy.minimum_metric_value, the model ships: the booster, its feature list, the feature store JSON and the Python inference handler are tarred up and uploaded to deploy.sagemaker_bucket, and the endpoint <model.name>-endpoint is created or updated to serve them. The pipeline blocks until the endpoint is InService, a few minutes for an update, so budget for that in whatever schedules it. On the first run, the container image is looked up by training.s3_region and only the four US regions are mapped, and the tarball is built inside the gem's install directory, so that path has to be writable. The feature store is built from the dataset you just trained on, so its values are as fresh as your last dataset build.

The credentials running the pipeline need to write to the bucket, create and update SageMaker models, endpoint configs and endpoints, and pass the execution role; whatever calls the reranker at query time needs sagemaker:InvokeEndpoint. The endpoint is a real instance billing around the clock, so instance_type is a cost decision as much as a latency one.

For local development, or if you would rather not run SageMaker, there is a local deploy mode that saves the booster to disk:

1deploy:
2 enabled: false
3 local: true
4 local_path: "tmp/ltr/models/us_patient.ubj"


Local deploy is independent of the SageMaker gate and does not have one of its own: if local is on, the booster is written, whatever the score. It reports under its own key, and result[:deployed] stays false because that one is the SageMaker answer:

1result[:deployed] # => false
2result[:local_deployment] # => { deployed: true, reason: :local_saved, path: "tmp/ltr/models/us_patient.ubj" }


Step 5: Rerank

The inference payload is a list of items, each with an id and the fields the model needs:

1items = hits.map do |hit|
2 {
3 "id" => hit["_id"],
4 "fields" => [
5 { "name" => "relevancy", "value" => hit["_score"] },
6 { "name" => "bm25_name", "value" => hit["fields"]["bm25_name"] },
7 { "name" => "bm25_description", "value" => hit["fields"]["bm25_description"] },
8 { "name" => "bm25_ingredients", "value" => hit["fields"]["bm25_ingredients"] },
9 ]
10 }
11end


The SageMaker client takes that list wrapped in an items key, because that is the JSON the endpoint expects. The local reranker takes the bare list. Ltr::Inference::Base picks the backend for you, local when deploy.local is on and SageMaker otherwise, but it passes your payload through untouched, so hand it the shape of the backend it is going to pick:


1response = Ltr::Inference::SagemakerClient.call({ "items" => items }, config)
2response = Ltr::Inference::LocalReranker.call(items, config)
3
4# => {
5# "items" => [
6# { "item" => "prod_812", "score" => 2.31 },
7# { "item" => "prod_104", "score" => 1.97 },
8# { "item" => "prod_559", "score" => 0.42 },
9# ...
10# ],
11# "took_seconds" => 0.006
12# }
13
14reranked_ids = response["items"].map { |item| item["item"] }


Items come back sorted by score, ids as strings. Either way, the _rank and _relative_score features are derived on the fly from the fields you send, using the model's own feature list to figure out which ones it needs. A field you leave out reaches XGBoost as NaN, which it handles as missing rather than failing. A field you send with a nil value does not: locally it becomes 0.0, so drop the field rather than sending nil. On SageMaker the handler first tries the feature store shipped in the tarball for any field you did not send, so you only have to send the search-time features, and a value you do send always wins. Locally there is no feature store lookup, so merge the feature store features into the fields yourself.

The local reranker caches the booster per path for the life of the process. If you redeploy a new .ubj under a running app, call Ltr::Inference::LocalReranker.reset! or restart.

The SageMaker client has aggressive default timeouts, half a second to connect and half a second to read, configurable under inference.sagemaker. A reranker that times out should fall back to the original order, not hold up the page, and that fallback is your job, not the gem's: a timeout raises, and the caller catches it. Ours looks like this:

1def rerank(hits)
2 response = Ltr::Inference::SagemakerClient.call({ "items" => items_for(hits) }, config)
3 order = response["items"].map { |item| item["item"] }
4 hits.sort_by { |hit| order.index(hit["_id"]) || order.size }
5rescue Seahorse::Client::NetworkingError, Aws::SageMakerRuntime::Errors::ServiceError => e
6 Ltr.config.report_error(e, query: params[:q])
7 hits
8end


Note that the AWS SDK's default retry policy still applies to the runtime client, so a bad half second can turn into a couple of them before that rescue fires. Budget for that or turn retries down on the client.




Running it on a schedule


The gem is a library, not a daemon. There is no scheduler in it. What it does ship is a rake task per pipeline. Each one takes the override file either as a task argument or from the environment, and the argument wins if you pass both:


1bin/rails "ltr:aggregate[config/ltr/us_patient.yml]"
2bin/rails "ltr:optimize[config/ltr/us_patient.yml]"
3bin/rails "ltr:train[config/ltr/us_patient.yml]"
4
5# or, which reads better in a crontab
6LTR_CONFIG=config/ltr/us_patient.yml bin/rails ltr:train


aggregate logs the row count and where the dataset landed, optimize logs the features and hyperparameters it chose, and train logs the metrics and the reason on each deploy gate. In our app, cron runs the three nightly, in that order, once per model, and spinning up a new model for a new search engine is a new override file and three more cron entries.


The first time through, the things worth checking are the row count from aggregate (zero or tiny means the mapping or the action names are off), the dropped-feature lines from train, and the deploy reasons, which will be one of :disabled, :below_threshold or :threshold_met for SageMaker and :not_configured or :local_saved for local.


Did It Work?

Our users have not sent many messages saying "good job on search". What we have instead is a Slack channel dedicated to negative search feedback that has gone almost silent. We get maybe one message a month, and our NDCG dashboards have moved in the direction we want them to. In search, silence is the compliment.

The first release was the biggest single jump, but we have been iterating since 2024 and the gains have kept coming. Having judgment lists changed how we work as much as it changed the results. Every change to retrieval, whether it is a new synonym filter or a tweak to query understanding, gets run through the judgment lists offline for an NDCG score before it goes near an A/B test. We go into production informed instead of hopeful.


Conclusion

Learning to rank is not new, but doing it end to end in Ruby without a Python sidecar felt new to us, and it worked. If I had to boil the whole project down: a shared metric beats gut feel, your best training data is probably something you already collect, and you do not have to leave Ruby to do machine learning.

Thank you for reading this post! If you want the version with all the diagrams and a couple of bad jokes, watch the talk. Dave and Lucas have also written about the other stages of our retrieval pipeline, query understanding and personalization, here on Builder's Corner.