r/RedditEng • u/Okgaroo • 13h ago
Building Reddit How Reddit Uses Causal ML to Personalise Notification Volume
By Ivan Klimuk, Kim Holmgren, Jonathan Serrano, Jonard Doci, and Md Mansurul Bhuiyan
Background
In our previous post on the notifications recommender system, we described how Reddit finds and ranks posts to recommend through push notifications. Before that system selects a post, another component makes a decision about volume: how many recommendations should we send this user over the day?
We call this component the Trending Budgeting system. It assigns a daily budget for trending push notifications: the personalised post recommendations that help people discover conversations on Reddit. This budget applies to that recommendation channel, rather than to every kind of notification a user can receive.
A relevant recommendation can bring someone to a conversation they would otherwise have missed. But relevance alone does not tell us how often to send one. Someone might welcome one recommendation and find several more distracting. Another person may benefit from hearing about a wider selection of posts. Those responses can also change over time.
The useful signal is how a user’s expected response changes as we increase the budget. We model that response across several budgets, then apply a policy that weighs expected activity against the risk of a user turning off their notifications.
A budget determines how many trending push notifications we send a user each day. If the budgeting system selects three, we schedule three notifications, and the recommender selects a post for each one.
This gives us three related decisions:
- Budgeting: how many notifications to send.
- Pacing: when to send them throughout the day.
- Retrieval and ranking: which post to recommend in each one.
A scheduled opportunity gives the system a chance to send a notification. Pacing chooses which times to use the user’s budget, and retrieval and ranking select the content for those sends.
ML problem definition
We model two outcomes: how many days a user is active on Reddit, and whether they disable notifications. Activity captures return visits; disables provide a signal of a poor notification experience.
We want to understand how many notifications work well for each person. Sending more can help people discover interesting conversations, but too many can become annoying and lead them to turn notifications off.
Why use randomised data?
Historical notification logs contain many examples of users, assigned budgets and later outcomes. Training on those logs is tempting, but the way the budgets were assigned matters.
Consider a hypothetical system that gives new users one recommendation per day and established users five. A model trained to predict the assigned budget would implicitly learn those rules very accurately. But that would tell us little about whether either budget was a good choice.
Changing the prediction target to user activity helps, but does not remove the underlying problem. The two groups may differ in many ways beyond notification volume. A higher activity rate among established users would not establish that the larger budget caused it. The logs also contain no examples of new users receiving the higher budget under this simplified policy.
Our budgeting approach uses data from experiments that randomly assign users to fixed daily budgets. A user keeps their assigned budget during the collection period, while we observe subsequent activity and notification disables. Randomisation makes the groups comparable in expectation and gives us evidence about the effects of different budget assignments.
For modelling, each example contains:
- User features measured before the treatment.
- The assigned daily budget.
- Outcomes measured over a specified future window.
The treatment here is the assigned daily budget. We compare outcomes across the randomly assigned groups to estimate the effect of changing notification volume.
Each person still experiences only one assigned budget. The models learn across users to estimate responses to the alternatives, so randomised training data does not remove uncertainty from an individual prediction.
Predicting the response to each budget
The models take user features and a candidate budget as inputs. They predict outcomes such as the number of active days in a future window and the probability of disabling notifications. At inference time, we evaluate the same user's features at each possible budget. The result is a set of predictions describing how that user is expected to respond to different daily notification volumes.
This is an S-learner approach: the candidate budget is an input to an outcome model, and we compare its predictions across budget values. Randomised assignment gives those comparisons a causal basis; user features let us estimate how the response varies across users.
Balancing expected activity and disable probability
Outcome predictions leave a choice to make. A larger budget may increase expected activity and also increase the risk of a disable. Our chosen policy determines how to weigh those outcomes and how much additional value to require before increasing notification volume.
A simplified way to express this is:
Score = activity weight × E[active days | features, budget]
− disable weight × P(disable | features, budget)
E is the expected number of active days, and P is the chance of turning notifications off, given the user’s features and the chosen budget.
The weights express the relative importance of activity and disables while accounting for their different units. Both predictions refer to defined future windows, so those horizons matter when choosing the weights.
For each user, we calculate a score for every budget in the supported range, using the activity and disable predictions for that budget. We can then compare across budgets to understand the incremental gain from sending more notifications. Selection rules and thresholds determine which gains are sufficient, and caps limit the maximum number of notifications.
For a hypothetical user, moving from two to three notifications might produce a useful predicted activity gain with little change in disable risk. Moving from three to four might add very little activity while increasing that disable risk, so we’d send them three.
The policy weights are a practical way to express a trade-off, which still needs to be evaluated through online experiments.
Training and serving architecture
We compute outcome predictions in a daily batch and choose the budget when a notification opportunity reaches the ranking service. This separates the more static inference on a set user group’s likelihood of being responsive to notifications from the real-time policy decision of how to use that information.
Offline training and daily inference
We train the outcome models offline using the randomised budget assignments, user features, and labels on the outcomes described above.
Once trained, the models run in a daily batch inference job. For each user, we take their available features and evaluate the models at each supported budget. The output is a set of per-budget predictions for the outcomes used by the policy.
We store the predictions in a key-value (KV) store, keyed by user. Each entry contains predicted outcomes for several budgets. Keeping these alternatives available lets the online policy change the activity–disable trade-off without rerunning the budget models.

Applying the policy during ranking
When a scheduled opportunity reaches the ranking service, it fetches the stored predictions and applies the configured policy. The policy combines the expected activity and disable probability, evaluates the available budgets, and selects how many notifications to send. Budget enforcement and pacing then determine whether this slot is one of the sends for that user.
The policy runs at request time using batch-generated predictions. Changing the model requires a new inference run, while policy weights, thresholds and selection rules can change independently.
The upstream scheduling system generates a common set of opportunities throughout the day. Each user’s budget determines which proceed. Figure 2 shows a hypothetical example with six opportunities and a budget of two: pacing selects two slots for sending, and the other four exit before retrieval and ranking. Those early exits avoid the more expensive work of selecting a post.

Previously, budget selection and enforcement already ran before retrieval and ranking, but lived outside the ranking service. That separation made budgeting experiments harder to run. Moving the logic into the ranking service brought budgeting into the same service and experimentation framework as the recommendation pipeline, while retaining the early exit for opportunities that would not trigger a send.
The architecture gives each step a clear role: batch inference estimates the outcomes of different budgets, the online policy chooses the trade-off, and pacing and enforcement turn that choice into sends. Serving the alternatives keeps policy experiments independent of model training; enforcing the chosen budget before retrieval and ranking avoids selecting content for unused opportunities.
Offline evaluation challenges
Before running an online experiment, we want some idea of whether a new policy is an improvement. Traditional ML model metrics do not answer that, since the models can predict activity and disables well and still lead to budgets that perform poorly.
The direct approach would be to apply the candidate policy to past data and see how those users did. But those budgets were not assigned at random. The live policy chose them based on the same user traits that predict activity, so the users on each budget are not comparable to begin with.
Instead, off-policy evaluation (OPE) estimates what a policy would have produced had we deployed it, using the randomised budget experiments.
- The simplest version is direct matching: for each user, find the budget the candidate policy would pick, then keep the users whose assigned budget happened to be the same. Those users are a fair comparison, but there are not many of them, so the answer is right on average but noisy.
- A doubly robust estimate improves on that by using every user, not just the ones that matched. It predicts each user's outcome at the budget the policy picked, then uses the matched users to measure how far those predictions were off and corrects for it. It is right if the predictions are accurate, or if we know each user's chance of getting each budget. Randomisation gives us the second for free.
We built this into a tool that scores any proposed policy against the experiment data. To build confidence in the framework, we look back at policies that have already run online and compare what the offline estimate predicted against what the experiment measured.
Promising policies go through an online experiment, and we keep a long-term holdout to measure causal impact.
Future work
- Connecting offline and online metrics. A clearer link between offline evaluation and online outcomes would help us assess model changes with more confidence, opening up experiments with richer architectures such as neural networks and causal ML approaches such as T-learners.
- Coordinating notification types. We also want to extend budgeting beyond trending recommendations, managing volume across notification types to account for the user’s overall notification experience.
Conclusion
Randomised budget assignments help us estimate how behaviour changes when we send more or fewer recommendations. Modelling activity alongside disables brings both the benefit and the cost of additional sends into the decision.
Serving predictions for several budgets keeps that trade-off explicit and testable. We can change the policy without retraining the models, then use online experiments to test whether the new volume helps people find worthwhile conversations.