ATRIUMsearch → argument graph
DataVideo · 43:13 — 44:43

In benchmark tests the RFM foundation model improves on the best-ever published supervised models by about 5% out-of-the-box, and by 12% when fine-tuned, translating into large real-world business impact at companies like Reddit and DoorDash.

Jure cites benchmark results showing the foundation model beats all previously published supervised models by ~5% (12% after fine-tuning), and describes real production wins at Reddit (near double-digit CTR increase) and DoorDash (hundreds of millions in revenue impact). ✦ AI generated

Jure Leskovec · The TWIML AI Podcast · 2026-05-21 · original ↗

starts at this moment · 43:13

Elicited by

So speaking of performance, you talked a little bit about some of the challenges with collecting benchmarks. But how do you find performance relative to those benchmarks and also, you know, more importantly, in the real world?

The foundation model by itself improves state-of-the-art over all supervised models ever published on this benchmark... the foundation model improves that I think for about 5% relative accuracy, and then if you further tune the model, meaning if you would fine-tune it, do some gradient-based updates, then the performance goes to 12% over the state-of-the-art.

verbatim transcript · starts at 43:13

Transcript · around this moment

43:13a white paper on Kumar RFM 2 uh that that people can read with a bunch of different benchmarks. Um what we see is that the foundation model um by itself improves uh state-of-the-art uh over all supervised models ever published on this on this benchmark. Right? So so the baseline is very high. It's like just build the best model you can and see how high you can get. um uh uh the

43:41foundation model improves that I think for about 5% relative uh the accuracy um and then if you further tune the model meaning if you would fine-tune it do some grain and base updates then the performance goes to 12% uh over the state-of-the-art and those are quite sizable sizable gains especially if you think about putting this in production in uh recommener systems or fraud detection where you every single digit

44:10performance in increase in accuracy can mean millions tens of millions uh in uh in business impact. Maybe the second thing I would say is where we see these methods also shine is with noisy and incomplete data cold start problems because of the relationships uh because of the relational structure the model is able to much better kind of hone in and be much more robust to the data

44:38missingness data corruption and things like that. So we've also done quite a lot of analysis around understanding and and uh like how this performs in on real world data uh sparse data small amounts of data noise incompleteness uh irrelevant columns and things like that. >> And and when you mention cold start like that suggests hey I want to start identifying fraudulent transactions but I have no labels I just have a bunch of

45:08data. Can you tell me where I should start looking? Like does it work for that kind of problem? >> Uh yeah, maybe I should quantify what cold start means. Usually cold start would mean um when a new user shows a new product shows up, right? So you still need to have some historical labels. I'm not you still need some historical labels, but usually you know prediction is easy once you have a lot

Around this claim