Kumo's relational foundation model beats the best supervised models on benchmark tasks by about 5% relative accuracy out of the box, rising to a 12% improvement after fine-tuning, gains large enough to translate into tens of millions of dollars of business impact in production systems.
Leskovec cites benchmark results showing RFM2 beats state-of-the-art supervised models by 5% zero-shot and 12% after fine-tuning, with real deployments at Reddit, DoorDash, and Coinbase showing large revenue and accuracy gains. ✦ AI generated
Jure Leskovec · The TWIML AI Podcast · 2026-05-21 · original ↗
starts at this moment · 43:13
“And so speaking of performance, uh you talked a little bit about some of the challenges with uh collecting benchmarks. But how do you find performance relative to those benchmarks and uh also you know more importantly in the real world?”
what we see is that the foundation model um by itself improves uh state-of-the-art uh over all supervised models ever published on this on this benchmark. Right? So so the baseline is very high. It's like just build the best model you can and see how high you can get. um uh uh the foundation model improves that I think for about 5% relative uh the accuracy um and then if you further tune the model meaning if you would fine-tune it do some grain and base updates then the performance goes to 12% uh over the state-of-the-art
verbatim transcript · starts at 43:13
43:13a white paper on Kumar RFM 2 uh that that people can read with a bunch of different benchmarks. Um what we see is that the foundation model um by itself improves uh state-of-the-art uh over all supervised models ever published on this on this benchmark. Right? So so the baseline is very high. It's like just build the best model you can and see how high you can get. um uh uh the
43:41foundation model improves that I think for about 5% relative uh the accuracy um and then if you further tune the model meaning if you would fine-tune it do some grain and base updates then the performance goes to 12% uh over the state-of-the-art and those are quite sizable sizable gains especially if you think about putting this in production in uh recommener systems or fraud detection where you every single digit
44:10performance in increase in accuracy can mean millions tens of millions uh in uh in business impact. Maybe the second thing I would say is where we see these methods also shine is with noisy and incomplete data cold start problems because of the relationships uh because of the relational structure the model is able to much better kind of hone in and be much more robust to the data
44:38missingness data corruption and things like that. So we've also done quite a lot of analysis around understanding and and uh like how this performs in on real world data uh sparse data small amounts of data noise incompleteness uh irrelevant columns and things like that. >> And and when you mention cold start like that suggests hey I want to start identifying fraudulent transactions but I have no labels I just have a bunch of
45:08data. Can you tell me where I should start looking? Like does it work for that kind of problem? >> Uh yeah, maybe I should quantify what cold start means. Usually cold start would mean um when a new user shows a new product shows up, right? So you still need to have some historical labels. I'm not you still need some historical labels, but usually you know prediction is easy once you have a lot
- ·Outperforms all published supervised models on benchmarks
- ·5% relative accuracy improvement zero-shot
- ·12% improvement after fine-tuning
- ·Gains translate to tens of millions of dollars
- ·Deployed at Reddit, DoorDash, and Coinbase
- ·Revenue and accuracy gains in production