Philipp Herzig explains that while LLMs excel at unstructured text and images, enterprise planning requires predictions — demand forecasting, cash flow prediction, classification of customer payment behavior, regression of payment delays. LLMs are not designed for these tasks because they generate one token at a time in sequence-to-sequence modeling. Classical ML approaches like XGBoost work but don't scale because they require hiring data scientists and training separate models per country. SAP's RPT (Relational Pre-trained Transformer) applies the transformer architecture to structured tabular data, enabling accurate predictions with small amounts of data.
transcript
Philipp Herzig: Now, first of all, from a business motivation, it's a great question, right, Sarah? I mean, first from the business motivation point of view, right? Again, LLMs, unstructured world, that's all good, right? But most of the time, if you want to plan forward, right? If you want to make good decisions in a company, you need predictions, right? You need predictions in terms of, oh, what's my demand, right? For, oh, is this depending on the seasonality effects and so on? What's my demand forecast maybe, right, for my products in the retail store? Or what's my demand, right, for my products so I can plan accordingly my manufacturing, right, if I'm a manufacturing customer. Or you want to predict your cash flow, right? You want to predict, and that has a bunch of input variables like, oh, what are actually my day of sales outstanding, right? And that is determined based on Are customers paying yes or no? That's a classification question. And if you then say, okay, if a customer is not paying within the payment terms, what's the payment delay? A classical regression question. And so on and so forth. The problem is, of course, still today, if we look at these predictive questions, right? And then you wanna maybe do a what if analysis from it, right? Now, if you want to do these predictions, quite frankly, then the challenge is large language models are not made for this, right? The way how they generate just one token after another, essentially, in a sequence-to-sequence modeling. I mean, they're language models, right? And they do this phenomenally well. But if you still want to do these predictors, you have to go back to these classical machine learning approaches. You use XGBoost or AutoGluon and many of these AutoML approaches that are still out there. The problem is just it doesn't scale. So we haven't seen in the predictive space the same level of democratization. You still need to hire a very good talented data scientist, right? And then if you, for example, if you're a large company, we did this, for example, at a pharmaceutical company, if you just want to solve the payment delay prediction problem I've mentioned, right? They are running in 90 countries around the world and they need these two models. So you end up with 180 models you need to train. You need to curate the data, you need to train the models, figure out, right, what the right model is, feature engineering, like the classical, machine learning kind of approach, right, that was used in the past. And what we said all the time is, okay, look, we have all this data stored in these tables, right? Thousands of tables, right, where all this information is stored. Can we not apply the same idea that large language models or multimodal models did for the unstructured world, actually for the structured in order to start predicting things? So you can just basically provide a little bit of context, a small amount of data, not a large amount of data, because that will always a problem, small amount of data, and then starting making high accurate predictions, so to speak, in that domain. And that led, actually, it was two years of research. We published that also at NeurIPS and a bunch of other conferences. We call this RPT1, so RAPID1 stands for relational pre-trained transformers. It's still based on the transformer architecture, but with a very different architecture. We released this and we see some, meanwhile, some very, very good results from that in various domains where, as I said, classification and regression, sometimes time series, and so on are concerned. And we believe this will be huge because it obviously will allow way more people from a business impact to make these predictions, which large language models have a really hard time with.