Using dataset embeddings as prompts instead of model prompts enables privacy-preserving generation of neural network weights from sensitive data that cannot be shared, opening model generation to regulated industries.
Borth describes the vision of dataset prompting: instead of needing an existing model to anchor weight generation, a single embedding of a dataset could serve as the prompt, enabling financial and healthcare institutions to generate models from sensitive data without exposing the data itself. ✦ AI generated
Damian Borth · The TWIML AI Podcast · 2026-07-27 · original ↗
starts at this moment · 32:08
“So give me a model that works well on this data.”
And if this could be done in a privacy preserving way, it would be even better, right? Because then suddenly people could use this without revealing individual members of this data set. So there's work that does something similar already and Kais you know um Zo and Andreas and others were working on that and we kind of moved this forward because we have model zoos right we have data sets and we have models so we have kind of images and we have models and why not trained and aligned space like clip did with text and images. Then we could use actually a kind of a data set encoder as a model prompt that could move into this space and point into this space instead of a model prompt. We use a data set prompt. So you you're a bank, you are a you know financial institution, healthcare provider, whatever. You don't reveal your data. You you have your data set of I don't know 100 samples, thousand samples. You create one data set embedding. So you cannot infer you know the individual members or samples. you give this embedding to us, we provide you the weights, we give you the weights, you're much faster in continuing training. So this would be this would be really interesting, right? Because that opens up to all the you know potential data sets that are not on hugging phase
verbatim transcript · starts at 32:08
32:08There's one missing piece though and we have a paper crank in review that might solve that. Uh so if you think about that we need a model as a prompt to get an anchor to sample other models can be a domain change. So what would be really really nice if you would not need this model as a prompt but you could prompt with your data set. >> So give me a model that works well on
32:30this data. >> Exactly. And if this could be done in a privacy preserving way, it would be even better, right? Because then suddenly people could use this without revealing individual members of this data set. So there's work that does something similar already and Kais you know um Zo and Andreas and others were working on that and we kind of moved this forward because we have model zoos right we have
32:55data sets and we have models so we have kind of images and we have models and why not trained and aligned space like clip did with text and images. Then we could use actually a kind of a data set encoder as a model prompt that could move into this space and point into this space instead of a model prompt. We use a data set prompt. So you you're a bank, you
33:20are a you know financial institution, healthcare provider, whatever. You don't reveal your data. You you have your data set of I don't know 100 samples, thousand samples. You create one data set embedding. So you cannot infer you know the individual members or samples. you give this embedding to us, we provide you the weights, we give you the weights, you're much faster in continuing training. So this would be this would be really
33:46interesting, right? Because that opens up to all the you know potential data sets that are not on hugging phase >> and is the the privacy preserving angle there because you your process would start with that embeddings anyway. So it doesn't matter who produces it or uh is that a compromise that you could do it with the processes but you could probably get more out of it if you had the actual
34:12data. >> Good question. So um it comes by the method that we used because uh imagine you have a data set with 10,000 1 million images. You need some kind you cannot prompt with all individual images. You need some. >> So you need a prompt with a thing, >> one thing or one vector, right? And you need to aggregate this knowledge. So um like you know you do a sentence and you
34:35have a text embedding to prompt your image that you generate. So it comes with that obviously you have to check for you know um task membership attacks you have to add some noise etc. But the idea is really like once we have this one embedding can we generate from this one data set embedding now the tokens or the embeddings that generate the tokens for the weights and this would open up
34:58and then the question is this would open up to data sets that people are not willing to share or not allowed to share because of regulation a financial institution they would maybe love to share but they're not allowed and I have a you know external PhD student with the Deutsche Bundes the national bank of Germany they cannot share But you know they might use those embeddings for that