ATRIUMsearch → argument graph
PredictionVideo · 110:27 — 113:11

Pre-training data filtering is an overlooked but promising approach to reducing misuse risk in open-weight models, and FAR.AI plans to open-source safe filters and fund $2M in grants for this work.

Gleave argues that simply filtering dangerous knowledge out of pre-training data — like not training models on anthrax papers — is a surprisingly effective and underused safety measure that doesn't degrade most capabilities. He notes this has been validated by independent researchers and the UK's AI Security Institute, and that FAR.AI plans to open-source safe filters while privately sharing more sensitive ones with developers, alongside launching a $2 million grant program for open-weight safety research. ✦ AI generated

Adam Gleave · The Cognitive Revolution · 2026-07-30 · original ↗

starts at this moment · 110:27

I think for that we're going to need new approaches, but one I'm most optimistic about in the short term is pre-training filtering. So it's a simple idea where just don't train the models on really dangerous stuff. If you don't need your model to help people make anthrax to don't train it on the anthrax papers. A a tiny number of users might be a little bit sad that it can't answer questions about this but most people won't even notice. But it has a big impact on the sort of misuse potential of the model. And this has been validated in a number of scientific papers. This is independent researchers. UK's AI security inropic has been sponsoring some research into this. Open a actually used this in their GPD OSS release. So, it's been tested quite well, but it's not become kind of common practice. And this is something we're actively excited about scaling to make sure this does work and a near frontier approach. And going back to what you were saying earlier about could we start sharing some of these data sets or filters with people. Our plan is to open source as much as we think it doesn't have a misuse potential and then privately share with developers of things that that do have misuse potential that could be really useful to just lower the cost of these kinds of interventions.

verbatim transcript · starts at 110:27

Transcript · around this moment

(01:10:48) a technical expertise angle and also just a compute angle to do finetuning against these models. It usually takes us at least a few weeks to get a new open weight model hooked into our infrastructure for fine-tuning and you need a sort of minimum number of GPUs. So there's a bit of a deterrence effect here, but it's not going to stop really capable, well resourced attackers. So I

(01:11:07) think for that we're going to need new approaches, but one I'm most optimistic about in the short term is pre-training filtering. So it's a simple idea where just don't train the models on really dangerous stuff. If you don't need your model to help people make anthrax to don't train it on the anthrax papers. A a tiny number of users might be a little bit sad that it can't answer questions

(01:11:27) about this but most people won't even notice. But it has a big impact on the sort of misuse potential of the model. And this has been validated in a number of scientific papers. This is independent researchers. UK's AI security inropic has been sponsoring some research into this. Open a actually used this in their GPD OSS release. So, it's been tested quite well, but it's not become kind of common practice. And

(01:11:51) this is something we're actively excited about scaling to make sure this does work and a near frontier approach. And going back to what you were saying earlier about could we start sharing some of these data sets or filters with people. Our plan is to open source as much as we think it doesn't have a misuse potential and then privately share with developers of things that that do have misuse potential that could

Around this claim