Attacker-borne, intentionally misaligned models are the harder danger but will eventually appear; current alignment techniques are not merely surface-thin and do exert meaningful influence.
Notes the flip side of helpfulness — explicitly training a misaligned model would be easy to use but harder to build, given compute shortages and alignment-encouraging training data, vindicating existing alignment techniques.
transcript
Nathan Lambert: The other side of the helpfulness example above is that it is clear someone could make this happen much more easily if they wanted to by explicitly training a misaligned model. To reiterate, this would be making a system that is easier to use for finding exploits at inference, but I think it'll be harder to train said model. I think this'll take longer than most commentators expect, as nearly all the strong public models and data industry existing to date encourage alignment (and it seems very hard for bad actors to get enough compute to train these models end-to-end, as all leading companies are in a compute shortage as well). We should take a moment to appreciate that the alignment techniques we are employing on current models have a meaningful influence and are not merely surface thin as some have worried. Downstream models have a propensity for mirroring their teacher's character.