Looking good and being correct are not the same thing in text-to-image generation — models can produce photorealistic images while failing to follow specific instructions about composition, identity, number of subjects, or diversity.
Text-to-image models have reached impressive visual quality, but they still fail at precision tasks like generating distinct faces, maintaining correct subject counts, and following specific compositional instructions.
transcript
Sam Sherington: Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those details.