The core limitation of current text-to-image models is not image quality but controllability — they fail to generate distinct identities when asked for multiple people, producing near-identical faces instead.
Despite photorealistic outputs, current T2I models struggle with identity preservation across multiple subjects. Training objectives optimize for realism and prompt matching but do not explicitly encourage facial diversity, leading to the same face duplicated across groups.
transcript
Fati Periqi: This is not an image quality problem the problem that I mentioned before like we are asking the model to generate faces and and certain number of faces and it keeps generating same faces identical almost identical faces over and over again. So quality wise, image quality wise, if you look at the pixels and noise and everything, it looks realistic. But the missing piece was that those models, base models, amazing models had not really learned to create truly distinct identities because existing training objectives focus heavily on the realism and matching the user prompt, but they don't explicitly encourage diversity between people.
explains mechanism · 2provides context · 4supports · 2