ATRIUMsearch → argument graph
ClaimVideo · 0:00 — 0:21

Looking good and being correct are not the same thing in text-to-image generation — models can produce photorealistic images while failing to follow specific instructions about composition, identity, number of subjects, or diversity.

Text-to-image models have reached impressive visual quality, but they still fail at precision tasks like generating distinct faces, maintaining correct subject counts, and following specific compositional instructions. ✦ AI generated

Sam Sherington · The TWIML AI Podcast · 2026-08-12 · original ↗

starts at this moment · 0:00

Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those details.

verbatim transcript · starts at 0:00

Transcript · around this moment

0:00Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those

0:21details. Push towards higher resolution or local generation, and quality, speed, and memory quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high-quality rendition of the image you actually ask for. One researcher at the forefront of this work is Fati Periqi, vice president of technology at Qualcomm, whose team

0:46presented more than 20 papers at this year's CVPR, the computer vision and pattern recognition conference. Here's Fati explaining why image generation still has plenty of hard problems to solve. Maybe we are also asking a single model to solve too many difficult problems at once. Think about what happens when you generate a scene with several people. Now the model has to understand the problem. that I'm asking to model and then decide how many people

Around this claim