ATRIUMsearch → argument graph
ContextVideo · 0:00 — 0:46

Image generation has progressed to producing plausible-looking images, but a significant gap remains between plausible outputs and precise, controllable results that faithfully match user prompts.

Text-to-image models now produce visually impressive images, but they struggle with specific compositions, distinct identities, higher resolutions, and reliable controllability. The next frontier is closing this gap. ✦ AI generated

Sam Sharenton · The TWIML AI Podcast · 2026-08-12 · original ↗

starts at this moment · 0:00

Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those details. Push towards higher resolution or local generation, and quality, speed, and memory quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high-quality rendition of the image you actually ask for.

verbatim transcript · starts at 0:00

Transcript · around this moment

0:00Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those

0:21details. Push towards higher resolution or local generation, and quality, speed, and memory quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high-quality rendition of the image you actually ask for. One researcher at the forefront of this work is Fati Periqi, vice president of technology at Qualcomm, whose team

0:46presented more than 20 papers at this year's CVPR, the computer vision and pattern recognition conference. Here's Fati explaining why image generation still has plenty of hard problems to solve. Maybe we are also asking a single model to solve too many difficult problems at once. Think about what happens when you generate a scene with several people. Now the model has to understand the problem. that I'm asking to model and then decide how many people

1:14should appear uh determine where they should be placed like the composition of the scene reason about their interactions because if there's a person if there's another person most likely there there is some connection preserve the identity of the person we can give okay this is my daughter this is my son and I want them to be in the picture not like any random person and finally render everything together in all a

Around this claim