ATRIUMsearch → argument graph
MechanismVideo · 23:21 — 24:19

Separating planning from rendering in image generation — using an architect to determine composition and an artist to generate pixels — mirrors how human artists work and produces more natural, controllable results.

R2C splits image generation into two stages: an architect that plans composition and subject placement, and an artist that renders the final image, allowing each component to be independently optimized for better results. ✦ AI generated

Fatih Porikli · The TWIML AI Podcast · 2026-08-12 · original ↗

starts at this moment · 23:21

Elicited by

Talk us through the R2C paper and how this model breaks down the problem of text to image generation.

In this paper, we build on that idea. Instead of asking model to do everything, we separate planning from rendering similar to how a human artist would work. So we have two components there. Architect which doesn't generate pixels but instead it creates the structure or composition for the scene. Then it decides where for instance if there's a person is going to be in this image where it should appear and if there are multiple people how they should be arranged, that it should look natural and realistic then artist starts with that composition structure and generates the final photorealistic image while of course from one objective also preserve identities if we provide identity.

verbatim transcript · starts at 23:21

Transcript · around this moment

23:24>> there's a planning, right? >> There's a planning aspect to it. Yeah. Yeah. >> In this paper, we build on it. Build on that idea. Instead of asking model to do everything, we separate uh planning from rendering similar to how a human artist would work. Um so we have two components there. Architect which is uh it doesn't generate pixels but instead it creates the structure or composition for the

23:52scene. Then uh kind of it decides where for inance if there's a person is going to be in this image where it should appear and if there are multiple people how they should be arranged you know that it should look natural and you know realistic then artist starts with that composition structure and generates the final photorealistic image while of course uh from one objective also preserve identities if we provide

24:19identity. So yeah this is like a um planning an architect then an artist type of framework. >> The underlying technical approach is also using RL just like disco. It's also gpo based approach. >> You are right. Um we also have this reward function leveraging on different objectives that we want to optimize. In this goal we had in intraim image and intra group like interim image at human

Around this claim