Separating image generation into a planning stage (architect) and a rendering stage (artist) mirrors how human artists work and produces more natural compositions.
R2CAN divides image generation into an architect that plans scene structure and an artist that renders photorealistic images, inspired by how human artists approach composition before details. This separation enables better control and more natural-looking results. ✦ AI generated
Fati Periqi · The TWIML AI Podcast · 2026-08-12 · original ↗
starts at this moment · 22:53
“Talk us through that.”
maybe it's easier to approach some of these red AI challenges like a human being. We do not like try to solve everything ourselves. Well, not just ourselves, but even like even in this domain, you know, art like you an artist approaching this problem wouldn't necessarily think about it pixel by pixel. They think about like if you know what's the subject, what's the background and there's a planning, right? In this paper, we build on it. Build on that idea. Instead of asking model to do everything, we separate planning from rendering similar to how a human artist would work. So we have two components there. Architect which is it doesn't generate pixels but instead it creates the structure or composition for the scene. Then kind of it decides where for inance if there's a person is going to be in this image where it should appear and if there are multiple people how they should be arranged you know that it should look natural and you know realistic then artist starts with that composition structure and generates the final photorealistic image while of course from one objective also preserve identities if we provide identity.
verbatim transcript · starts at 22:53
22:53agentic flow again uh maybe it's easier to approach some of these red AI challenges like a human being. We do not like try to solve everything ourselves. Well, not just ourselves, but even like even in this domain, you know, art like you an artist approaching this problem wouldn't necessarily think about it pixel by pixel. They think about like if you know what's the subject, what's the background and
23:24>> there's a planning, right? >> There's a planning aspect to it. Yeah. Yeah. >> In this paper, we build on it. Build on that idea. Instead of asking model to do everything, we separate uh planning from rendering similar to how a human artist would work. Um so we have two components there. Architect which is uh it doesn't generate pixels but instead it creates the structure or composition for the
23:52scene. Then uh kind of it decides where for inance if there's a person is going to be in this image where it should appear and if there are multiple people how they should be arranged you know that it should look natural and you know realistic then artist starts with that composition structure and generates the final photorealistic image while of course uh from one objective also preserve identities if we provide
24:19identity. So yeah this is like a um planning an architect then an artist type of framework. >> The underlying technical approach is also using RL just like disco. It's also gpo based approach. >> You are right. Um we also have this reward function leveraging on different objectives that we want to optimize. In this goal we had in intraim image and intra group like interim image at human
24:54perception score and count accuracy. Here in R2 we also have uh well we have the composition can be put the right face into right place in the composition type of objective which is a part of the reward function for GRPO and uh also pose of the phase because now we are composing I mean we don't want one people to look this way the other people look the other way if let's say we are
- ·Human artists plan composition before details
- ·Architect handles scene structure, not pixels
- ·Artist renders final photorealistic image
- ·Enables natural compositions and identity preservation
- ·Architect plans subject placement and arrangement
- ·Artist generates photorealistic output from structure
- ·Mimics human creative process
- ·Better control over scene composition
- ·Takes composition from architect
- ·Generates photorealistic final image
- ·Preserves provided identity information
- ·Focuses solely on visual output