ATRIUMsearch → argument graph
ExampleVideo · 20:00 — 24:40

Claude Opus 5 produced dramatically better results in the Standard Agent Zen Garden challenge compared to other models, especially in generating detailed 3D interactive scenes from a single prompt.

In a blind test where AI models were given a single prompt to build an interactive 3D Zen garden, Claude Opus 5 produced results that were noticeably superior — featuring interactive elements like lighting lanterns and feeding koi. ✦ AI generated

Scott · Syntax · 2026-07-27 · original ↗

starts at this moment · 20:00

Opus 5 whooped some ass in this challenge. Like, this this was a beatdown. Um, and when I looked at all of these, I could not believe how much better Opus Fives was. So, like if you come to Opus F. Okay. Sounds beautiful. [...] this one is so much better. You can light lanterns and stuff in it. You can feed koiish. Uh, it's crazy how much better this one is.

verbatim transcript · starts at 20:00

Transcript · around this moment

19:58>> Cool. All right. I think uh let's keep it moving. We got a lot of stuff to talk about today. >> Oh, is it is this me? Sorry. Am I supposed to be >> Wes with the critique about professionalism supposed to be really good at these transition folks? >> And it says Wes throw to Scott. Scott [laughter] tell us about standard agent Zen Garden. That sounds interesting.

20:24Oh, it is interesting, Wes. Let me tell you, it is interesting. Uh, basically, all models, all AI models were given a single prompt to build a Zen garden that you could walk through, and then there is a blind test to determine which one is better. So, I wanted to throw this out here. I know one shots are maybe not the the most interesting thing, but in the same regard, I think this was pretty

20:52fascinating to see. So, uh there they give you a couple at least here you can do and then you can explore all of them afterwards. So, the idea is is that you were given the the task of creating a garden that you could walk through and then you would evaluate which models did a better job of them. This is at model zenyphen garden.standard aages.ai. We'll have the link in the the video comment.

21:17So this is one of them. This is the moon water garden. And you can see the job is done. Okay. Then we can say actually it's got like a little compass here and stuff. Then we can say see the next garden. And then it's going to ask you to compare them. So here's whispering stones. It's actually interesting these little um dialogues they came up with too. Oh,

21:37this one has sound apparently. I don't know why it has sound. Um, how are we feeling about these two Zen Garden attempts? >> Pretty cool. >> Uh, I I loaded it up and both of yours look cooler than mine. >> Yes, I think that's a thing. So, if we say >> enough, it's it's really annoying. >> Yeah, you're making me sleepy. >> Yes, obviously Garden 2 I thought was

21:58better even though it had some white noise. I'm actually surprised the sound got picked up onto the stream. I don't think I have loop back going. Either way, this was my pick. It was Grock 4.5 that did the second one where GPT56 Luna did the first one. So, uh, no surprise that Luna wasn't very good. I actually wanted to even throw to you guys about that. I know Wesid hasn't used GPT56

22:23yet. I haven't had great experiences with Soul myself, but um, maybe it's just because I'm doing more UI kind of UX things. Either way, the one that I thought was the most interesting throughout this because you can go look at the rest of the gardens after the fact, the thing that I was struck most by this, which will also lead into the idea of Opus 5 being released, is that

22:51Opus 5 whooped some ass in this challenge. Like, this was this was a beatdown. Um, and when I looked at all of these, I could not believe how much better Opus Fives was. So, like if you come to Opus F. Okay. Sounds beautiful. >> Let me let me mute this tab. Yes, I'm Yes. Yes. I'm muting the tab. >> Yeah. And can you can you also give us a

Around this claim