Building a judgment agent that matches expert consensus requires a three-part flywheel: distilling expert reasoning through techniques like consequence mapping and cross-expert debate, pressure-testing the resulting rubric with a second group of experts applying it to real labels, and iterating until disagreement reveals and resolves ambiguity.
Robbie Goldfarb describes Forum AI's three-step method—expert reasoning distillation, pressure-testing rubrics against real labels, and iterating until consensus emerges—for building AI judges that actually match how human experts judge nuanced cases. ✦ AI generated
Robbie Goldfarb · The Cognitive Revolution · 2026-06-26 · original ↗
starts at this moment · 41:02
“How do you how do you actually do this? How do you select the uh experts? How do you um obtain the necessary data in order to create this distillation and how do you apply it?”
the first piece is what we call um high high level conversations. We we call it um expert reasoning distillation... For example, consequence mapping, thought process mapping, um, edge case testing where we pressure taste the edge cases. We do, uh, crossexpert debates where we'll actually have them discuss things.
verbatim transcript · starts at 41:02
41:02um expert reasoning distillation. And we've come up with a number of creative techniques for having conversations with these experts that look nothing like traditional data labeling. Um, and we could talk through some of them. For example, consequence mapping, thought process mapping, um, edge case testing where we pressure taste the edge cases. We do, uh, crossexpert debates where we'll actually have them discuss things. And through that we kind of take a first
41:29crack at developing a set of rubrics and context graphs to represent how they think about a given factor. Let's say political bias. But then and this is what we found to be the most important step. We pressure test it. So we get another group of experts to then go and look at that guidance and actually try and apply it with real labels. And what will almost always happen is experts
41:55will disagree, right? They'll apply the date labels differently. And where experts disagree indicates where there is ambiguity in your rubric. So then we go back to step one and you sort of have this flywheel where you kind of have these highle conversations to try and wrestle with ambiguity to actually applying it. And eventually when you go through that enough times you get to a point where you have um first of all you
42:21have kind of rubrics and context that represent expert consensus. But then you also have a gold like golden data that you can then use to step three which is calibrate a judge for in this case assessing factual accuracy. Um I could go into more detail but that that that's the gist. So does this then become a RL signal? Is that kind of ultimately how all this work gets translated into improved
42:51frontier model performance? >> Yep. Increasingly we're doing more of that. I think until now a lot of the work we've done now is we've done over the past several months is more just evaluation. um both kind of like largecale comprehensive evaluations like we launched our public facing news bench benchmark to kind of give more consumers and and executive level a sense of where um a system is at. Um but of course step