GPT-3 was surprisingly capable at legal reasoning, going largely unnoticed and unused despite its potential.
Winston Weinberg discovered that GPT-3 could produce lawyer-quality legal answers when tested on landlord-tenant questions, with 86% of answers deemed acceptable by attorneys — a finding that even surprised OpenAI's leadership. ✦ AI generated
Winston Weinberg · No Priors · 2025-10-31 · original ↗
plays this moment only · 0:39 — 2:00
And what had happened was he showed me GPT-3, which at the time was public and I was, first of all, just incredibly surprised that no one was talking about GPT-3 and no one was using it in any way, shape, or form. And he showed me that and I showed him kind of my legal workflows. And we started, the kind of aha moment was we went on r/legaladvice, which is basically a subreddit where people ask a bunch of legal questions and almost every single answer is 'So who do I sue?' Almost every single time. And we took about 100 landlord-tenant questions and we came up with kind of some chain of thought prompts. And this is before anyone was talking about chain of thought or anything like that. And we applied it to those landlord-tenant questions and we gave it to three landlord-tenant attorneys. And we just said nothing about AI. We just said, here's a question that a potential client asked and here's an answer. Would you send this answer without any edits to that client? Would you be fine with that? Is it ethical? Is it a good enough answer to send? And 86 out of 100 was yes.
verbatim transcript · starts at 0:39
(00:00:06) 2025 has been another remarkable year in AI. (00:00:08) This week on No Priors, we're sharing our favorite moments from the podcast from the year so far. (00:00:13) We've talked to visionary leaders at... (00:00:15) Harvey, OpenAI, Glean, A Bridge, and more. (00:00:17) We also talked to legends of science like Dr. (00:00:19) Fei-Fei Li and Noubar Afeyan. (00:00:21) But first, let's start with a moment that captures the magic of leaning into new capabilities at the right time. (00:00:26) Harvey CEO Winston Weinberg discovered an extraordinary opportunity hidden in plain sight. (00:00:31) Gabe and I actually had met a couple of years before, and I definitely didn't know anything about the startup world and didn't have a plan of doing a startup. (00:00:39) And what had happened was he showed me GPT-3, which at the time was public and (00:00:45) And I was, first of all, just incredibly surprised that no one was talking about GPT-3 and no one was using it in any way, shape, or form. (00:00:52) And he showed me that and I showed him kind of my legal workflows. (00:00:57) And we started, the kind of aha moment was we went on r/legaladvice, which is basically a subreddit where people ask a bunch of legal questions and almost every single answer is (00:01:11) So who do I sue? (00:01:12) Almost every single time. (00:01:13) And we took about 100 landlord-tenant questions and we came up with kind of some chain of thought prompts. (00:01:20) And this is before anyone was talking about chain of thought or anything like that. (00:01:24) And we applied it to those landlord-tenant questions and we gave it to three landlord-tenant attorneys. (00:01:29) And we just said nothing about AI. (00:01:31) We just said, here's a question that a potential client asked and here's an answer. (00:01:36) Would you send this answer without any edits to that client? (00:01:39) Would you be fine with that? (00:01:40) Is that ethical? (00:01:41) Is it a good enough answer to send? (00:01:44) And 86 out of 100 was yes. (00:01:48) And actually, we cold emailed the general counsel of OpenAI, and we sent him these results. (00:01:53) And his response basically was, oh, I had no idea the models were this good at legal. (00:01:58) And we met with the C-suite of OpenAI a couple weeks after. (00:02:02) Now, from legal reasoning to spatial intelligence, the legendary Dr. (00:02:06) Fei-Fei Li opened our eyes to an entirely different dimension of AI capability. (00:02:10) I think from a neural and cognitive science point of view that spatial intelligence is a really hard problem that evolution has to solve for animals. (00:02:22) And what's really interesting is I think animals have solved it to an extent, but not fully solved it. (00:02:28) It's one of the hardest problem because what is the problem animal has to solve? (00:02:35) Animals have to evolve the capability of collecting lights in something, which we call eyes mostly. (00:02:45) And then with that collection of eyes, it has to reconstruct a 3D world in their mind somehow so that they can navigate. (00:02:56) and they can do things, and of course, they can interact. (00:03:00) For humans, we're the most capable animal in terms of manipulation. (00:03:04) We can do a lot of things. (00:03:06) And all this is spatial intelligence. (00:03:09) To me, that's just rooted in our intelligence. (00:03:14) What is interesting is it's not a fully solved problem, even in animals. (00:03:20) We, for example, for humans, right? (00:03:25) If I ask you to close your eyes right now and draw out or build a 3D model of the environment around you, it's not that easy. (00:03:35) We don't have that much capability to generate extremely complicated 3D model till we get trained. (00:03:44) You know, there are some of us, whether they're architects or designers or just people with a lot of training and a lot of talent, (00:03:53) And that's a hard thing to do. (00:03:56) And imagine you do it at your fingertip much more easily and allow much more fluid interactivity and editability. (00:04:07) That would just be (00:04:09) a whole different world for people, no pun intended. (00:04:13) Data is the beast feeding the AI train, and thus Merck War CEO, Brendan Foody, is working with major AI labs on how to build what's next. (00:04:21) He gives a clear prediction about what's coming for the workforce. (00:04:25) I think displacement in a lot of roles is going to happen very quickly, and it's going to be very painful. (00:04:36) and a large political problem. (00:04:38) Like I think we're going to have a big populist movement around this and all the displacement that's going to happen. (00:04:43) But one of the most important problems in the economy is figuring out how to respond to that, right? (00:04:50) Like how do we figure out what everyone who's working in customer support or recruiting should be doing in a few years? (00:04:57) How do we reallocate wealth (00:04:59) once we have, once we approach super intelligence, especially if the value and gains of that are more of a power law distribution. (00:05:10) And so I spend a lot of time thinking about like how that's going to play out. (00:05:14) And I think it's really at the heart of it. (00:05:15) What do you think happens eventually? (00:05:17) X percent of people get displaced from like color work. (00:05:20) What do you think they do? (00:05:21) I think there's going to be a lot more in the physical world. (00:05:24) I think that (00:05:26) What does the physical world mean? (00:05:30) Well, it could be everything ranging from people that are creating robotics data to people that are waiters at restaurants or are just like therapists because people want like human interaction. (00:05:46) Like whatever that looks like. (00:05:48) I think all of, I think that automation in the physical world is going to happen (00:05:54) Lot slower than what's happening in the digital world, just because of so many of the self-reinforcing gains and... (00:06:05) a lot of self-improvement that can happen in the virtual world, but not physical one. (00:06:10) Which brings us to one of the biggest questions of our time. (00:06:13) How do we navigate the geopolitical implications of superintelligence? (00:06:17) Dan Hendricks, the director of the Center for AI Safety, has an answer. (00:06:21) Let's think of what happened in nuclear strategy. (00:06:24) Basically, a lot of states deterred each other from doing a first strike because they could then retaliate. (00:06:30) They had a shared vulnerability. (00:06:32) So they were, we're not going to do this really (00:06:35) aggressive action of trying to make a bid to wipe you out because that will end up causing us to be damaged. (00:06:42) And we have a somewhat similar situation later on when AI is more salient, when it is viewed as pivotal to the future of a nation. (00:06:51) When people are on the verge of making a superintelligence more, when they can, say, automate pretty much all AI research, I think states would try to deter each other from trying to leverage that to (00:07:05) develop it into something like a super weapon that would allow the other countries to be crushed, or use those AIs to do some really rapid automated AI research and development loop that could have it bootstrapped from its current levels to something that's super intelligent, vastly more capable than any other system out there. (00:07:24) I think that later on, it becomes so destabilizing that China just says, we're going to do something preemptive, like do a cyber attack on your data center. (00:07:33) And the US might do that to China. (00:07:35) And Russia, coming out of Ukraine, will reassess the situation, get situationally aware, think, oh, what's going on with the US and China? (00:07:45) Oh my goodness, they're so head on AI. (00:07:47) AI is looking like a big deal. (00:07:48) Let's say it's later in the year when a big chunk of software engineering is starting to be impacted by AI. (00:07:54) Oh, wow, this is looking pretty relevant. (00:07:56) Hey, if you try and use this to crush us, we will prevent that by doing a cyberattack on you. (00:08:01) And we will keep tabs on your projects because it's pretty easy for them to do that espionage. (00:08:06) Noubaraf Feyan has been thinking about how biotech gets built and how to change the game for three decades. (00:08:11) His breakthroughs have impacted global health. (00:08:13) He's the founder and CEO of Flagship Pioneering and the co-founder of Moderna. (00:08:17) He wants to make entrepreneurship a scientific effort, not a random one, and he thinks AI can help. (00:08:22) The motivation for Flagship (00:08:25) stems from what I was doing before, which was that I started a company in 1987 when 24-year-old immigrants didn't start companies in this country, but instead it was kind of like former Merck senior executives or IBM senior executives were the only ones who were entrusted with the massive amounts of venture capital, namely two, $3 million per round used to go into venture capital. (00:08:47) So this was very early days. (00:08:49) And I had the kind of chance, (00:08:52) opportunity to start a company right out of my graduate school and ended up raising quite a bit of venture money and eventually kind of went down a path of entrepreneurship. (00:09:02) Along the way, one of the things that interested me was why it is that kind of the entrepreneurial process was supposed to be random, improvisational, kind of idiosyncratic, almost emotional, gamey, (00:09:18) All of those things I kind of thought was a bit of a put-off when it comes to actually doing things in a serious, professional way. (00:09:26) And I kind of used to go around in the very early 90s saying, why isn't... (00:09:30) entrepreneurship a profession. (00:09:32) And if it was going to be a profession, how could it be a profession? (00:09:35) What do you mean by gaming? (00:09:36) Because it's like supposed to fail most of the time and once in a while you win and then you celebrate the win. (00:09:42) And what I mean is like it. (00:09:44) It's random. (00:09:45) But not only random, but there's like winners and losers. (00:09:48) and keeping score. (00:09:49) I don't know, it's maybe the wrong word, but I just mean like people even call gamification in the software space. (00:09:55) There is a version of this, like I don't mind being playful because if you're overly serious, sometimes you miss things, but it can't just all be played. (00:10:04) We take hard-earned money, we deploy it to do things that are damn near impossible. (00:10:10) Once in a while, we reduce them to practice so they become not only possible but valuable. (00:10:14) And yet, (00:10:15) People treat it like, oh, well, it didn't work. (00:10:18) There's 20 different things we tried. (00:10:19) One of them worked. (00:10:20) That, I don't know, as an engineer by background, as a scientist, I just thought that what we do, especially, listen, in healthcare, especially in climate, especially in agriculture, food security, you can't think of this as shots on goal and this. (00:10:35) Now, you've got to say, hey, we can get better at this. (00:10:38) Reasoning is the biggest paradigm shift in AI architecture since the Transformer. (00:10:42) Brandon McKenzie and Eric Mitchell from OpenAI explained a crucial insight about (00:10:45) reasoning models? (00:10:46) I can give maybe very concrete cases for like the visual reasoning side of things. (00:10:53) There's a lot of cases where, and back to also the model being able to estimate its own uncertainty, you'll give it some kind of question about an image and the model will very transparently tell you, I shouldn't have thought like, I don't know, I can't really see the thing you're talking about very well. (00:11:06) Or like it almost knows like that its vision is not very good. (00:11:11) But what's kind of magical is when you give it access to a tool, it's like, okay, well, I got to figure something out. (00:11:16) Let's see if I can manipulate the image or crop around here or something like this. (00:11:20) And what that means is that it's much more productive use of tokens as it's doing that. (00:11:26) And so your test time scaling slope goes from something like this to something much deeper. (00:11:31) And we've seen exactly that. (00:11:34) The test time scaling slopes for without tool use and with tool use for visual reasoning specifically are very noticeably different. (00:11:41) like for like writing code for something like there are a lot of things that an LLM could try to figure out on its own, but would require a lot of attempts and self verification that you could write a very simple program to do in like a verifiable and, you know, much faster way. (00:12:04) So (00:12:06) I do some research on this company and use this type of valuation model to tell me what the valuation should be. (00:12:15) You could have the model try to crank through that and fit those coefficients or whatever in its context, or you could literally just have it write the code to just do it the right way and just know what the actual answer is. (00:12:28) And so, yeah, I think part of this is you can just allocate compute a lot more efficiently because you can (00:12:35) defer stuff that the model doesn't have comparative advantage to doing to a tool that is really well suited to doing anything. (00:12:42) Sometimes the most profound moments in AI development aren't the grand theoretical breakthroughs. (00:12:46) They're based on taste, data generation, and grinding work. (00:12:49) The visceral experience of watching something you hoped would work actually come to life. (00:12:52) Issa Falford from OpenAI captures that moment perfectly. (00:12:56) Here, she's describing the training that went into deep research. (00:12:59) It really was one of those things where we thought that training on browsing tasks would work. (00:13:05) It felt like we had good conviction in it, but actually the first time you train a model on a new data set using this algorithm and seeing it actually working and playing with the model was pretty incredible, even though we thought it would work. (00:13:20) So honestly, just that it worked (00:13:24) so was pretty surprising, even though we thought it would, if that makes sense. (00:13:28) Yeah, it's a risk-real experience of like, oh, the path is paved with strawberries or whatever. (00:13:33) Exactly. (00:13:34) But then sometimes some of the things that it fails out are also surprising. (00:13:37) Like sometimes it will make a mistake where it will do such smart things and then make a mistake where I'm just thinking, why are you doing that? (00:13:44) Stop. (00:13:45) So I think there's definitely a lot of room for improvement. (00:13:47) But yeah, we've been impressed with the model so far. (00:13:50) One of the biggest surprises of AI and a core principle for us here at Conviction is how it can make bad markets suddenly good ones. (00:13:57) The right technology can meet the right moment in unexpected ways. (00:14:00) Arvind Jain built Glean and what everyone said was a graveyard market, enterprise search. (00:14:05) It was like a graveyard of all these companies that tried to solve the problem and it didn't. (00:14:10) Part of it was just that I think search is a hard problem. (00:14:13) In an enterprise, even getting access to all the data that you want to search, (00:14:18) It was such a big problem. (00:14:19) In the pre-SaaS world, there was no way to sort of go into those data centers, figure out where the servers were, where the storage systems were, try to connect with information in them. (00:14:29) It was a big challenge. (00:14:30) So SaaS actually solved that issue. (00:14:32) So like search products, like most of them, most of those companies started in the pre-SaaS world, they failed because you just couldn't build a turnkey product. (00:14:39) But SaaS actually allowed you to actually build something, you know, which is my insight was that like, look, you know, the enterprise world has changed. (00:14:47) We have these SaaS (00:14:48) systems now, and SaaS systems don't have versions. (00:14:51) Like everybody, all customers have the same version. (00:14:55) They're open, they're interoperable. (00:14:57) You can actually hit them with APIs and get all the content. (00:15:00) I felt that the biggest problem was actually solved, which was that I could actually easily go and bring all the enterprise information and data in one place and build this unified search system on top. (00:15:11) So that was actually a big unlock. (00:15:12) And by the way, the origins of Glean is, so at Rubrik, we had this problem. (00:15:16) We grew fast. (00:15:17) We had a lot of information (00:15:18) across 300 different SaaS systems and nobody could find anything in the company. (00:15:22) And people were complaining about it in surveys. (00:15:25) And I always run IT in my startups. (00:15:28) And so there's a complaint that came to me, I had to solve it. (00:15:31) So I tried to buy a search product and I realized there's nothing to buy. (00:15:34) I mean, that's really the origins of how Gleam got started as a company. (00:15:38) And so that was one big issue. (00:15:41) So SaaS made it easy to actually connect your enterprise data and knowledge to a search system. (00:15:46) So that actually made it possible (00:15:48) for us to, for the very first time, build a turnkey product. (00:15:51) But there are a lot of other advances as well. (00:15:53) One is, like, look, businesses have so much information and data. (00:15:56) One interesting fact, one of our largest customers, they have more than 1 billion documents inside their company. (00:16:03) Now, hear this, when Elad and I, when we were working on search at Google in 2004, the entire internet was actually 1 billion documents. (00:16:11) There's a massive explosion of content inside businesses, so you have to build scalable systems. (00:16:17) And you couldn't build a system like that before in the pre-cloud era. (00:16:21) Perhaps no story captures the human impact of this AI moment and its potential better than what's happening in healthcare. (00:16:27) Here's Shiv Rao, CEO and founder of Abridge. (00:16:30) It's pretty heroic in general for a doctor to give you feedback like, Hey, this sucked and you got to do better. (00:16:35) You didn't recognize the way I said this medication or (00:16:40) I'm A gastroenterologist and I would never, sequence my problems in my assessment and plan section of my note this way. (00:16:46) It doesn't serve me well and makes me look like terrible as a doctor or whatever. (00:16:49) We get that feedback. (00:16:50) We love it. (00:16:50) It's oxygen. (00:16:52) But then we also get the feedback that's like, Hey, this is amazing and I'm not gonna retire anymore. (00:16:56) And I've got like years, decades left in my career now, thanks to this technology. (00:17:01) But in this Channel Love Stories, all of that feedback, that positive feedback, we just get it like programmatically funneled. (00:17:06) So any one of our people inside of the company can always go into that channel and it's like purpose, it's like fulfillment immediately. (00:17:14) Like you immediately understand why we're all working so hard and why it makes sense because like being on this, (00:17:20) very telephone pole-like journey these last couple years is obviously, like, it's news for so many of us, and we're all kind of building new muscles, but it's a lot of pressure. (00:17:31) But this is my favorite bit of feedback. (00:17:32) So this love story comes from a doctor at Tanner Health, which is a rural health system. (00:17:37) And she wrote to us. (00:17:38) She wrote, I was sitting at dinner last week, and my son asked me, 'Mommy, why aren't you working right now?' I literally took my phone out and explained to him that Abridge is a new tool that lets mommy come home early and eat dinner with her family. (00:17:50) I started to tear up and looked over at my husband, who then said, 'Mommy's going to be able to eat dinner with us every night now.' (00:17:57) And we get feedback like that every day. (00:18:00) And so there's dopamine hits in hypergrowth, and those are awesome, but I think that they get us through sprints. (00:18:08) But I think it's the oxytocin hits like this. (00:18:11) It's the purpose, it's the fulfillment. (00:18:12) It's like, that's, I think, what I think we're really after in this company. (00:18:15) And so everybody's mission driven out there, but I think this mission, (00:18:20) Like, it hits me at least a little bit different. (00:18:22) These conversations remind us that we're living through a hinge moment in history. (00:18:26) Stay tuned as we have more conversations with the builders and thinkers leading the way for the rest of the year. (00:18:31) If you like what we're doing, leave us a review on Apple Podcasts or Spotify, comment on YouTube, or let us know who we should have as a guest. (00:18:38) Thanks for listening. (00:18:41) Find us on Twitter @NoPriorsPod. (00:18:44) Subscribe to our YouTube channel if you wanna see our faces, follow the show on Apple Podcasts, (00:18:49) Spotify, or wherever you listen. (00:18:50) That way you get a new episode every week. (00:18:52) And sign up for emails or find transcripts for every episode at no-priors.com.
- ·GPT-3 could produce lawyer-quality legal answers
- ·86% of answers deemed acceptable by attorneys
- ·Finding surprised even OpenAI's leadership
- ·Capability went largely unnoticed at the time
- ·Tested 100 landlord-tenant questions from Reddit
- ·Used chain-of-thought prompts before technique was named
- ·Three attorneys reviewed answers blind (no AI disclosure)
- ·Asked: would you send this to a client without edits?