Using accumulated human tax-expert feedback on model errors as ground truth and reward signal made OpenAI's reinforcement fine-tuning (RFT) especially effective for Sphere's tax-determination task, since it targeted precisely the hard cases the model had previously missed.
Alex explains that Sphere's existing pipeline of tax-expert feedback on model mistakes provided ideal ground truth and signal for OpenAI's RFT alpha program, yielding accuracy improvements now used in production. ✦ AI generated
Alex Boucott · The TWIML AI Podcast · 2026-06-09 · original ↗
starts at this moment · 32:53
“You also experimented with using fine-tuning, uh RFT in particular, for your process. Can you talk a little bit about where it fits in?”
What we had that was very useful was feedback from our human tax experts every time the model T RAM had gotten something wrong on a determination... we had the ground truth, we had signal, and we knew that these were hard problems that the model had missed previously. And so, that was a really good recipe for RFT, and we saw uh improvements with, uh, during the alpha program with with OpenAI on RFT.
verbatim transcript · starts at 32:53
32:53need to provide a grader. And, um, what we had that was very useful was feedback from our human tax experts every time the model T RAM had gotten something wrong on a determination. So, as the tax experts are reviewing, when the model is incorrect, they leave feedback, and they they give that feedback, um, similar to how they would give feedback to like a colleague who had maybe, you know, a a
33:19more junior colleague that had made this determine text blurb about, you know, what they thought. And explain it, you know, an an explanation in a way where you want that person to get better, and you want that person to have this, you know, extra context that maybe isn't clear from just the legislation. So, some, you know, background information about how Alabama treats a certain vocabulary word, something like that. And what we
33:43found was that was a very, so I guess twofold, we had already like a set of questions that we knew the model struggled with today because it had missed them, and then we had a way to give really great signal through the feedback and through the fact that, of course, we had the correct answer. Like they they the tax experts fixed the issue, of course, um, and then they also
34:07leave the feedback. So, we had the ground truth, we had signal, and we knew that these were hard problems that the model had missed previously. And so, that was a really good recipe for RFT, and we saw, uh, improvements with, uh, during the alpha program with with OpenAI on RFT, and that's what we use in production today is a well, a different model that that we've worked with them
34:28to to RFT, um, but we've seen performance or accuracy improvements. And that really is the key for us is accuracy. We track it very closely. I'm always checking in on it. We want to know how accurate is the model being. And accurate means, you know, how often is the tax expert having to make an adjustment to the model's work. >> And I'm curious your experience with like
- ·Human tax experts flagged every model determination error
- ·Errors gave ground truth plus reward signal for RFT
- ·Feedback targeted hard cases model previously missed
- ·Improvements seen during OpenAI's RFT alpha program
- ·Sphere joined OpenAI's RFT alpha program early
- ·Expert-labeled error data fed reinforcement fine-tuning
- ·Resulting accuracy gains now used in production