Mechanism◆Video · 32:53 · 2m
Using accumulated human tax-expert feedback on model errors as ground truth and reward signal made OpenAI's reinforcement fine-tuning (RFT) especially effective for Sphere's tax-determination task, since it targeted precisely the hard cases the model had previously missed.
Alex Boucott · The TWIML AI Podcast