
Most eval advice assumes you can check the answer. But a lot of high-stakes LLM work has no definitive answer- like qualitative coding, thematic analysis, narrative findings. These are outputs where two expert humans will legitimately disagree, and where "correct" is a judgement rather than a match. In UK government social research, those outputs feed policy decisions, so "the model seemed confident" is not an acceptable standard of proof. This talk covers how we've approached validation in that environment: separating the parts of a pipeline that do have ground truth from the parts that never will, building statistical validation harnesses for the former, and designing human-in-the-loop review that catches failure in the latter. I'll cover what we measure, what we deliberately don't automate, and the failure modes we've hit. Attendees will leave with a practical framework for deciding which parts of a subjective pipeline can be evaluated algorithmically, which need human judgement by design, and how to defend either choice to a sceptical client.
Oliver Gooding is AI & Automation Lead at IFF Research, a UK social research agency working for central government clients including DWP, the Home Office and HMRC. He built and leads IFF's AI strategy and governance framework, and has spent the last two years shipping production LLM tooling into government research projects, from automated figure checking to LLM-assisted content analysis. He has a background in social research methodology and an MSci in Physics from Bristol.