Louis Yiven Zhu

Oxford Internet InstituteAI Evaluation

00Who am I

I’m Louis, an MSc student at the Oxford Internet Institute, working on the science of AI evaluation. My research asks what a benchmark score is really evidence of. I am increasingly interested in when those scores predict performance beyond the test, especially when AI systems use tools or work with people. I also study how measured capability enters economic measurement, labour-market evidence and governance.

ice hockeyalpine ski racingpiano

01Research

AI evaluations produce scores. The harder question is what those scores allow us to claim. My work examines their reliability, what they measure, and whether they predict performance in a different setting. I draw on psychometricsThe statistics of human testing, which asks whether a test is reliable, comparable, and measuring what it claims., statistical modelling and economics, with a growing interest in evaluations of interactive and tool-using systems.

  1. Q1Can the score be trusted?
  2. Q2What does the score measure?
  3. Q3What does it predict outside the evaluation?
  4. Q4When should a decision rely on it?

NextI am developing a study of whether pre-deployment evaluations predict independently assessed task performance outside the original test conditions. Professional work is one possible setting. I am particularly interested in what changes when tools, interaction and human judgement become part of the system being evaluated.

Selected work

under reviewQ2 · measure

One Capability or Many?

Do economically framed benchmarks measure a distinct capability? No distinct factor emerges under a pre-specified rule, yet a multi-factor model predicts them better out of sample.

421 model configurationsHash-pinned frontier model configurations in the frozen twelve-benchmark battery, the full sample. The paper also reports a 103-model economically dense subset. ΔMSE 0.037 incremental predictionIn the leave-one-benchmark-out test, a multi-factor model improves prediction of held-out economic scores over a single general index by a pooled 0.037 in MSE, bootstrap 95% interval [0.019, 0.055]. Modest but reliable.

pre-specified hypotheses · analysis plan deposited retrospectively · under review, NeurIPS 2026 TAE workshop
preprint · analysis plan · code & data

under reviewQ1 · trust

Three Ways Classical Test Theory Misleads for LLM Judges

Examines when reliability statistics answer the wrong question in LLM-judge pipelines, with an empirical item bank and simulation studies.

3 statistics misindexedKR-20, the dependability index and Livingston–Lewis classification accuracy, each answering a question the judge pipeline does not ask.

under review, NeurIPS 2026 workshop on reliable evaluation (JUDGe)
code & data · preprint soon

under reviewQ4 · decide

The Price of Intelligence

Constructs a quality-adjusted price index for AI inference from public data, examining how the way capability is measured changes the account of falling prices.

3,208 models21,024 posted-price observations across 3,208 models and 86 providers over 31 months, February 2024 to August 2026. The capability table scores 782 models; the crosswalk links 500 of them to the price panel. 87% hidden declineMatched-model methods measure inference prices falling at 0.10 log points a year; the quality-adjusted index falls at 0.73. The difference arrives through new models rather than price cuts on old ones. 0.998 ranks intactExcluding contamination-flagged benchmarks leaves model rankings intact at 0.998 yet moves the index by 0.49 log points a year. Agreement on rankings does not license claims about levels.

pre-registered validity audit · under review, NeurIPS 2026 EconML workshop
preprint · dataset · Hugging Face · pre-registration

workshop paper

From Advisor to Voting Teammate

Studies how an AI agent’s institutional role and access to information affect group decisions, in an agent-based simulation.

1.1M simulation runsRuns of the agent-based model across authority structures and information conditions. A simulation, not a human-participant experiment.

Tian, Zhang, Zhu and colleagues · published · Workshop on Human-Agent Collaboration, CHI 2026

in preparation

The Science of Evaluations

A collaborative account of open problems in AI evaluation. I contribute to the sections on validity and the evidence needed to support evaluation claims.

EvalEval Coalition · Hugging Face, Edinburgh, EleutherAI · core contributor

Further work

Other research outside the evaluation programme

Earlier and dormant work · adversarial CAPTCHAs that humans pass and AI agents fail, with UCL Computer Science and Holistic AI · a study of cognitive load, XAI and trust in human–AI teams with UCL Computer Science, study design complete

02Experience

University of Oxford UCL LSE ETH Zurich MIT Holistic AI Hugging Face
2026–

Research Contributor · Relit, ETH Zurich

Benchmark construction and perturbation testing for legal NLP, with Prof Elliott Ash

2026

Departmental Associate & Summer Research Intern · UCL Science & Technology Studies

Designed a new module on AI, digital labour and the future of work as a funded UCL MAPS commission, building its evidence base and syllabus from the review up. Supervised by Dr Joanna Octavia, May to September 2026

2026–

Supervised research project · LSE Department of Statistics

AI benchmark measurement, a latent-variable analysis of a twelve-benchmark battery with held-out prediction, extended from assessed coursework into the One Capability manuscript. Supervised by Dr Marcos E. Barreto

2026–

Core Contributor · EvalEval Coalition

Contributing the treatment of validity and evidentiary standards to the Science of Evaluations paper, with Hugging Face, Edinburgh and EleutherAI, and a source-to-claim verification workflow for the draft

2026

Teaching Assistant & Course Developer · UCL STS

Responsible Innovation in Practice and Governance of Emerging Technologies

2025–26

Student AI Researcher · Holistic AI

Adversarial robustness of LLM and vision–language agents

2025–26

Student Researcher · UCL Computer Science

Agent-based modelling of AI authority in group decisions, with Prof Maarten Speekenbrink; the From Advisor to Voting Teammate workshop paper at CHI 2026

2025

Research Contributor · Institute of Economic Affairs

UK graduate premium, contributing to Julian Jessop’s analysis of administrative earnings data

2025

Research Contributor · UCL Institute for Global Prosperity

AI & Youth, with Honorary Professor Noreena Hertz

03Talks & community

Talks & presentations

  • Workshop on Human-Agent Collaboration, CHI 2026 — paper presentation
  • UCL Centre for Responsible Innovation — invited to present the MMLU validity work
  • Explore Econ 2025, UCL — mandatory AI risk disclosure as a regulatory mechanism

Roles & service

  • Invited Reviewer — NeurIPS 2026 workshops, Trust-AI-Eval (TAE) and EconML
  • AI for Good, University of Oxford — Director of Engagement
  • Oxford Artificial Intelligence Society — Sponsorship Lead
  • Thinking About Thinking — Fellow, 2026–27
  • UCL Investment Society — Chairman · UCL Political Science & Economy Society — President, to 2026

04Recognition

05Ask me anything

Ask Louis
ready

06Contact

For research, collaboration, or data and code requests.