There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
Key takeaway
Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods.
Why it's on PDOOM
PDOOM selected this story for its capabilities signals: AGI, Benchmark.
AGIBenchmarkFrom ArXiv cs.AI
Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the e
By V. S. Raghu Parupudi