Tracking AI existential risk. Source-backed context. Original reporting always linked.
MONITORING CORE AI-RISK FEEDS · UPDATED HOURLY
Research

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

Zac Boring August 25, 2026 1 min read
Read original source →

Key takeaway

Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods.

Why it's on PDOOM

PDOOM selected this story for its capabilities signals: AGI, Benchmark.

AGIBenchmarkFrom ArXiv cs.AI

Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the e

By V. S. Raghu Parupudi

Read the full article at ArXiv cs.AI →