Judging by the Cover

Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer and Jeff M. Phillips.
arXiv 2026, under review at ICLR 2027.
PaperCodeData

01

The shortcut

Could you pass a truth quiz without reading a single question? On TruthfulQA, a benchmark the AI world uses to decide which models can be trusted, you can. The true answers simply look different: hedged a little more, phrased a little longer, and far more often opening with a flat “No”.

25% of true answers open with a negation, against 6% of false ones.

Every card is a real TruthfulQA pair with its cue words marked. The cells are the six surface features of each answer: a darker cell means a larger value than most TruthfulQA answers have. The third row is the part of the gap between the two answers that the audit can exploit: it counts only where the pair leans the same way the whole dataset does, and it is what Audit-Prune scores on its first pass. It rescores after every removal and later adds pairs back, so some high-scoring pairs end up kept. Shown: the paper’s four examples, then twenty short pairs drawn at random from neutral topics, ten removed and ten kept.

02

Six surface features

SURFACE6 is a checklist of six text-only cues: whether an answer opens with a negation, how many negations it holds, how much it hedges, its word count, its average token length and its type–token ratio. No embeddings, no pretrained model. A plain logistic regression on these six tells true answers from false ones on binary-choice TruthfulQA without ever seeing the question.

68.9% accuracy (AUC 0.715) where chance is 50%. That is higher than 16 of the 18 models on the llm-stats TruthfulQA leaderboard (September 2026).

The two numbers measure different things: the leaderboard scores a model’s answers to the questions, and this classifier never sees a question. A question-blind probe landing in the same range is itself the sign of leakage. Negation carries most of the signal. The probe on this page is the same classifier fitted on all 1,580 TruthfulQA answers, running in your browser. It reads style and knows nothing about truth, so treat its score as a measure of how an answer is dressed.

03

Not just TruthfulQA

The same six features, applied to 13 more benchmarks. Two hallucination benchmarks, HaluEval QA and MedHallu, leak more than TruthfulQA itself. MultiNLI, SNLI, MultiRC, SelfCheckGPT and FEVER carry a clear surface signal. BoolQ and PIQA sit at the edge of detectability.

AUC 0.973 on HaluEval QA and 0.821 on MedHallu, against 0.715 on TruthfulQA.

Bars start at chance (AUC 0.5) and carry 95% confidence intervals. Fifteen datasets in all, counting TruthfulQA and the cleaned TruthfulQA-476, which drops to near chance. HaluEval QA’s leak is driven by answer length. For paired datasets both answers of a pair are audited, so TruthfulQA’s 790 pairs are 1,580 texts.

04

Cleaning the leak

Audit-Prune greedily removes the pairs that most reinforce the asymmetry the audit classifier exploits. After every removal it refits the standardization, the classifier and the audit score, and it stops once the audit AUC is at or below a threshold θ. Then it adds back any removed pair whose return keeps the AUC at or below θ. It needs no knowledge of how the dataset was built, only whether leakage remains.

At θ = 0.53, 476 of 790 pairs are kept (60.3%). A fixed removal order stalls at AUC 0.583.

Move the threshold. The upper chart is how many pairs survive; the lower one is how well the surface audit still does on them, where 0.5 is chance. The fixed-prefix baseline removes pairs in one precomputed order and never gets below an audit AUC of 0.583. Refitting at every step is what lets Audit-Prune reach 0.528. Rank agreement is Spearman ρ between model rankings on the kept pairs and on the full benchmark, over 14 open-weight models.

05

TruthfulQA-476

The release: a cleaned 476-pair subset of binary-choice TruthfulQA that works as a drop-in replacement. The surface audit falls to the edge of statistical detectability while the ranking of real models is largely preserved. Classifiers trained on it are also harder to fool. On a held-out test set whose surface cues are flipped, accuracy rises significantly for five of nine classifier families, and it stays unchanged within noise on a natural test set.

Model ranking preserved: Spearman ρ = 0.915 (95% CI [0.81, 0.99]), Kendall τ = 0.827.

SurfaceFlipped-131 holds misconception questions whose false answer carries the true-looking cues (a negation lead, hedging, extra length) and whose true answer is a bare assertion. Natural-131 holds open-ended factual questions across twelve neutral domains, written without regard to surface form. Each was generated by a frontier LLM, checked by an independent LLM judge, then verified by hand. The five families listed improve with McNemar p < 0.05.

What cleaning changes
Measure TruthfulQA TruthfulQA-476
Pairs 790 476
Surface-audit AUC after cleaning: 95% CI [0.503, 0.557], p = 0.048 0.715 0.528
Surface-audit accuracy 0.689 0.522
Accuracy on SurfaceFlipped-131, by training set
Classifier TruthfulQA TruthfulQA-476
Llama-3.2-3B 0.267 0.450
SmolLM2-1.7B 0.427 0.634
Phi-3.5-mini 0.435 0.580
BGE-large 0.344 0.458
ModernBERT-base 0.618 0.733

06

Paper, data and code