Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng et al. · arXiv · 2026
Abstract
Benchmarks assembled from claims published after a model's training cutoff are assumed to be free of contamination. This paper tests that assumption and finds it only partly true: between 17 and 29 percent of post-cutoff claims are still effectively contaminated, often because they can be settled by combining public knowledge that predates the cutoff. The resulting inflation on fact-checking scores reaches eleven Macro-F1 points.
Why it matters
A caution worth internalising whenever you read a benchmark result. 'Published after the cutoff' is a weaker guarantee than it sounds, because reasoning over old knowledge can answer new questions.
https://arxiv.org/abs/2607.23514