Measure & mitigate at inference TimeSPEC
All Leaks Count, Some Count More
Backtesting assumes models use only pre-cutoff evidence, but pretrained LLMs routinely leak later knowledge into their rationales. We introduce Shapley-DCLR to measure how much decision-critical reasoning is contaminated at the claim level, and TimeSPEC, an inference-time supervisor that regenerates predictions from temporally filtered evidence without retraining.
Preprint · Under review at EMNLP 2026