Publications

First Author Publications

  1. tempo2026_preview.png
    TEMPO
    TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
    Zeyu Zhang and Bradly C. Stadie
    Temporal leakage lets LLMs recall post-cutoff outcomes during backtesting, so reported accuracy is often invalid. TEMPO trains temporal discipline with a two-mode RL objective that first drives leaked claims to zero, then optimizes prediction quality once reasoning stays pre-cutoff, with theory and experiments showing large leakage reductions without sacrificing performance when valid signals exist.
    Preprint, Under review at NeurIPS 2026
  2. timespec2026_preview.png
    TimeSPEC
    All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
    Zeyu Zhang, Ryan Chen, and Bradly C. Stadie
    Backtesting assumes models use only pre-cutoff evidence, but pretrained LLMs routinely leak later knowledge into their rationales. We introduce Shapley-DCLR to measure how much decision-critical reasoning is contaminated at the claim level, and TimeSPEC, an inference-time supervisor that regenerates predictions from temporally filtered evidence without retraining.
    Preprint, Under review at EMNLP 2026
  3. CUOLR2023_preview.png
    Off-Policy LTR
    Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective
    Zeyu Zhang, Yi Su, Hui Yuan, and 5 more authors
    Most off-policy LTR methods assume a specific click model and need custom debiasing for each setting. We cast ranking under general stochastic click models as an MDP and propose CUOLR, a click-model-agnostic offline RL approach that learns from logged clicks without prior knowledge of the click process, outperforming prior methods across large-scale benchmarks.
    In Advances in Neural Information Processing Systems, 2023

Co-authored Publications

  1. lamp2025_preview.png
    LAMP
    LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
    Ryan Chen, Youngmin Ko, Zeyu Zhang, and 5 more authors
    Fluent LLM explanations need not track what actually drives the prediction. LAMP treats the model’s own stated factors as coordinates and fits a local linear decision surface, without gradients or internal access, to test whether predictions move consistently when those factors change, enabling lightweight auditing of proprietary models (AISTATS 2026 Spotlight).
    Spotlight paper in International Conference on Artificial Intelligence and Statistics, 2026