Zeyu (Steven) Zhang

PhD Student, Statistics and Data Science, Northwestern University

prof_pic.jpg

🏛️ Statistics & Data Science

🎓 Northwestern University

📧 zeyuzhang2028@u.northwestern.edu

I am a third-year Ph.D. student in the Department of Statistics and Data Science at Northwestern University. My research focuses on knowledge leakage in large language models, with particular emphasis on settings that involve complex reasoning and chain-of-thought processes. I study how and when sensitive or unintended information may be exposed during multi-step inference, and I am interested in principled methods for detecting, characterizing, and mitigating such leakage. In parallel, I work on fine-tuning language models using reinforcement learning–based approaches, including RLHF, Direct Preference Optimization (DPO), and related algorithms, with the goal of improving alignment, robustness, and controllability.

Outside of research, I am deeply committed to an active lifestyle. I follow a structured workout program and train almost every day, viewing physical discipline as a natural complement to intellectual rigor. I enjoy a wide range of sports, with basketball being my favorite, and I am particularly drawn to long-distance running for both its physical demands and its meditative rhythm. These activities keep me grounded, energized, and continuously motivated—both on and off the track.

news

May 25, 2026 📄 New preprint: All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting is now on arXiv. Code is available here.
May 13, 2026 📄 New preprint: TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting is now on arXiv. Code is available here.
Jan 22, 2026 🎉 Thrilled to announce that our paper LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models has been accepted as a Spotlight at AISTATS 2026! Grateful to all co-authors for the amazing collaboration. 🚀

latest posts

selected publications

  1. tempo2026_preview.png
    TEMPO
    TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
    Zeyu Zhang and Bradly C. Stadie
    RL fine-tuning that enforces temporal discipline before optimizing task performance.
    Preprint, Under review at NeurIPS 2026
  2. timespec2026_preview.png
    TimeSPEC
    All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
    Zeyu Zhang, Ryan Chen, and Bradly C. Stadie
    Claim-level leakage metric (Shapley-DCLR) plus an inference-time architecture that keeps predictions pre-cutoff.
    Preprint, Under review at EMNLP 2026
  3. lamp2025_preview.png
    LAMP
    LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
    Ryan Chen, Youngmin Ko, Zeyu Zhang, and 5 more authors
    Local linear probes that reveal how LLM decisions shift under controlled input changes.
    Spotlight paper in International Conference on Artificial Intelligence and Statistics, 2026
  4. CUOLR2023_preview.png
    Off-Policy LTR
    Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective
    Zeyu Zhang, Yi Su, Hui Yuan, and 5 more authors
    A reinforcement-learning view that unifies off-policy learning-to-rank methods.
    In Advances in Neural Information Processing Systems, 2023