cv
Basics
| Name | Zeyu (Steven) Zhang |
| Label | PhD Student |
| zeyuzhang2028@u.northwestern.edu | |
| Url | https://ZeyuZhang1901.github.io |
| Summary | PhD student at Northwestern University working on trustworthy and reliable LLMs: temporal knowledge leakage in LLM backtesting, black-box interpretability, and RL-based post-training. |
Work
-
2025.09 - Present Graduate Researcher - Trustworthy LLM Backtesting (TEMPO, TimeSPEC & Leakage-Adjusted Evaluation)
Northwestern University
Detecting, attributing, quantifying, and eliminating temporal knowledge leakage in LLM backtesting. Three first-author papers under review or in submission (NeurIPS 2026, EMNLP 2026, TMLR).
- Formalized temporal knowledge leakage: models exploit post-cutoff pre-training knowledge on historical prediction tasks, inflating reported accuracy and invalidating evaluation
- Proved the leakage inflation is non-identifiable from passive backtest scores and that one external reference (clean-control difference-in-differences or cutoff regression discontinuity) restores identification, yielding a leakage-adjusted score with confidence intervals
- Introduced Shapley-DCLR, a claim-level Shapley-weighted metric quantifying what fraction of decision-critical reasoning is contaminated
- Designed TimeSPEC, an inference-time architecture interleaving temporally filtered retrieval with claim-level supervision, requiring no retraining
- Developed TEMPO, GRPO-based post-training with a two-mode reward; proved monotonic leakage decrease and convergence to the leak-free optimum
- Reduced leakage from 2-13% to 0.6-3.7% across three tasks and two models while improving task performance by 6-13% where valid signals exist
-
2024.06 - 2025.01 Graduate Researcher - Factual Memorization in LLMs
Northwestern University
Analysis of LLM memorization behaviors under SFT and DPO. Qualifying examination paper.
- Analyzed LLM memorization under SFT and DPO fine-tuning, reproducing and extending Stanford's FineTuneBench
- Proposed a formal distinction between passive memorization (exposure) and positive memorization (direct QA supervision)
- Showed injected future-dated facts do not generalize beyond training, and that a temporal-consistency system prompt mitigates the resulting overfitting
-
2024.01 - 2025.01 Graduate Researcher - Black-Box LLM Interpretability (LAMP)
Northwestern University
Co-developed LAMP for interpreting black-box LLMs. Accepted as Spotlight at AISTATS 2026.
- Co-developed LAMP, which treats an LLM's self-reported explanations as a coordinate system and fits locally linear surrogate decision surfaces
- Designed perturbation-based probing that extracts decision surfaces without gradients, logits, or internal activations, enabling audits of proprietary LLMs
- Ran experiments across sentiment analysis, controversial-topic detection, and safety-prompt auditing; surfaces align with human and expert judgments
-
2022.05 - 2023.12 Research Collaborator - Unified Off-Policy Learning to Rank (CUOLR)
Princeton University (remote)
Mentored by Prof. Mengdi Wang and Prof. Huazheng Wang. Published at NeurIPS 2023.
- Unified ranking under general stochastic click models as a Markov Decision Process, enabling principled offline RL for off-policy learning to rank
- Proposed CUOLR, a click-model-agnostic algorithm requiring no explicit debiasing or prior knowledge of the click process
- Achieved state-of-the-art performance on large-scale benchmarks with consistent robustness across heterogeneous click models
Education
-
2023.09 - 2028.06 Evanston, IL, USA
Doctor of Philosophy
Northwestern University
Statistics and Data Science
- GPA: 3.95/4.00
- Advisor: Prof. Bradly C. Stadie
- Expected graduation: June 2028
-
2020.09 - 2023.06 Hefei, Anhui, P.R.China
Certificate (Minor)
University of Science and Technology of China (USTC)
Artificial Intelligence
- Talent Program in Artificial Intelligence
- Outstanding Undergraduate Honorary Rank (top 5%)
-
2019.09 - 2023.06 Hefei, Anhui, P.R.China
Bachelor of Engineering
University of Science and Technology of China (USTC)
Electronic Information Engineering
- GPA: 3.93/4.30, Rank: 5/213 in School of Information Science and Technology
- China National Scholarship, Ministry of Education of the PRC (top 1%)
- Wang Xiaomo Talent Program in Cyber Science and Technology
Awards
- 2026.01.01
Spotlight Presentation, AISTATS 2026
International Conference on Artificial Intelligence and Statistics
LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models.
- 2020.10.01
China National Scholarship
Ministry of Education of the P.R.China
Top 1% of students. Academic year 2020-2021.
- 2020.09.01
Outstanding Student Scholarship (Class A)
University of Science and Technology of China
Top 1% of students.
- 2021.11.01
The National Undergraduate Electronic Design Contest
Ministry of Education of China
2nd Prize Nationally, 1st Prize in Anhui Province.
- 2020.10.01
Scholarship for Talent Program in Basic Disciplines (Class A)
University of Science and Technology of China
Top 3% of students.
Publications
-
2026.07.01 Temporal Leakage in LLM Backtesting: From Statistical Modeling to Leakage-Adjusted Evaluation
In submission to TMLR
Zeyu Zhang, Bradly C. Stadie
-
2026.05.01 TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
Preprint, under review at NeurIPS 2026 (arXiv:2605.18843)
Zeyu Zhang, Bradly C. Stadie
-
2026.02.01 All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
Preprint, under review at EMNLP 2026 (arXiv:2602.17234)
Zeyu Zhang, Ryan Chen, Bradly C. Stadie
-
2026.01.15 LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
International Conference on Artificial Intelligence and Statistics (AISTATS), 2026. Spotlight
Ryan Chen, Youngmin Ko, Zeyu Zhang, Catherine Cho, Sunny Chung, Mauro Giuffré, Dennis L. Shung, Bradly C. Stadie
-
2023.12.01 Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective
Advances in Neural Information Processing Systems (NeurIPS), 2023
Zeyu Zhang, Yi Su, Hui Yuan, Yiran Wu, Rishab Balasubramanian, Qingyun Wu, Huazheng Wang, Mengdi Wang
Skills
| Programming | |
| Python | |
| PyTorch | |
| Hugging Face Transformers | |
| R | |
| C/C++ | |
| Bash | |
| Git | |
| LaTeX |
| ML / LLM | |
| RL post-training (GRPO, DPO, RLHF) | |
| LLM evaluation & benchmarking | |
| Retrieval-augmented pipelines | |
| Offline reinforcement learning | |
| API-based LLM systems (OpenAI platform) |
Languages
| Chinese | |
| Native speaker |
| English | |
| Fluent |