cv
Basics
| Name | Zeyu (Steven) Zhang |
| Label | PhD Student |
| zeyuzhang2028@u.northwestern.edu | |
| Url | https://ZeyuZhang1901.github.io |
| Summary | PhD student at Northwestern University working on trustworthy and reliable LLMs: temporal knowledge leakage in LLM backtesting, black-box interpretability, and RL-based post-training. |
Work
-
2025.09 - Present Graduate Researcher - Trustworthy LLM Backtesting (TEMPO, TimeSPEC & Leakage-Adjusted Evaluation)
Northwestern University
Detecting, attributing, quantifying, and eliminating temporal knowledge leakage in LLM backtesting. Three first-author papers: Findings of EMNLP 2026, plus two under review or in submission (NeurIPS 2026, TMLR).
- Formalized temporal knowledge leakage: models exploit post-cutoff pre-training knowledge on historical prediction tasks, inflating reported accuracy and invalidating evaluation
- Proved the leakage inflation is non-identifiable from passive backtest scores (recency alone reproduces the leakage signature: four flagship models fail the naive pre/post check on questions they cannot have memorized) and derived a concentration law: inflation = honest uncertainty x extraction factor, so partial memorization is disproportionately rewarded
- Showed one external reference restores measurement: a known cutoff (regression discontinuity) or a matched clean control (difference-in-differences) yields a leakage-adjusted score with confidence intervals; validated by planting leakage in twin models (dose recovery, null on the clean twin) and deployed on frontier models, detecting one cutoff-localized signature and clearing five models whose advantages were recency alone
- Introduced Shapley-DCLR, a claim-level Shapley-weighted metric quantifying what fraction of decision-critical reasoning is contaminated
- Designed TimeSPEC, an inference-time architecture interleaving temporally filtered retrieval with claim-level supervision, requiring no retraining
- Developed TEMPO, GRPO-based post-training with a two-mode reward; proved monotonic leakage decrease and convergence to the leak-free optimum
- Reduced leakage from 2-13% to 0.6-3.7% across three tasks and two models while improving task performance by 6-13% where valid signals exist
-
2024.06 - 2025.01 Graduate Researcher - Factual Memorization in LLMs
Northwestern University
Analysis of LLM memorization behaviors under SFT and DPO. Qualifying examination paper.
- Analyzed LLM memorization under SFT and DPO fine-tuning, reproducing and extending Stanford's FineTuneBench
- Proposed a formal distinction between passive memorization (exposure) and positive memorization (direct QA supervision)
- Showed injected future-dated facts do not generalize beyond training, and that a temporal-consistency system prompt mitigates the resulting overfitting
-
2024.01 - 2025.01 Graduate Researcher - Black-Box LLM Interpretability (LAMP)
Northwestern University
Co-developed LAMP for interpreting black-box LLMs. Accepted as Spotlight at AISTATS 2026.
- Co-developed LAMP, which treats an LLM's self-reported explanations as a coordinate system and fits locally linear surrogate decision surfaces
- Designed perturbation-based probing that extracts decision surfaces without gradients, logits, or internal activations, enabling audits of proprietary LLMs
- Ran experiments across sentiment analysis, controversial-topic detection, and safety-prompt auditing; surfaces align with human and expert judgments
-
2022.05 - 2023.12 Research Collaborator - Unified Off-Policy Learning to Rank (CUOLR)
Princeton University (remote)
Mentored by Prof. Mengdi Wang and Prof. Huazheng Wang. Published at NeurIPS 2023.
- Unified ranking under general stochastic click models as a Markov Decision Process, enabling principled offline RL for off-policy learning to rank
- Proposed CUOLR, a click-model-agnostic algorithm requiring no explicit debiasing or prior knowledge of the click process
- Achieved state-of-the-art performance on large-scale benchmarks with consistent robustness across heterogeneous click models
Teaching
-
2026.09 - 2026.12 Evanston, IL
Teaching Assistant - Applied Multivariate Analysis (STAT 348)
Northwestern University
Teaching assistant for Applied Multivariate Analysis with Prof. Thomas Severini (second appointment).
- Hold in-class discussion sessions to answer student questions
- Grade all homework assignments
-
2026.04 - 2026.06 Evanston, IL
Teaching Assistant - Data Science 3 with Python (STAT 303-3)
Northwestern University
Teaching assistant for advanced data science course with Prof. Emre Besler.
- Evaluated quizzes, homework, and exams throughout the term
-
2026.01 - 2026.03 Evanston, IL
Teaching Assistant - Data Science 2 with Python (STAT 303-2)
Northwestern University
Teaching assistant for Data Science 2 with Prof. Emre Besler.
- Graded all homework assignments and in-person exams
-
2025.09 - 2025.12 Evanston, IL
Teaching Assistant - Applied Multivariate Analysis (STAT 348)
Northwestern University
Teaching assistant for Applied Multivariate Analysis with Prof. Thomas Severini.
- Held weekly TA discussion sessions
- Graded all homework assignments
-
2025.04 - 2025.06 Evanston, IL
Teaching Assistant - Data Science 3 with Python (STAT 303-3)
Northwestern University
Teaching assistant for advanced data science course with Prof. Emre Besler.
- Supported instruction on non-linear models and tree-based methods
- Graded homework and exams
-
2025.01 - 2025.03 Evanston, IL
Teaching Assistant - Data Science 2 with Python (STAT 303-2)
Northwestern University
Teaching assistant for Data Science 2 with Prof. Emre Besler.
- Graded homework assignments and exams
-
2024.09 - 2024.12 Evanston, IL
Teaching Assistant - Introduction to Probability and Statistics (STAT 210)
Northwestern University
Teaching assistant for introductory statistics course with Prof. Maxim Sinitsyn.
- Held discussion sessions twice a week
- Graded midterms and final exam
Education
-
2023.09 - 2028.06 Evanston, IL, USA
Doctor of Philosophy
Northwestern University
Statistics and Data Science
- GPA: 3.95/4.00
- Advisor: Prof. Bradly C. Stadie
- Expected graduation: June 2028
-
2020.09 - 2023.06 Hefei, Anhui, P.R.China
Certificate (Minor)
University of Science and Technology of China (USTC)
Artificial Intelligence
- Talent Program in Artificial Intelligence
- Outstanding Undergraduate Honorary Rank (top 5%)
-
2019.09 - 2023.06 Hefei, Anhui, P.R.China
Bachelor of Engineering
University of Science and Technology of China (USTC)
Electronic Information Engineering
- GPA: 3.93/4.30, Rank: 5/213 in School of Information Science and Technology
- China National Scholarship, Ministry of Education of the PRC (top 1%)
- Wang Xiaomo Talent Program in Cyber Science and Technology
Awards
- 2026.01.01
Spotlight Presentation, AISTATS 2026
International Conference on Artificial Intelligence and Statistics
LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models.
- 2020.10.01
China National Scholarship
Ministry of Education of the P.R.China
Top 1% of students. Academic year 2020-2021.
- 2020.09.01
Outstanding Student Scholarship (Class A)
University of Science and Technology of China
Top 1% of students.
- 2021.11.01
The National Undergraduate Electronic Design Contest
Ministry of Education of China
2nd Prize Nationally, 1st Prize in Anhui Province.
- 2020.10.01
Scholarship for Talent Program in Basic Disciplines (Class A)
University of Science and Technology of China
Top 3% of students.
Publications
-
2026.08.04 Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores
Preprint, in submission to TMLR (arXiv:2608.02985)
Zeyu Zhang, Bradly C. Stadie
-
2026.05.01 TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
Preprint, under review at NeurIPS 2026 (arXiv:2605.18843)
Zeyu Zhang, Bradly C. Stadie
-
2026.02.01 All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
Findings of EMNLP 2026, to appear (arXiv:2602.17234)
Zeyu Zhang, Ryan Chen, Bradly C. Stadie
-
2026.01.15 LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
International Conference on Artificial Intelligence and Statistics (AISTATS), 2026. Spotlight
Ryan Chen, Youngmin Ko, Zeyu Zhang, Catherine Cho, Sunny Chung, Mauro Giuffré, Dennis L. Shung, Bradly C. Stadie
-
2023.12.01 Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective
Advances in Neural Information Processing Systems (NeurIPS), 2023
Zeyu Zhang, Yi Su, Hui Yuan, Yiran Wu, Rishab Balasubramanian, Qingyun Wu, Huazheng Wang, Mengdi Wang
Skills
| Programming | |
| Python | |
| PyTorch | |
| Hugging Face Transformers | |
| R | |
| C/C++ | |
| Bash | |
| Git | |
| LaTeX |
| ML / LLM | |
| RL post-training (GRPO, DPO, RLHF) | |
| LLM evaluation & benchmarking | |
| Retrieval-augmented pipelines | |
| Offline reinforcement learning | |
| API-based LLM systems (OpenAI platform) |
Languages
| Chinese | |
| Native speaker |
| English | |
| Fluent |