status: mid post-training ยท reward signal: curiosity ยท KL penalty: sleep
๐ training history๐ evals๐ค ask
๐คpolicy
About me
I'm a junior at
Stanford studying Mathematics and Computer Science. I work on reinforcement learning and post-training, and on what those methods do to a model's safety: whether its reasoning stays honest, and which training choices quietly trade oversight for capability. This fall I'm joining
Redwood Research part-time as an AI safety research intern, after a summer on the Data Platform team at
Databricks.
I grew up in Hong Kong and went to
Eton College. I sing baritone with the
Mendicants, Stanford's oldest a cappella group. I speak English, Cantonese, and Mandarin natively, plus some German and beginner Japanese.
๐ฅreward model
What I optimize for
- Models that stay overseeable as they get stronger. Chain-of-thought monitoring only works if the reasoning we read is the reasoning that drove the answer. I want to know which training choices preserve that, and which erode it while the text still looks fine.
- Post-training that doesn't trade safety for capability by accident. Reward design, distillation, and RL loops each have side effects on honesty and controllability. Those effects should be measured, not assumed.
- Interpretability artifacts wired into training carefully. A probe that is safe to read can be dangerous to optimize against. The access a training loop gives an internal signal should be a deliberate decision.
- Measurements that survive scrutiny. Pre-registered designs, placebo controls, bootstrap intervals, and numbers that regenerate from committed artifacts. A safety metric with a hidden confound is worse than no metric.
โgradient updates
Research
Four papers are under review. Titles, summaries, and PDFs are on the papers page. Venues will be added once decisions are out.
๐rollouts
Projects
A resumable agent that mines recent ML papers, walks the citation graph to find proposed extensions nobody has attempted, and ranks the open problems. SQLite plus vector search, content-addressed caching, best-first frontier.
Agent that evaluates live Polymarket prices and flags mispriced contracts from incoming information and microstructure signals. Finalist at the NVIDIA, Vercel, and Brex hackathon.
agents ยท Bridgewater AI HackathonSelected as 1 of 24. A multi-agent system that maps causal graphs for policy-impact scenarios by decomposing macro questions into testable sub-claims.
ai / music ยท TreeHacks runner-upMultimodal music coach that analyzes instrument audio and video with NVIDIA vision and audio models. Devpost
BPE tokenizer trainer in Rust with Python bindings. A 32k vocabulary on 12.9 GB of FineWeb in 38 s against 257 s for HuggingFace tokenizers, with every merge byte-identical and a fraction of the memory. GitHub
A language model built end to end in PyTorch: tokenizer, training loop, FlashAttention as Triton kernels, distributed data parallel with optimizer sharding, profiling, scaling laws, and SFT plus GRPO post-training.
LSTM, GNN, and CNN models over 545k+ observations for spatiotemporal air quality forecasting, with robustness testing under distribution shift.
๐training history
Experience
- Sep 2026 โ Jan 2027
- Jun โ Sep 2026
- Sep 2024 โ Jun 2026
Director of Hackspace, BASESStanford's largest student-run entrepreneurship programs. Sponsor relations and coordination for hackathons serving 200+ participants. - Sep 2025 โ now
Course Grader, Applied Matrix Theory, Stanford MathematicsEvaluate proof-based linear algebra for 200+ students. - Jun โ Sep 2025
Undergraduate Researcher, Chiu Lab, StanfordBioengineering REU. Built CryoViT, CNN and vision transformer models that segment cell structures in noisy 3D cryo-EM data, applied to Alzheimer's research. - Jan โ Jun 2025
Teaching Assistant, Programming Abstractions, Stanford CSLed weekly C++ sections on recursion, complexity, and data structures.
๐evals
Honors
YC FellowSummer 2026 ยท mentored by Harshita Arora
TreeHacks Summer Fellow2026
MATS 11.0accepted
Rabi ScholarColumbia University ยท top 10 scientific admits nationally
Gladstone Memorial Prize, King's ScholarEton College ยท valedictorian prize
PrefectEton College
British Mathematical Olympiad Round 2top 50 nationally, twice
UK Chemistry Olympiadgold medal three times ยท IChO team reserve
World Science Scholars1 of 48 selected globally
๐คinference
Ask my post-trained self
A small open-weights model with a short bio in its context. Ask it about my work. It can be wrong, so trust the sections above over it.