Abraham Yeung

Abraham Yeung

Math & CS at Stanford. Reinforcement learning, post-training, and AI safety.

status: mid post-training  ยท  reward signal: curiosity  ยท  KL penalty: sleep

๐Ÿ“ˆ training history๐Ÿ“Š evals๐Ÿค– ask

๐Ÿค–policy

About me

I'm a junior at Stanford studying Mathematics and Computer Science. I work on reinforcement learning and post-training, and on what those methods do to a model's safety: whether its reasoning stays honest, and which training choices quietly trade oversight for capability. This fall I'm joining Redwood Research part-time as an AI safety research intern, after a summer on the Data Platform team at Databricks.

I grew up in Hong Kong and went to Eton College. I sing baritone with the Mendicants, Stanford's oldest a cappella group. I speak English, Cantonese, and Mandarin natively, plus some German and beginner Japanese.

๐Ÿฅ•reward model

What I optimize for

โˆ‡gradient updates

Research

Four papers are under review. Titles, summaries, and PDFs are on the papers page. Venues will be added once decisions are out.

๐Ÿ“rollouts

Projects

Research Frontier Mineragents ยท in progress

A resumable agent that mines recent ML papers, walks the citation graph to find proposed extensions nobody has attempted, and ranks the open problems. SQLite plus vector search, content-addressed caching, best-first frontier.

Prediction Markets Agentmarkets ยท top 5 of 100+

Agent that evaluates live Polymarket prices and flags mispriced contracts from incoming information and microstructure signals. Finalist at the NVIDIA, Vercel, and Brex hackathon.

Multi-Agent Causal Modelingagents ยท Bridgewater AI Hackathon

Selected as 1 of 24. A multi-agent system that maps causal graphs for policy-impact scenarios by decomposing macro questions into testable sub-claims.

Maestroai / music ยท TreeHacks runner-up

Multimodal music coach that analyzes instrument audio and video with NVIDIA vision and audio models. Devpost

gigabpesystems ยท Rust ยท PyPI

BPE tokenizer trainer in Rust with Python bindings. A 32k vocabulary on 12.9 GB of FineWeb in 38 s against 257 s for HuggingFace tokenizers, with every merge byte-identical and a fraction of the memory. GitHub

Transformer from scratchsystems ยท CS 336

A language model built end to end in PyTorch: tokenizer, training loop, FlashAttention as Triton kernels, distributed data parallel with optimizer sharding, profiling, scaling laws, and SFT plus GRPO post-training.

AQI Forecastingtime series ยท CS 229

LSTM, GNN, and CNN models over 545k+ observations for spatiotemporal air quality forecasting, with robustness testing under distribution shift.

๐Ÿ“ˆtraining history

Experience

๐Ÿ“Ševals

Honors

๐Ÿค–inference

Ask my post-trained self

A small open-weights model with a short bio in its context. Ask it about my work. It can be wrong, so trust the sections above over it.