Khang Nguyen
MS student in Mathematical Data Science at USC. Reinforcement learning theory with William Chang's group (BruinML, UCLA).
About
I'm a master's student in Mathematical Data Science at USC, finishing in December 2026, and I do reinforcement learning theory with William Chang's group at UCLA. Most of my work is on bandits and online learning: best-arm identification, adversarial and corrupted MDPs, continuous action spaces, and how fast entropy regularization approximates equilibria in continuous games. Alongside the theory I build things, most recently the LLM agent behind Coursistant's classroom assistant.
Research interests
Bandits and online learning, best-arm identification, adversarial and corrupted MDPs, continuous action spaces, equilibria in continuous games.
Publications and preprints
Published or accepted
- Tight rates of approximation of mixed Nash equilibria by entropy regularization in continuous games First author. Proved most of the results, wrote most of the paper and the rebuttal, ran the numerics. proceedingsPDF
- Gap-Independent Regret for Multi-Agent Combinatorial Semi-Bandits Tightened most of the proofs; wrote the experiment code. code
- A pathological property of nonlocal discrete operators Equal contribution (alphabetical). Independently proved the central algebraic identity. DOI
- Minimum-Energy Optimal Gaze Control for K-Eye Systems with Three-Axis Rotation First author.
Under review
- Hybrid Offline-Online Follower Manipulation in General-Sum Stackelberg Games
- Learning Adversarial Continuous MDPs with Bandit Feedback and Unknown Transitions
- Policy Optimization for Corrupted Markov Decision Processes First author.
- Best Arm Identification in Lipschitz Bandits: Uniform and Adaptive Strategies First author.
- Pure Exploration for Curriculum Bandits
- ZoomQ: Adaptive Action Discretization for Continuous-Action RL
Selected projects
- Coursistant classroom assistant. Sole AI engineer for an LLM agent covering the course lifecycle through chat: 25+ function-calling tools, citation-aware RAG, every state-changing action gated behind an out-of-band human confirmation, 300+ regression tests.
- Tiered document extraction. A PDF and document extraction pipeline that escalates from cheap parsers to OCR and vision models only when a cheaper tier fails, built for arbitrary input formats.
- Offline RL for pitch selection. Conservative Q-learning over 8,400+ Statcast pitches; traced an apparent 0.94 runs-per-pitch policy improvement to Q-function miscalibration rather than a real gain.
- CUDA convolutional network. A CNN written from scratch in CUDA C/C++ with forward and backward passes and a CPU-vs-GPU benchmark suite on MNIST.
Teaching
Teaching assistant, DSCI 560 Data Science Professional Practicum (algorithmic trading), USC, Spring 2026.