Publications
* equal contribution · full list on Google Scholar
2026
-
Entropy-Preserving Reinforcement Learning. In ICLR 2026.Policy gradient methods quietly collapse policy entropy during training; we introduce mechanisms that keep policies diverse, performant, and retrainable on new tasks.
2025
-
Reinforcement Learning for Long-Horizon Interactive LLM Agents. arXiv preprint, 2025.RL for long-horizon, multi-turn, tool-using LLM agents. A 32B Qwen2.5 LoRA fine-tune reached 71% on AppWorld, 9 points above OpenAI o1, after training on just 24 scenarios. -
Robust Autonomy Emerges from Self-Play. In ICML 2025.Gigaflow, a batched driving simulator that trains on 42 years of driving experience per hour on a single 8-GPU node: robust and naturalistic driving emerges from 1.6 billion km of self-play, without ever seeing human data.
2024
2023
-
QuadSwarm: A Modular Multi-Quadrotor Simulator for Deep Reinforcement Learning with Direct Thrust Control. In ICRA 2023 Workshop on the Role of Robotics Simulators for UAVs.A fast modular multi-quadrotor simulator with direct thrust control and domain randomization for sim-to-real reinforcement learning.
2021
2020
-
Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning. In ICML 2020.Reinforcement learning framework with the highest single-machine training throughput at the time of publication, about 10x faster than traditional synchronous RL implementations. SOTA results in challenging VizDoom and DMLab environments.

