Hi there, I'm Aleksei Petrenko

prof_pic.jpg

Making computers do things at which, at the moment, people are better. (Rich and Knight, 1991)

I am a research scientist at Apple, currently working with Vladlen Koltun on RLVR post-training for interactive LLM agents.

I received my PhD in Computer Science in 2022 from the University of Southern California. I was part of the Robotic Embedded Systems Lab, advised by prof. Gaurav Sukhatme.

During my PhD I worked at NVIDIA on high-throughput simulation and RL for dexterous robots, and at Intel on massively parallel 3D rendering and high-throughput reinforcement learning. Before going to academia I spent 8 years in industry, working on software R&D, machine learning, algorithms, 3D graphics, computer vision, and virtual reality.

Research interests

I study computationally efficient methods for training in simulation using reinforcement learning, as well as problems of sim-to-real transfer. Recently I’ve been working on:

  • Highly optimized systems for deep reinforcement learning, such as RL algorithms and simulators.
  • Advanced training scenarios: population-based training and self-play.
  • RL post-training with verifiable tasks for LLM agents.
  • Reinforcement learning in robotics: dexterous manipulation, autonomous driving, drones.

In the past I also worked on quality-diversity methods, exploration in RL, memory in embodied agents, and stochastic future prediction.

Selected publications All publications →

  1. Entropy-Preserving Reinforcement Learning
    Entropy-Preserving Reinforcement Learning. A Petrenko*, B Lipkin*, K Chen, E Wijmans, M Cusumano-Towner, R Giryes, P Krähenbühl. In ICLR 2026.
    Policy gradient methods quietly collapse policy entropy during training; we introduce mechanisms that keep policies diverse, performant, and retrainable on new tasks.
  2. Reinforcement Learning for Long-Horizon Interactive LLM Agents
    Reinforcement Learning for Long-Horizon Interactive LLM Agents. K Chen*, M Cusumano-Towner*, B Huval*, A Petrenko*, J Hamburger, V Koltun, P Krähenbühl. arXiv preprint, 2025.
    RL for long-horizon, multi-turn, tool-using LLM agents. A 32B Qwen2.5 LoRA fine-tune reached 71% on AppWorld, 9 points above OpenAI o1, after training on just 24 scenarios.
  3. Robust Autonomy Emerges from Self-Play
    Robust Autonomy Emerges from Self-Play. M Cusumano-Towner*, D Hafner*, A Hertzberg*, B Huval*, A Petrenko*, E Vinitsky*, E Wijmans*, T Killian, S Bowers, O Sener, P Krähenbühl, V Koltun. In ICML 2025.
    Gigaflow, a batched driving simulator that trains on 42 years of driving experience per hour on a single 8-GPU node: robust and naturalistic driving emerges from 1.6 billion km of self-play, without ever seeing human data.
  4. DexPBT: Scaling up Dexterous Manipulation for Hand-Arm Systems with Population Based Training
    DexPBT: Scaling up Dexterous Manipulation for Hand-Arm Systems with Population Based Training. A Petrenko, A Allshire, G State, A Handa, V Makoviychuk. In RSS 2023.
    Large-scale reinforcement learning for high-DoF hand-arm systems.
  5. DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality
    DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality. A Handa*, A Allshire*, V Makoviychuk*, A Petrenko*, R Singh*, J Liu*, D Makoviichuk, K Van Wyk, A Zhurkevich, B Sundaralingam, Y Narang, J Lafleche, D Fox, G State. In ICRA 2023.
    Learning dexterous in-hand manipulation in vectorized simulation and deploying policies on the real robot.
  6. Megaverse: Simulating Embodied Agents at One Million Experiences per Second
    Megaverse: Simulating Embodied Agents at One Million Experiences per Second. A Petrenko, E Wijmans, B Shacklett, V Koltun. In ICML 2021.
    The fastest (at the time of release) embodied simulator for AI research. 1,000,000+ FPS of immersive experience on a single machine.
    Megaverse: Simulating Embodied Agents at One Million Experiences per SecondMegaverse: Simulating Embodied Agents at One Million Experiences per Second
  7. Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning
    Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning. A Petrenko, Z Huang, T Kumar, G Sukhatme, V Koltun. In ICML 2020.
    Reinforcement learning framework with the highest single-machine training throughput at the time of publication, about 10x faster than traditional synchronous RL implementations. SOTA results in challenging VizDoom and DMLab environments.
    Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningSample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning