01
Topics of interest
Supervision and Reward Cost: Methods for obtaining useful learning signals cheaply, including learned reward models, human feedback, human-in-the-loop interventions, and strategies for avoiding supervision as the bottleneck.Exploration and Safety on Hardware: Algorithms and systems for efficient exploration under safety constraints, reset-free or autonomous RL, safe online adaptation, and approaches that improve wall-clock efficiency during real-world training.Leveraging Priors for Real-World RL: Approaches that use pretrained policies, VLAs, offline datasets, world models, or sim-to-real transfer to reduce the number of required on-robot interactions.Post-Deployment Adaptation: Methods for finetuning deployed policies, improving robustness to distribution shift, learning new behaviors from experience, and adapting generalist policies without catastrophic forgetting.Long-Horizon and Contact-Rich RL: Techniques for improving sample efficiency in long-horizon manipulation, dexterous control, contact-rich tasks, and settings with delayed rewards, compounding errors, or difficult exploration.Benchmarks, Metrics, and Lessons Learned: Benchmarks, evaluation protocols, shared platforms, negative results, system-level insights, and analysis of what worked, what failed, and why in real-world RL experiments.