Practical Model-Based Policy Optimization for Efficient Reinforcement Learning
Recent advancements in probabilistic model-based reinforcement learning (MBRL) have accelerated learning by generating samples from the model. However, this approach often suffers from time inefficiencies caused by frequent model updates. To address this issue, researchers from the Chinese Academy of Sciences propose a novel model-based policy optimization (PMBPO) framework that enhances the reliability of generated samples and reduces computation time. PMBPO achieves remarkable results on Mujoco control benchmarks and a quadruped robot control scenario, reducing one-step computation time by 90% and surpassing state-of-the-art MBRL approaches in learning efficiency and control capability.
Key Takeaways:
- The PMBPO framework enhances the reliability of generated samples by introducing an expressive probabilistic model that focuses on system dynamic features over continuous time steps.
- PMBPO reduces the one-step computation time by 90% compared to state-of-the-art MBRL approaches.
- Evaluated on five Mujoco control benchmarks and one quadruped robot control scenario, PMBPO achieves 70% more cumulative rewards than state-of-the-art MBRL approaches.
- PMBPO extends the feasibility of MBRL in practical control scenarios.
- The code of PMBPO is available at https://github.com/mrjun123/PMBPO.
- The research concludes that PMBPO reduces the computational burden of real-time learning while surpassing recent MBPO approaches in learning efficiency and control capability.
Statistics:
- PMBPO reduces one-step computation time by 90% compared to state-of-the-art MBRL approaches.
- PMBPO achieves 70% more cumulative rewards than state-of-the-art MBRL approaches.
- Evaluated on five Mujoco control benchmarks and one quadruped robot control scenario.
- The code of PMBPO is available at https://github.com/mrjun123/PMBPO.
- The research is supported by the National Natural Science Foundation of China (NSFC) and the Major Program of Science and Technology of Shenzhen.
Sources:
- VerticalNews. "Chinese Academy of Sciences Discusses New Findings on Technology." June 2, 2025.
- IEEE Transactions on Automation Science and Engineering. "Practical Reinforcement Learning Using Time-efficient Model-based Policy Optimization." Vol. 22, No. 3, 2025, pp. 14436-14447.
- Chinese Academy of Sciences, Shenzhen Institutes of Advanced Technology, Shenzhen 518055, People's Republic of China.