diff --git a/mjx/training_apg.ipynb b/mjx/training_apg.ipynb index bb7d35ed..598f06d7 100644 --- a/mjx/training_apg.ipynb +++ b/mjx/training_apg.ipynb @@ -1396,7 +1396,7 @@ "source": [ "We see that PPO struggles to learn locomotion in this setup, even with over 10x the number of simulator steps. \n", "\n", - "Rather than indicating a shortcoming of PPO, this study demonstrates a policy-learning setup that effectively leverages FoPG methods. FoPG's perform well here because it involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n", + "Rather than indicating a shortcoming of PPO, this study demonstrates a policy-learning setup that effectively leverages FoPG methods. It involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n", "\n", "In contrast, RL algorithms such as PPO benefit from [policy-learning setups](https://colab.research.google.com/github/google-deepmind/mujoco/blob/main/mjx/tutorial.ipynb) that have less structured rewards. Unlike FoPG methods, they [benefit greatly](https://www.science.org/doi/abs/10.1126/scirobotics.adg1462) from sparse, non-differentiable rewards such as a large penalty for falling." ]