clean up text
This commit is contained in:
@@ -1396,7 +1396,7 @@
|
||||
"source": [
|
||||
"We see that PPO struggles to learn locomotion in this setup, even with over 10x the number of simulator steps. \n",
|
||||
"\n",
|
||||
"Rather than indicating a shortcoming of PPO, this study demonstrates a policy-learning setup that effectively leverages FoPG methods. FoPG's perform well here because it involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n",
|
||||
"Rather than indicating a shortcoming of PPO, this study demonstrates a policy-learning setup that effectively leverages FoPG methods. It involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n",
|
||||
"\n",
|
||||
"In contrast, RL algorithms such as PPO benefit from [policy-learning setups](https://colab.research.google.com/github/google-deepmind/mujoco/blob/main/mjx/tutorial.ipynb) that have less structured rewards. Unlike FoPG methods, they [benefit greatly](https://www.science.org/doi/abs/10.1126/scirobotics.adg1462) from sparse, non-differentiable rewards such as a large penalty for falling."
|
||||
]
|
||||
|
||||
Reference in New Issue
Block a user