From 568e1f2d59d1ee4d7768bb448be2ab9411a75f4f Mon Sep 17 00:00:00 2001 From: Andrew Date: Thu, 18 Apr 2024 04:50:21 +0200 Subject: [PATCH] refine text --- mjx/training_apg.ipynb | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/mjx/training_apg.ipynb b/mjx/training_apg.ipynb index 7d4cb990..66a9b8d6 100644 --- a/mjx/training_apg.ipynb +++ b/mjx/training_apg.ipynb @@ -1382,7 +1382,7 @@ "source": [ "We see that PPO struggles to learn locomotion in this setup, even with over 10x the number of simulator steps. \n", "\n", - "This doesn't mean that PPO is a bad algorithm! Rather, this study demonstrates one way we can *design policy-learning setups that leverage FoPG methods*. FoPG's perform well in this setup because it involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n", + "Rather than indicating a shortcoming of PPO, this study demonstrates a policy-learning setup that effectively leverages FoPG methods. FoPG's perform well here because it involves learning small, accurate perturbations from a good baseline - optimizing to the local minima in a deep valley. FoPG's are precise enough to guide these perturbations using nuanced reward signals, such as our foot placement spline.\n", "\n", "In contrast, RL algorithms such as PPO benefit from [policy-learning setups](https://colab.research.google.com/github/google-deepmind/mujoco/blob/main/mjx/tutorial.ipynb) that have less structured rewards. Unlike FoPG methods, they [benefit greatly](https://www.science.org/doi/abs/10.1126/scirobotics.adg1462) from sparse, non-differentiable rewards such as a large penalty for falling." ]