More Training Is Not Better¶
The same task as First Training, five times longer. Watch the curve peak and then slide.
| Goal | Read peak versus final on a curve. Pick checkpoints by evidence, not run length. |
| Task | GoatRacer-Fiesta-3obs-v0 |
| Default run | 1024 environments, 300 iterations, seed 42 |
| Time | Medium |
| Completes when | The run finishes with a final mean reward of at least 40. |
Same task, five times longer¶
This run is First Training with one change: 300 iterations instead of 60. If more training meant more skill, the curve should end higher. It will not. This run is broken on purpose.
Everything else is identical. Same digital twin, same 1024 worlds, same three observations, same reward. The only knob that moved is how long PPO keeps going.
What to watch for¶
A real 300-iteration run of this exact config, 1024 environments, seed 42:
- Peak 86.7 at iteration 65.
- The last 50 iterations average 50.4.
The curve climbs like before and peaks near iteration 60, in the mid 80s. Then it slides. By iteration 300 it churns around 50. That is a policy measurably worse than the one you had 200 iterations earlier. The slide reproduces run after run. It is not bad luck.
Why it gets worse¶
PPO never stops changing the weights. Every update chases the most recent batch of rollouts. Once the task is learned, that chase walks the policy away from the peak instead of holding it. Nothing in the loop knows it was done. Only you can see that, on the curve.
The fix¶
Training saves a checkpoint every few iterations. The fix is not a smarter run. It is picking the checkpoint from the peak instead of taking whatever the run ends on. Read the curve, pick the peak, evaluate that. That habit is the whole lesson. It is how the keeper policies were chosen.
The knobs¶
The knobs are the same as First Training. The default iteration count is 300, the top of the range. Keep the seed and change only the length to see the effect.
What completes the lab¶
The run must finish with a final mean reward of 40 or more. The bar is set below the post-collapse tail on purpose. The checkpoint teaches peak versus final, not pass-or-fail suspense. A two-question quiz checks the idea.