First Training¶
1024 copies of the car learn waypoint driving on a cloud GPU. You watch the curve live.
| Goal | See the RL loop whole: observations to actions to reward. Read a healthy curve. |
| Task | GoatRacer-Fiesta-3obs-v1 |
| Default run | 1024 environments, 60 iterations, seed 42 |
| Time | Medium |
| Completes when | The run finishes. |
The parallel worlds¶
The button on the Play tab starts a real run on a GPU. The simulator loads the car's digital twin and makes 1024 copies. Each copy gets its own course: 10 waypoints, about 5 m apart. There is no screen and no game window. The simulation runs headless and sends back numbers.
All 1024 cars live in one tensor on one GPU. One step of the simulator moves every car at once. More parallel worlds means the reward averages over more attempts. A crash in one world is a row reset, not a tow truck.
What the car sees¶
The policy gets three numbers on each step.
| Observation | Meaning |
|---|---|
| Distance to the next waypoint | In meters. |
| cos of the heading error | The heading error is one angle. |
| sin of the heading error | Sent as cos and sin so there is no jump from +180 to -180 degrees. |
That is the whole world, on purpose. This lab strips perception away so you can watch pure learning. The vision labs put the eyes back later.
The loop¶
The policy is a small neural network. Three numbers in, two numbers out: steer and drive. The simulator applies the actions. Physics moves the car. The run scores the step.
What the reward pays for¶
| Term | What it pays for |
|---|---|
| Progress | Moving toward the next waypoint. |
| Heading | A small bonus for facing the waypoint. |
| Goal bonus | Arriving on a waypoint. The next waypoint then becomes the target. |
The launch form exposes these as reward knobs, with a steer-rate penalty and a top-speed cap beside them. Each knob is clamped to a safe band.
| Reward knob | Range |
|---|---|
| Progress weight | 0.0 to 3.0 |
| Heading weight | 0.0 to 0.3 |
| Goal bonus | 0.0 to 30.0 |
| Steer-rate weight | 0.0 to 0.3 |
| Top speed | 30 to 120 rad/s |
TODO (Rob): confirm the default values of the five reward knobs as shown on the launch form.
The knobs¶
| Knob | Values | What it does |
|---|---|---|
| Parallel environments | 256, 512, 1024, 2048, 4096 | More worlds gives a smoother curve per iteration, and a higher cost per iteration. Past the point where the GPU is full, more worlds buys smoother, not better. |
| Iterations | 20 to 300 | Rounds of drive-then-update. 60 is the healthy recipe for this task. |
| Seed | 0 to 9999 | The random start. Same seed, same starting weights, same curve. |
Two real runs at 256 and 4096 worlds with the same seed and 60 iterations landed on the same plateau, 63 and 64. The 4096 run reached 60 by iteration 17. The 256 run reached 60 by iteration 20. The trainer logged about 10,500 steps per second at 256 worlds and about 110,000 at 4096.
The training knobs¶
A second row of knobs shapes how the policy learns, not what it is paid for. The defaults are the recipe every keeper policy was trained with.
| Knob | Default | How it breaks |
|---|---|---|
| Clip range | 0.2 | Too wide, and a good run collapses on the next update. |
| Entropy bonus | 0.01 | Zero, and the policy locks onto the first habit that scores. |
| Initial action noise | 1.0 | Too low, nothing new is tried. Too high, it crashes before it learns. |
| Learning rate | 5e-4, adaptive | 10x: the reward falls off a cliff. 0.1x: ten times the iterations for the same curve. |
| Gamma | 0.99 | Low: greedy for the next waypoint. High: slow, noisy credit. |
| Epochs per batch | 8 | Many: it overfits the batch and the KL target throttles the learning rate. |
Change one knob, keep the seed, and compare the curves. That is the whole method.
How to read the curve¶
From a real run of this exact task:
- Reward starts near zero.
- It passes 10 by iteration 10.
- It passes 70 by iteration 20.
- It levels off near 80.
If your curve climbs and flattens like that, the policy learned. Where it flattens is where more iterations stop buying skill.
What you get back¶
- The reward curve.
- A 3D playback of the trained policy threading its waypoints, in your browser: pick a camera, scrub, slow it down.
- A
policy.onnxfile. This is the same format the real car loads.
The ONNX file is the policy. The Jetson on the car runs it with no PyTorch, no GPU, and no training code. The network for this lab is a small MLP: the three observations in, dense layers with ELU, and two outputs. Open the file in Netron to see the raw graph.
What completes the lab¶
This lab completes when the run finishes. There is no reward bar. The run is the exercise. A three-question quiz checks the ideas above.