Skip to content

First Training

1024 copies of the car learn waypoint driving on a cloud GPU. You watch the curve live.

Goal See the RL loop whole: observations to actions to reward. Read a healthy curve.
Task GoatRacer-Fiesta-3obs-v1
Default run 1024 environments, 60 iterations, seed 42
Time Medium
Completes when The run finishes.

The parallel worlds

The button on the Play tab starts a real run on a GPU. The simulator loads the car's digital twin and makes 1024 copies. Each copy gets its own course: 10 waypoints, about 5 m apart. There is no screen and no game window. The simulation runs headless and sends back numbers.

All 1024 cars live in one tensor on one GPU. One step of the simulator moves every car at once. More parallel worlds means the reward averages over more attempts. A crash in one world is a row reset, not a tow truck.

What the car sees

The policy gets three numbers on each step.

Observation Meaning
Distance to the next waypoint In meters.
cos of the heading error The heading error is one angle.
sin of the heading error Sent as cos and sin so there is no jump from +180 to -180 degrees.

That is the whole world, on purpose. This lab strips perception away so you can watch pure learning. The vision labs put the eyes back later.

The loop

The policy is a small neural network. Three numbers in, two numbers out: steer and drive. The simulator applies the actions. Physics moves the car. The run scores the step.

What the reward pays for

Term What it pays for
Progress Moving toward the next waypoint.
Heading A small bonus for facing the waypoint.
Goal bonus Arriving on a waypoint. The next waypoint then becomes the target.

The launch form exposes these as reward knobs, with a steer-rate penalty and a top-speed cap beside them. Each knob is clamped to a safe band.

Reward knob Range
Progress weight 0.0 to 3.0
Heading weight 0.0 to 0.3
Goal bonus 0.0 to 30.0
Steer-rate weight 0.0 to 0.3
Top speed 30 to 120 rad/s

TODO (Rob): confirm the default values of the five reward knobs as shown on the launch form.

The knobs

Knob Values What it does
Parallel environments 256, 512, 1024, 2048, 4096 More worlds gives a smoother curve per iteration, and a higher cost per iteration. Past the point where the GPU is full, more worlds buys smoother, not better.
Iterations 20 to 300 Rounds of drive-then-update. 60 is the healthy recipe for this task.
Seed 0 to 9999 The random start. Same seed, same starting weights, same curve.

Two real runs at 256 and 4096 worlds with the same seed and 60 iterations landed on the same plateau, 63 and 64. The 4096 run reached 60 by iteration 17. The 256 run reached 60 by iteration 20. The trainer logged about 10,500 steps per second at 256 worlds and about 110,000 at 4096.

The training knobs

A second row of knobs shapes how the policy learns, not what it is paid for. The defaults are the recipe every keeper policy was trained with.

Knob Default How it breaks
Clip range 0.2 Too wide, and a good run collapses on the next update.
Entropy bonus 0.01 Zero, and the policy locks onto the first habit that scores.
Initial action noise 1.0 Too low, nothing new is tried. Too high, it crashes before it learns.
Learning rate 5e-4, adaptive 10x: the reward falls off a cliff. 0.1x: ten times the iterations for the same curve.
Gamma 0.99 Low: greedy for the next waypoint. High: slow, noisy credit.
Epochs per batch 8 Many: it overfits the batch and the KL target throttles the learning rate.

Change one knob, keep the seed, and compare the curves. That is the whole method.

How to read the curve

From a real run of this exact task:

  • Reward starts near zero.
  • It passes 10 by iteration 10.
  • It passes 70 by iteration 20.
  • It levels off near 80.

If your curve climbs and flattens like that, the policy learned. Where it flattens is where more iterations stop buying skill.

What you get back

  • The reward curve.
  • A 3D playback of the trained policy threading its waypoints, in your browser: pick a camera, scrub, slow it down.
  • A policy.onnx file. This is the same format the real car loads.

The ONNX file is the policy. The Jetson on the car runs it with no PyTorch, no GPU, and no training code. The network for this lab is a small MLP: the three observations in, dense layers with ELU, and two outputs. Open the file in Netron to see the raw graph.

What completes the lab

This lab completes when the run finishes. There is no reward bar. The run is the exercise. A three-question quiz checks the ideas above.