Labs¶
The labs are the whole climb. Drive the car, teach it to see, put your own code on the wheel, then train a policy that drives without you.
This reference covers the three Train labs. Each one starts a real training run on a GPU.
| Lab | Task | Run | What you learn |
|---|---|---|---|
| First Training | Waypoint driving, 3 observations | 1024 environments, 60 iterations | The RL loop and a healthy curve. |
| More Training Is Not Better | The same task | 1024 environments, 300 iterations | Peak versus final. Pick checkpoints by evidence. |
| Train for the Real Car | The car's own 17-observation policy | 1024 environments, up to 1500 iterations | The policy that ships, and the SIL check. |
Terms used on these pages¶
| Term | Meaning |
|---|---|
| Policy | The neural network that drives. Observation in, action out. |
| Observation | The numbers the policy sees on each step. |
| Action | Two numbers: steer and drive. |
| Reward | The score for one step. Training pushes the policy toward actions that score well. |
| Iteration | One round of drive-then-update. |
| Environment | One copy of the car in the simulator. Also called a world. |
| Seed | The random start. The same seed gives the same starting weights and the same curve. |
| Checkpoint | A saved copy of the policy weights. Training saves one every few iterations. |
| PPO | The training algorithm. It trains an actor (the policy) and a critic (a judge). |
| ONNX | The exported policy file format. The car loads it directly. |
| SIL | Software-in-the-loop. The car's real runner software drives a simulated car. |
Other labs¶
The Vision labs (AprilTags, Seeing Depth, Blocked Path, Open-Path Centroid, Obstacle Avoidance), Hello Racer, and Motor Lab run in your browser against a simulated car. They are not covered in this reference yet.
TODO (Rob): add reference pages for the Vision labs and Motor Lab when their content is final.