Skip to content

Labs

The labs are the whole climb. Drive the car, teach it to see, put your own code on the wheel, then train a policy that drives without you.

This reference covers the three Train labs. Each one starts a real training run on a GPU.

Lab Task Run What you learn
First Training Waypoint driving, 3 observations 1024 environments, 60 iterations The RL loop and a healthy curve.
More Training Is Not Better The same task 1024 environments, 300 iterations Peak versus final. Pick checkpoints by evidence.
Train for the Real Car The car's own 17-observation policy 1024 environments, up to 1500 iterations The policy that ships, and the SIL check.

Terms used on these pages

Term Meaning
Policy The neural network that drives. Observation in, action out.
Observation The numbers the policy sees on each step.
Action Two numbers: steer and drive.
Reward The score for one step. Training pushes the policy toward actions that score well.
Iteration One round of drive-then-update.
Environment One copy of the car in the simulator. Also called a world.
Seed The random start. The same seed gives the same starting weights and the same curve.
Checkpoint A saved copy of the policy weights. Training saves one every few iterations.
PPO The training algorithm. It trains an actor (the policy) and a critic (a judge).
ONNX The exported policy file format. The car loads it directly.
SIL Software-in-the-loop. The car's real runner software drives a simulated car.

Other labs

The Vision labs (AprilTags, Seeing Depth, Blocked Path, Open-Path Centroid, Obstacle Avoidance), Hello Racer, and Motor Lab run in your browser against a simulated car. They are not covered in this reference yet.

TODO (Rob): add reference pages for the Vision labs and Motor Lab when their content is final.