Skip to content

Train for the Real Car

Train the car's own policy. Then prove the car's software can drive it with a software-in-the-loop (SIL) check.

Goal A policy that drives the car's own runner software, from launch to the SIL verdict.
Programs Wide Clearance (GoatRacer-Waymores-Fiesta-Gump-v2) and Tight Progress (GoatRacer-Waymores-Fiesta-Gump-v4)
Default run 1024 environments, 1500 iterations, seed 42
Time About 25 minutes of training on the lab GPU at the default knobs.
Completes when A SIL check reports PASS.

What the policy observes

The real car has no waypoints. It has a lidar. So this policy observes 17 numbers.

Group Count Meaning
Sectors 12 The lidar fan ahead of the car, split into 12 wedges, one distance each. The fan spans plus or minus 95 degrees. 24 rays are averaged into the 12 sectors.
Gap 3 Where the widest opening is, as a direction, and how deep it goes.
Last commands 2 The steer and drive the policy sent one tick ago.

The action is still two numbers: steer and drive. Everything the policy will ever know on the car is in that list. If a fact is not in the observation, the policy cannot use it.

The fan stops short of the car's own struts. Nothing the car's body does to the scan can reach the network.

Actor and critic

PPO trains two networks at once. The actor is the policy. It picks the action. The critic never drives. It estimates how much reward is still coming. That estimate is the yardstick for each update.

Only the actor ships. The critic is a training tool.

The asymmetric critic

The critic never leaves the simulator, so it can see things the car cannot. Here it also sees the car's exact position on the track and the motion sensor (the IMU). The actor sees only the 17 numbers above.

A better yardstick gives cleaner learning signals. The actor still trained on exactly what the car can measure. Nothing it learned depends on a sensor the car does not have, or on one that reads differently in the real world.

The actor is a GRU network with 128 hidden units. The GRU adds memory: part of its output rides along to the next tick.

Forward only

This car cannot drive backward. A negative drive command only brakes. The policy learns to steer out of trouble, never to reverse into what it cannot see.

The two programs

The launch form has a program picker. Both programs share the car's contract: 17 numbers in, steer and drive out, no IMU. Either can drive the runner and ship to a car. A program is named by the two things that differ: the track family it trains on, then the reward.

Program Task Tracks Reward
Wide Clearance -Gump-v2 The narrow set (1.2 to 1.6 m wide) plus Oval, Sprint, and Laguna Seca. The clearance reward. This is the baseline on the car today. It parks on a 1 m track.
Tight Progress -Gump-v4 The tight 1 m family plus the narrow set. The parking incentive is removed. Clearance is read on a 3 m horizon. Standing still always costs. Reverse is taxed.

What the reward pays for

The base reward is the same family for both programs.

Term What it does
Progress Pays for signed distance along the track.
Speed Pays for speed when the path ahead is clear, up to a cap.
Clearance brake Charges for speed when the path ahead is short.
Action rate Charges for jerky steering.
Crash A large penalty. A crash ends the episode.

Tight Progress adds a stall penalty that is not gated on clearance, and a tax on negative drive.

TODO (Rob): confirm the reward term list and whether the lap bonus should be named here.

The knobs

Knob Values Default
Program Wide Clearance, Tight Progress Wide Clearance
Parallel environments 256, 512, 1024, 2048, 4096 1024
Iterations 100 to 1500, in steps of 50 1500
Seed 0 to 9999 42

The training knobs from First Training are here too. This task adds three domain-randomization knobs for the lidar.

Knob Range What it does
Ray noise 0.0 to 0.3 m Random error added to each ray.
Ray dropout 0.0 to 0.5 The chance a ray returns nothing.
Observation latency 0.0 to 1.0 The chance the policy sees a one-tick-old observation.

TODO (Rob): confirm the default values of the three lidar domain-randomization knobs.

The form shows the cost estimate in pod-minutes. A full-length session is about 55 pod-minutes.

The SIL check

Training used an idealized car. The policy was wired straight into the simulator. It read the lidar as numbers and wrote steer and drive as numbers. No messages, no clock, no robot software.

The real robot is not wired that way. On the car, the policy runner loads the exported policy, listens to the lidar over the robot's messaging bus, builds the 17 numbers, runs the network, and publishes steer and drive to the motor controller, all on a real clock. A mistake anywhere in that chain, and a perfect policy does nothing on the floor.

SIL is short for software-in-the-loop. The check takes the exact policy runner and connects it to ONE simulated car over the same messages it uses on the robot. Then it lets it drive for 1350 ticks on one track.

The check has a track picker. Unseen is the default: one of the holdout tracks the policy never trained on. A pass on a training track proves less than a pass on a track the policy has never seen. The send card and the car's note say which track the policy was proven on.

What a PASS means

A PASS means it drives. It does not mean it drives well.

The bench reports PASS only when the runner stayed alive, stayed armed, and covered ground. The score line under the verdict is the evidence.

Score Meaning
Laps-equivalent Distance covered, in laps of the track.
Laps Whole laps completed.
Crashes How many times the car hit a wall.
Mean mph Average speed over the check.

A short training run gives a policy that passes with a low score and a lot of crashes. Judge programs on laps and crashes on the same unseen track, never on the training reward. The curve peaks and slides. The holdout does not lie.

What completes the lab

A SIL PASS completes the lab. A passing policy can go to a car. See Send a policy to the car.