Skip to content

Send a policy to the car

A policy that passed the SIL check can go to a car. This page says what happens, what the verdict means, and what to expect on the floor.

Before you send

  • The run finished and exported policy.onnx.
  • The SIL check reports PASS. See Train for the Real Car.
  • A car is assigned to you. If the send card says no car is assigned, ask your instructor.
  • The car is switched on and has an internet connection.

What the SIL verdict means

Verdict Meaning
PASS The car's runner software loaded the policy, stayed alive, stayed armed, and covered ground for 1350 ticks on one simulated track. The chain works.
FAIL The runner stopped, disarmed, or the car did not move. The bench line says why.

PASS means it drives. It does not mean it drives well. Read the score line: laps-equivalent, laps, crashes, and mean mph. The Unseen track is the honest test. The send card and the car's note say which track the policy was proven on.

Send it

  1. Open the run in Gym, or the session panel after the SIL check.
  2. On the send card, pick the car and write a short note. Say what changed and what to look for on the floor.
  3. Click Send to car and confirm.

The policy ships as an over-the-air update. It is exported, versioned, and added to the car's policy library. The car gets it soon after. It appears on the policy page with a NEW badge and its version. That is how you know it arrived.

TODO (Rob): state how long a send takes to arrive on the car under normal conditions.

On the floor

  1. Power up the car. Kill switch in hand. See Safety first.
  2. Open the dashboard and go to the policy page.
  3. Select the new policy.
  4. Pick the runner mode.
Runner mode What it does
Observe-only The runner computes every tick and publishes nothing. The safe first run of a new policy.
eRPM The runner sends wheel-speed targets. This is the training semantic.
Current The runner sends torque. Experimental.

Start with observe-only. Watch the readout. When the commands look sane, switch to eRPM and select the RL Policy strategy on the dashboard.

What to expect

  • The policy sees only the lidar. It does not see lines, colors, or you.
  • It drives forward only. A negative drive command brakes.
  • It steers toward the widest gap ahead. On an open floor it may wander. Give it walls: a track of tubes, boxes, or boards.
  • Wide Clearance may park on a track about 1 m wide. Tight Progress was trained for that width.
  • A crash in simulation is a reset. A crash on the floor is a crash. Keep your thumb on B.

Stop it

  • Select Stop on the dashboard, or press B on the kill switch.
  • Stopping the policy does not start teleop. Select Teleop when you want the gamepad back.