Same traffic, different lights.
Sixteen junctions, 375 cars an hour on every entry road. Five minutes from the middle of the hour.
- On the map
- -
- Trips finished
- -
- On the map
- -
- Trips finished
- -
Scored on traffic it never saw.
Sixteen junctions with 375 cars an hour on every entry road, in a rush that turns around halfway through the hour. Near the point where the grid starts to jam, so results swing from seed to seed; the last column counts the seeds where the learned lights won outright.
| Controller | Mean trip | Spread | Stranded cars | Learned wins |
|---|---|---|---|---|
| Learned | 371.8 s | ± 132.0 | 167 | - |
| Max pressure | 448.6 s | ± 154.8 | 204 | 10 of 10 |
| Fixed-time | 536.6 s | ± 150.0 | 586 | 9 of 10 |
| Actuated | 549.2 s | ± 270.4 | 762 | 8 of 10 |
What actually mattered.
One change at a time from the main network: no view of the roads ahead, a separate network per junction, the pressure reward, the one-junction network used as it is, and training from scratch.
| Controller | Mean trip | Spread | Stranded cars | Learned wins |
|---|---|---|---|---|
| Blind to outgoing roads | 323.6 s | ± 95.3 | 67 | 0 of 10 |
| Learned | 371.8 s | ± 132.0 | 167 | - |
| Max pressure | 448.6 s | ± 154.8 | 204 | 10 of 10 |
| Pressure reward | 470.1 s | ± 169.8 | 492 | 9 of 10 |
| Fixed-time | 536.6 s | ± 150.0 | 586 | 9 of 10 |
| Actuated | 549.2 s | ± 270.4 | 762 | 8 of 10 |
| A network per junction | 750.2 s | ± 317.7 | 1381 | 10 of 10 |
| One-junction network | 818.0 s | ± 123.9 | 1880 | 10 of 10 |
| Cold start | 1171.4 s | ± 77.4 | 2632 | 10 of 10 |
Where each one breaks down.
- Learned
- Actuated
- Fixed-time
- Max pressure
| Controller | 250 | 300 | 350 | 375 |
|---|---|---|---|---|
| Learned | 193 s | 213 s | 262 s | 411 s |
| Actuated | 222 s | 246 s | 299 s | 622 s |
| Fixed-time | 258 s | 288 s | 390 s | 534 s |
| Max pressure | 248 s | 273 s | 331 s | 499 s |
- Learned
- Actuated
- Fixed-time
- Max pressure
| Controller | 300 | 350 | 400 | 450 |
|---|---|---|---|---|
| Learned | 195 s | 211 s | 247 s | 557 s |
| Actuated | 225 s | 246 s | 275 s | 745 s |
| Fixed-time | 246 s | 264 s | 326 s | 663 s |
| Max pressure | 269 s | 286 s | 315 s | 514 s |
A green wave, found by itself.
- Main-street trip
- 230 s
- All trips
- 134 s
- Offsets between greens
- 25, 30, 30, 30, 30 s
How long it took to learn.
- gamma 0.97, pressure
- gamma 0.97, queue
- gamma 0.9, queue
Mean trip time on two validation seeds, checked every 10 training episodes. Thin flat lines: the tuned hand-written controllers on the same traffic (Actuated 90 s, Max pressure 94 s, Fixed-time 106 s).
- warm start
- cold start
- one network per junction
Mean trip time on two validation seeds, checked every 10 training episodes. Thin flat lines: the tuned hand-written controllers on the same traffic (Fixed-time 438 s, Max pressure 470 s, Actuated 618 s).
Small network, strict rules.
What went wrong first.
- Done:Lights that never settledLooking only about 50 seconds ahead, the first network switched around 220 times an hour, losing 5 s to every change, and stalled far behind the hand-written controllers. Looking about two minutes ahead fixed it.
- Done:A car left at a red for 8 minutesCounting waiting cars, one lone car barely registers, so one network left it at a red while it cycled through empty phases. Max pressure starved phases too. Now any phase with a car waiting 5 minutes is served next, for every controller.
- Done:A flawed formulaA widely shared version of the pressure reward subtracts the free space on the outgoing roads instead of the cars on them, which makes letting a car through worth nothing. Trained the same way, it averaged 133 s per trip at a rush-hour junction, worse than a fixed timer (108 s); the correct formula averaged 89 s.
- Done:Pressure goes blind on a gridPressure, cars waiting minus cars ahead, works on a lone junction whose exits never fill. On a grid, a map jammed solid looks perfectly balanced to it. The grid network learns from the number of waiting cars instead.
- Done:Starting from nothing jams the gridWith sixteen junctions exploring at random, the grid locks up and there is little to learn from a map where nothing moves. Starting from the one-junction network and exploring gently fixed it; from scratch, the network was still losing to every hand-written controller after 100 hours of traffic.
- Done:One network per light learns slowerGiving every junction its own network, with the same head start and the same training, means each one learns from a sixteenth of the experience. Trip times roughly doubled (750 s against 372 s at rush hour) and it lost to every hand-written controller. Sharing one network lets a lesson learned at one corner help every other.