Traffic signal control

Traffic lights
that learn.

Every light on a simulated city grid runs the same small neural network. It learned by trial and error, then faced traffic it had never seen against the best hand-written signals, each tuned the same way.

Simulated in SUMO. Scored on traffic held out from training.

Learned lights20:00
Loading replay
One network, sixteen lights- cars
17%
Shorter trips than the best tuned controller, 16-junction grid at rush hour
10 of 10
Rush-hour test seeds where it beat tuned max pressure on the grid
30 s
Green-wave offset it found on its own on a one-way street (the textbook answer: 30 s)
Replays

Same traffic, different lights.

Pick a map and two controllers. Both sides replay exactly the same cars, so every difference comes from the lights. Stopped cars show in orange.

Sixteen junctions, 375 cars an hour on every entry road. Five minutes from the middle of the hour.

Neural network
Loading replay
On the map
-
Trips finished
-
Hand-written
Loading replay
On the map
-
Trips finished
-
0:00
Moving car Stopped car Green Yellow Red
Results

Scored on traffic it never saw.

Every number is a mean over 10 traffic seeds (5 on the 8x8 grid) that were never used for training, tuning or picking checkpoints. Trip time counts from each car's scheduled departure, so a controller can't look good by keeping cars off the map, and cars still stuck at the end count until the end.

Sixteen junctions with 375 cars an hour on every entry road, in a rush that turns around halfway through the hour. Near the point where the grid starts to jam, so results swing from seed to seed; the last column counts the seeds where the learned lights won outright.

Mean trip time, 10 held-out traffic seeds (lower is better)
ControllerMean tripSpreadStranded carsLearned wins
Learned371.8 s± 132.0167-
Max pressure448.6 s± 154.820410 of 10
Fixed-time536.6 s± 150.05869 of 10
Actuated549.2 s± 270.47628 of 10
Experiments

What actually mattered.

The main network was chosen before any of these tests ran. Each variant changes one thing; the hand-written controllers stay in the chart for scale.

One change at a time from the main network: no view of the roads ahead, a separate network per junction, the pressure reward, the one-junction network used as it is, and training from scratch.

Mean trip time, 10 held-out traffic seeds (lower is better)
ControllerMean tripSpreadStranded carsLearned wins
Blind to outgoing roads323.6 s± 95.3670 of 10
Learned371.8 s± 132.0167-
Max pressure448.6 s± 154.820410 of 10
Pressure reward470.1 s± 169.84929 of 10
Fixed-time536.6 s± 150.05869 of 10
Actuated549.2 s± 270.47628 of 10
A network per junction750.2 s± 317.7138110 of 10
One-junction network818.0 s± 123.9188010 of 10
Cold start1171.4 s± 77.4263210 of 10
Traffic volume

Where each one breaks down.

The same controllers on the grid from light traffic to the edge of gridlock, 5 seeds per point. Hand-written controllers keep the settings tuned at the scored volume, the way a city tunes its signals once.
Rush hour: mean trip by cars per hour per entry road
  • Learned
  • Actuated
  • Fixed-time
  • Max pressure
Controller250300350375
Learned193 s213 s262 s411 s
Actuated222 s246 s299 s622 s
Fixed-time258 s288 s390 s534 s
Max pressure248 s273 s331 s499 s
Flat traffic: mean trip by cars per hour per entry road
  • Learned
  • Actuated
  • Fixed-time
  • Max pressure
Controller300350400450
Learned195 s211 s247 s557 s
Actuated225 s246 s275 s745 s
Fixed-time246 s264 s326 s663 s
Max pressure269 s286 s315 s514 s
Green wave

A green wave, found by itself.

On a one-way street with straight traffic, the best timing is known: each junction turns green one block's travel time after the one before, 30 seconds here. Nobody told the network. Read the chart left to right: each line is a car, and a line that climbs without flattening never stopped. Tuned actuated control finds the same offsets, because on a one-way street reacting to the cars that arrive is enough.
A car on the main street Green Yellow Red
Main-street trip
230 s
All trips
134 s
Offsets between greens
25, 30, 30, 30, 30 s
Training

How long it took to learn.

Each training episode is one simulated hour of traffic with a fresh seed and volume.
One junction at rush hour
  • gamma 0.97, pressure
  • gamma 0.97, queue
  • gamma 0.9, queue

Mean trip time on two validation seeds, checked every 10 training episodes. Thin flat lines: the tuned hand-written controllers on the same traffic (Actuated 90 s, Max pressure 94 s, Fixed-time 106 s).

16-junction grid at rush hour
  • warm start
  • cold start
  • one network per junction

Mean trip time on two validation seeds, checked every 10 training episodes. Thin flat lines: the tuned hand-written controllers on the same traffic (Fixed-time 438 s, Max pressure 470 s, Actuated 618 s).

How it works

Small network, strict rules.

One network, every light
Each junction gives the same network 46 numbers about its own corner of the map: stopped cars, cars coming in three stretches of each approach, how full the roads it feeds are, and which phase it shows and for how long. Four numbers come back, one per phase: how good showing it next would be.
Four phases, fixed rules
North-south and east-west each get a through phase and a protected left-turn phase. A switch costs 3 s of yellow and 2 s of all-red, a green lasts at least 10 s, a green ends after 60 s if anyone else is waiting, and no car waits at a red for more than 5 minutes. Every controller lives by the same rules.
How it learns
Double DQN: the network guesses the value of each phase, acts, sees how many cars end up waiting, and corrects its guess, looking about two minutes ahead. On the grid it starts from the one-junction network, which skips the jams that random exploration causes.
The competition
Max pressure (serve the phase with the most cars waiting minus cars ahead), actuated signals (hold green while cars keep arriving) and fixed-time plans. Each gets the same tuning budget on separate traffic before the test.
Honest scoring
Training, checkpoint choice, tuning and scoring each use their own traffic seeds. Trip time counts from each car's scheduled departure, stranded cars count to the end of the run, and teleporting is off, so gridlock stays gridlock.
The simulator
SUMO, the open-source microscopic traffic simulator: every car accelerates, brakes, changes lanes and queues. Cars take one of several shortest routes and keep it, even when the road ahead is jammed.
Along the way

What went wrong first.

The fixes behind the results, in the order they turned up.
  1. Done:
    Lights that never settled
    Looking only about 50 seconds ahead, the first network switched around 220 times an hour, losing 5 s to every change, and stalled far behind the hand-written controllers. Looking about two minutes ahead fixed it.
  2. Done:
    A car left at a red for 8 minutes
    Counting waiting cars, one lone car barely registers, so one network left it at a red while it cycled through empty phases. Max pressure starved phases too. Now any phase with a car waiting 5 minutes is served next, for every controller.
  3. Done:
    A flawed formula
    A widely shared version of the pressure reward subtracts the free space on the outgoing roads instead of the cars on them, which makes letting a car through worth nothing. Trained the same way, it averaged 133 s per trip at a rush-hour junction, worse than a fixed timer (108 s); the correct formula averaged 89 s.
  4. Done:
    Pressure goes blind on a grid
    Pressure, cars waiting minus cars ahead, works on a lone junction whose exits never fill. On a grid, a map jammed solid looks perfectly balanced to it. The grid network learns from the number of waiting cars instead.
  5. Done:
    Starting from nothing jams the grid
    With sixteen junctions exploring at random, the grid locks up and there is little to learn from a map where nothing moves. Starting from the one-junction network and exploring gently fixed it; from scratch, the network was still losing to every hand-written controller after 100 hours of traffic.
  6. Done:
    One network per light learns slower
    Giving every junction its own network, with the same head start and the same training, means each one learns from a sixteenth of the experience. Trip times roughly doubled (750 s against 372 s at rush hour) and it lost to every hand-written controller. Sharing one network lets a lesson learned at one corner help every other.
Questions

Fair questions.