Follow two recorded performances, a comparison and an evaluator update.
A selected example from the recorded training of 32 agents. Policy copying was disabled.
A Recorded Exchange
The performances are replayed together from their recorded starting states. Each stops when that recorded performance ended.
The Sending Agent Compares Both Performances
| Evaluator | Performance A | Performance B |
|---|
| | |
|---|
A performance score is the sum of the evaluator's rewards over that recorded motion. The higher score determines the comparison label.
The Receiving Agent Revises the Evaluator
| Evaluator | Performance A | Performance B | Preferred |
|---|
| | | |
|---|
| | | |
|---|
These scores apply the recorded evaluators to the same two performances. Subsequent policy trials are scored with the updated evaluator.
Explore the Full Recorded Population
Choose a training excerpt or another exchange. Population playback follows the recording while the selected exchange above stays fixed.
Sending AgentReceiving AgentNew Policy KeptRecovering After a Fall
Recording and Results
These are four forty-second excerpts from a completed eighty-minute training run. The population recording shows actual training. The two performance replays show the completed records used in the selected exchange.
The recording comes from training seed 1001. The opening example was selected to show an evaluator switching from preferring a performance that ended in a fall to preferring a completed four-second performance. The run and opening exchange are selected examples. All 1,325 recorded exchanges remain available in the population explorer. The article reports results across all sixteen paired runs.
The recording reproduces all 120 saved training checkpoints and the final policies and evaluators exactly. CartPole advances every 0.02 seconds. Population playback is sampled every 0.1 seconds. After failure, an agent waits one second before starting again. A completed four-second performance restarts immediately.
Recording Provenance