PPO Social Learning

Loading saved experiment replays…

The camera follows each cart. These experiments allow unrestricted cart travel and pole rotation; crossing the usual CartPole boundaries does not end a trial.

0.00 s
About These Replays

Each replay shows a final, frozen PPO policy on a reserved near-upright start. All six agents from all five training populations are available. The first two test starts were included for every agent without selecting for success. Paired conditions use identical starts.

The percentages include all 720 near-upright test trajectories per condition. Upright posture means an angle RMS below five degrees during the final five seconds, with angle measured modulo a full rotation. Cart position is a separate measurement. These are not standard CartPole completion rates.

These policies have 57 parameters and were trained in fixed 200-step episodes. In the private control, each learner's initial reward function stays fixed. In the teacher condition, the teacher scores performances using a fixed reward function that favors uprightness. Learners receive labels identifying which performance the teacher scored higher and fit their own reward functions to those labels. The shuffled control randomizes preference labels. In the peer condition, learners receive the labels generated by other learners' reward functions.