Multi-agent reinforcement learning

Reward shaping PPO via self-play curriculum.

This project trained PPO agents for 2v2 Soccer-Twos on PACE HPC, studying how reward shaping, opponent mixtures, and separated policy/value networks affect stability, catastrophic forgetting, and tournament performance.

  • 29 training experiments
  • 62% -> 15% catastrophic forgetting
  • 1st 72-team class tournament
PPO self-play versus reward shaped training analysis

PPO Self-Play vs Reward Shaping

Training analysis comparing episode reward, policy loss, value loss, and entropy.

Method

The final agent used PPO with separated policy/value networks and a mixed opponent curriculum: 70% self-play opponents and 30% fixed heuristic baseline. Reward shaping added dense signals for approach, kicking, offense, defense, time penalty, and teammate separation.

PPO Ray RLlib Self-Play Reward Shaping PACE HPC Ablation
01

Self-Play Curriculum

Maintained rolling opponent snapshots while mixing in a fixed CEIA baseline opponent.

02

Reward Shaping

Added dense football-specific signals while keeping the sparse goal reward dominant.

03

Ablation Study

Tested 12+ strategies including BC pre-training, frame stacking, entropy schedules, and reward variants.

04

Failure Analysis

Logged opponent overfitting, distribution shift, asymmetric team performance, and catastrophic forgetting.

Results

The final experiment identified separated value/policy networks plus a self-play-dominant opponent mix as the most stable design, producing a 90% evaluation win rate against the CEIA baseline and winning the class tournament.

PPO self-play versus reward shaped training curves

Training Analysis

Reward shaping accelerated early learning, while the final self-play curriculum and separated networks reduced forgetting and stabilized policy improvement.