Introduction
Multi-agent autonomous racing requires agents to balance individual performance, cooperation with teammates and competition against opponents. In a 2-vs-2 setting, effective decisions may involve coordinated overtaking, defensive blocking or sacrificing short-term individual performance to improve the team outcome. Model-based reinforcement learning provides a promising framework for this problem by learning a predictive world model and optimizing policies through imagined multi-agent trajectories.
This thesis investigates joint latent imagination for 2-vs-2 autonomous racing, with the goal of learning coordinated team strategies without relying on an explicitly designed high-level tactical planner.
Goals
- 2-vs-2 world model: develop a latent predictive world model of four interacting racing vehicles, combining individual, teammate and opponent performance within the learning objective;
- Training strategy: investigate curriculum learning and self-play to improve stability and progressive skill acquisition;
- Emergent cooperation: evaluate whether coordinated overtaking, blocking, role differentiation and team-oriented sacrifices emerge without being explicitly prescribed;
- Validation: assess team performance, coordination, collisions, learning efficiency and robustness to different initial conditions and opponent strategies.
Requirements
- Knowledge of Python;
- Basic knowledge of machine learning and reinforcement learning (can be learnt during the project, familiarity with deep learning frameworks such as PyTorch is a plus)
Contact
Marco Doria Fragomeni: marco.doria@polimi.it
Stefano Arrigoni: stefano.arrigoni@polimi.it
