
AWS demonstrated training the Qwen3-VL-8B vision-language model using the open-source SkyRL framework on SageMaker HyperPod. By applying Group Relative Policy Optimization post-training to a supervised fine-tuning checkpoint, researchers increased the model's success rate in navigating visual mazes from 43.75% to over 95% on a fixed evaluation set.
Reinforcement learning for vision-language models
The training process utilized Group Relative Policy Optimization to refine the Qwen3-VL-8B model, which initially started from a supervised fine-tuning checkpoint. This multi-turn reinforcement learning approach allows the agent to learn from rewards accumulated over an entire sequence of steps within a visual maze, rather than receiving feedback only for single actions.
To manage the heavy computational demands of these multi-node runs, the framework ran on Amazon SageMaker HyperPod. This infrastructure utilizes cluster resiliency features to monitor node health and automatically replace faulty hardware, allowing long-running reinforcement learning jobs to resume from checkpoints instead of restarting from scratch after a failure.
Evaluation results and infrastructure limits
Testing conducted on a fixed 64-maze evaluation set showed that the reinforcement learning post-training significantly improved performance over the baseline. The solve rate rose from 43.75% to more than 95%. The system integrates with Ray clusters and uses managed dashboards for monitoring training dynamics and observability.
The reported success is based on a specific, fixed set of 64 mazes. The source does not provide performance data for unseen or procedurally generated environments outside of this evaluation set. While the infrastructure supports large-scale workloads, the training duration and total GPU-hours required for these results were not specified.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at AWS ML