Multi-Task Diffusion Policy for Robotic Manipulation
A single conditional diffusion model trained to solve Push-T and Block-Push tasks
I explored and extended a state-based diffusion policy to a multi-task setting. The goal was to train a single shared model to perform two distinct robotic manipulation tasks—Push-T (2D PyMunk) and Block-Push (PyBullet XArm)—by combining per-task processing with a unified noise-prediction network.
Push-T Environment
Block-Push Environment
Source code: Github
What I built
- Extended a diffusion policy to solve the PyMunk Push-T task and the PyBullet Block-Push task using a single shared 1D U-Net backbone.
- Designed per-task linear projection heads to map native observation features (5-dim for Push-T, 16-dim for Block-Push) into a shared 64-dimensional space.
- Implemented task conditioning by concatenating a one-hot task label to the projected observations to form a 130-dimensional global conditioning vector.
- Built per-task action encoders and decoders to manage action spaces across different environments.
- Created a custom collate function and weighted sampling strategies to balance heterogeneous datasets (24k Push-T samples vs 107k Block-Push samples).
- Utilized action chunking during live environment rollouts to execute a horizon of multiple predicted actions without replanning.
Why it matters
Training a single architecture to handle multiple disparate tasks is a major step toward generalist robot agents. By using task-conditioned diffusion models with shared representations, the network learns robust, unified features while seamlessly bridging task-specific observation and action spaces.