Multi-Task Diffusion Policy for Robotic Manipulation

A single conditional diffusion model trained to solve Push-T and Block-Push tasks

I explored and extended a state-based diffusion policy to a multi-task setting. The goal was to train a single shared model to perform two distinct robotic manipulation tasks—Push-T (2D PyMunk) and Block-Push (PyBullet XArm)—by combining per-task processing with a unified noise-prediction network.

Push-T environment GIF
Push-T Environment
Block-Push environment GIF
Block-Push Environment

Source code: Github

What I built

  • Extended a diffusion policy to solve the PyMunk Push-T task and the PyBullet Block-Push task using a single shared 1D U-Net backbone.
  • Designed per-task linear projection heads to map native observation features (5-dim for Push-T, 16-dim for Block-Push) into a shared 64-dimensional space.
  • Implemented task conditioning by concatenating a one-hot task label to the projected observations to form a 130-dimensional global conditioning vector.
  • Built per-task action encoders and decoders to manage action spaces across different environments.
  • Created a custom collate function and weighted sampling strategies to balance heterogeneous datasets (24k Push-T samples vs 107k Block-Push samples).
  • Utilized action chunking during live environment rollouts to execute a horizon of multiple predicted actions without replanning.

Why it matters

Training a single architecture to handle multiple disparate tasks is a major step toward generalist robot agents. By using task-conditioned diffusion models with shared representations, the network learns robust, unified features while seamlessly bridging task-specific observation and action spaces.