Soroush Mehraban

I received my PhD from the University of Toronto, where I was advised by Dr. Babak Taati and was a Faculty Affiliate Researcher at the Vector Institute. I am currently a Member of Technical Staff (Research) at Bagel Labs, where I work on Physical AI. Previously, I interned at Pickford AI, where I worked on video style transfer, 3D gesture generation, and text-to-3D scene generation. My research spans generative models, computer vision, and human motion analysis, including 3D human pose estimation, 3D human mesh recovery, action recognition, and gait assessment.

Email  /  CV  /  Scholar  /  Linkedin  /  YouTube  /  Github

profile photo

Research

My research focuses on generative models and their applications across a range of computer vision tasks. Below, I highlight some of my recent contributions.

SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati
arXiv
project page / arXiv / code / dataset / demo

SynthGait-19K is a physically grounded synthetic dataset of 19,272 walking videos derived from 6,427 motion-capture sequences across 437 subjects. It pairs diverse RGB videos with SMPL motion and six clinically meaningful gait parameters, and introduces GaitXFormer for direct gait-parameter estimation from video.

FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
Soroush Mehraban, Andrea Iaboni, Babak Taati
WACV2026
project page / arXiv

FastHMR introduces two merging strategies, Error Constrained Layer Merging (ECLM) and Mask guided Token Merging (Mask ToMe), to reduce computational cost and redundancy in transformer based 3D Human Mesh Recovery. ECLM selectively merges layers with minimal impact on MPJPE, while Mask ToMe merges background tokens that contribute little to prediction. A diffusion based decoder further enhances performance by using temporal context and pose priors. The method achieves up to 2.3x faster inference while slightly improving accuracy across benchmarks.

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli*, Soroush Mehraban*, Jacob Rommann, Harrison Sanborn, Babak Taati, Cole Clifford
arXiv
project page / arXiv

Puppeteer generates posture-aware, object-grounded co-speech gestures in a causal latent space.

PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
Soroush Mehraban*, Vida Adeli*, Jacob Rommann, Kyryl Truskovskyi, Harrison Sanborn, Babak Taati, Cole Clifford
ECCV 2026 (AI4VA Workshop)
project page / arXiv

PickStyle is a diffusion-based video style transfer framework that preserves video context while applying a target visual style. It uses low-rank style adapters and synthetic clip augmentation from paired images for training, and introduces Context-Style Classifier-Free Guidance (CS-CFG) to independently control content and style, achieving temporally consistent and style-faithful video results.

Synthetic faces with controllable pain expressions from Pain in 3D Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment
Xin Lei Lin*, Soroush Mehraban*, Abhishek Moturu, Babak Taati
ICPR 2026
project page / arXiv / code

Pain in 3D introduces 3DPain, a large-scale dataset of controllable synthetic pain faces with paired neutral references, action-unit annotations, and clinically grounded pain scores. It also presents ViTPain, a reference-guided vision transformer for identity-aware automated pain assessment.

dise Token Perturbation Guidance for Diffusion Models
Mohammad Javad Rajabi, Soroush Mehraban, Seyedmorteza Sadat, Babak Taati
NeurIPS 2025
GitHub / arXiv

A simple yet effective method based on token shuffling for extending the benefits of CFG to broader settings, including unconditional generation.

dise LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
Amirhossein Kazerouni, Soroush Mehraban, Michael Brudno, Babak Taati
ICCV 2025
project page / arXiv

LIFT enables unified implicit neural representations across diverse tasks by leveraging localized implicit functions and a hierarchical latent generator.

GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model
Vida Adeli, Soroush Mehraban, Majid Mirmehdi, Alan Whone, Benjamin Filtjens, Amirhossein Dadashzadeh, Alfonso Fasano, Andrea Iaboni, Babak Taati
WACV2026
project page / arXiv

GAITGen is a generative framework that synthesizes realistic gait sequences conditioned on Parkinson’s severity. Using a Conditional Residual VQ-VAE and tailored Transformers, it disentangles motion and pathology features to produce clinically meaningful gait data. GAITGen enhances dataset diversity and improves performance in parkinsonian gait analysis tasks.

dise STARS: Self-supervised 3D Action Recognition with Contrastive Tuning
Soroush Mehraban, Mohammad Javad Rajabi, Babak Taati
WACV2026
project page / arXiv

STARS enhances the Mask Autoencoder (MAE) approach in self-supervised learning by applying contrastive tuning. We also show that MAE approaches fail in few-shot settings and achieve improved performance by using the proposed method.

dise Benchmarking Skeleton-based Motion Encoder Models for Clinical Applications: Estimating Parkinson's Disease Severity in Walking Sequences
Vida Adeli, Soroush Mehraban, Irene Ballester, Yasamin Zarghami, Andrea Sabo, Andrea Iaboni, Babak Taati
FG, 2024
Code / arXiv

Evaluating recent motion encoders for the task of parkinsonism severity estimation (UPDRS III gait)

dise MotionAGFormer: Enhancing 3D Pose Estimation with a Transformer-GCNFormer Network
Soroush Mehraban Vida Adeli, Babak Taati
WACV, 2024
Code / video / arXiv

Estimating 3D locations of 17 main joints from a monocular video.


Website Template