Embodied AI · Robot Learning · Imitation Learning

Mousumi Das

M.S. in Electrical & Computer Engineering,
University of Southern California.

Mousumi Das MD

Hi there! I'm an M.S. student in Electrical & Computer Engineering at the University of Southern California. I build robots that learn efficiently and reliably — from expert demonstrations, reward signals and autonomous interaction.

I work with Jesse Thomason in the GLAMOR Lab on composing vision-language-action policies during execution, and with Erdem Bıyık in the LIRA Lab on demonstration-efficient imitation learning. Currently I am interning at the Georgia Tech Research Institute with Colin Usher on vision-based pick-selection for deformable objects.

I'm most excited about where this goes next: robots that adapt on their own, from interaction rather than new supervision.

News

  • Jul 2026

    DiPS, our work on dialogue policy selection for high-stakes persuasion agents, was accepted to SIGDIAL 2026!

  • Jun 2026

    Submitted SWAP, a stepwise action policy-routing framework for VLAs, to CoRL 2026.

  • May 2026

    Started as a Research Intern at the Georgia Tech Research Institute with Colin Usher, working on robotic pick-selection for deformable objects.

  • Jan 2026

    Joined the LIRA Lab at USC with Erdem Bıyık, working on teacher-aware imitation learning.

  • Jun 2025

    Joined the GLAMOR Lab at USC with Jesse Thomason, working on language-guided robots.

  • Jan 2025

    Started my M.S. in Electrical & Computer Engineering at the University of Southern California.

Research

Publications

Robot's-eye RGB view of deformable poultry carcasses on the conveyor. SAM2 instance segmentation producing 14 candidate carcass masks. Ranker scores each candidate and selects the top-ranked pick. Outcome frame showing the executed pick attempt.

Foundation-Model Pick Selection for Deformable Objects with Human-in-the-Loop Ranking

Mousumi Das, Colin Usher

It is hard to grasp multiple densely-packed deformable objects in clutter. We learn a pick ranker from operator demonstrations that continually improves from human corrections at deployment, on top of perception and grasping foundation models.

In preparation. Target: ICRA 2027.

Media
coming soon

Imitation Learning from Deliberative Teachers

Mousumi Das, Erdem Bıyık

Human demonstrators don't act randomly — they teach. TV-BC upweights the hard, informative demonstrations a teacher would emphasize, learning more sample-efficiently and generalizing better out-of-distribution than standard behavioral cloning.

In preparation. Target: ICLR 2027.

SWAP: Stepwise Action Policy Routing for Vision-Language-Action Models

Mousumi Das, Aditeya Prajapati, Abrar Anwar, Jesse Thomason

No single policy is best at every step of a task. SWAP learns to switch between vision-language-action policies mid-rollout, improving real-robot success by up to 33% over the best fixed policy.

Submitted to CoRL 2026.

DiPS architecture: dialogue history is encoded, an IQL policy selector picks a persona, and an LLM generates the operator response, with an LLM judge monitoring the goal.

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

Tianyi Zhang*, Mousumi Das*, Abrar Anwar, Jesse Thomason, David Traum

Persuasion isn't one-size-fits-all. In a wildfire-evacuation setting, DiPS learns to pick the right persuasion strategy turn-by-turn, outperforming zero-shot and RAG baselines with both simulated and real people.

SIGDIAL 2026 (Oral Presentation)

Media
coming soon

Hand Gesture Recognition using DenseNet201–Mediapipe Hybrid Modelling

Prachetas Padhi*, Mousumi Das*

A hybrid DenseNet201 + Mediapipe model for robust hand-gesture recognition.

ICACRS, IEEE 2022.

* denotes equal contribution.

Interests

  • Imitation Learning — sample-efficient learning from demonstrations that generalizes out-of-distribution.
  • VLA Compositionality — composing and routing vision-language-action policies at execution time.
  • Language Guided Robotics — language conditioned policies for embodied decision making and interaction.
  • Perception for Manipulation — grounded perception with foundation models for contact rich manipulation.

Selected Projects

Legged Locomotion · Granular Media

Stride Amplitude for Quadruped Sand-Slope Climbing

Pushing a leg harder into sand doesn't always climb faster. On a custom 2-DoF-per-leg quadruped and an adjustable sand ramp, we swept stride amplitude against slope to find a sweet spot where moderate strides gain traction without avalanching the substrate, tripling uphill displacement over the baseline.

Robotics · ASR · Vision

Voice & Vision Controlled Robotic Hand

A robotic hand that mirrors a user in real time using two input modes at once: speech recognition for voice commands and Mediapipe landmark tracking for live gesture replication. Detected finger-joint angles map to servo commands on a microcontroller, keeping end-to-end latency under 100 ms.

Multi-Robot · Computer Vision

Automated Sortation of Packages in a Warehouse

A swarm of mobile robots sort packages under an overhead camera, each one tracked by its own ArUco marker. A central planner assigns A* routes and relays them over WiFi to onboard microcontrollers, cutting sort-cycle time by 35% against a fixed-route baseline.

Get in Touch

I'm always happy to discuss robot learning and research collaborations.