Huang (Raven) Huang

I am a postdoctoral researcher at Stanford University, advised by Prof. Jiajun Wu and Prof. Fei-Fei Li. I received my Ph.D. in EECS from UC Berkeley, advised by Prof. Ken Goldberg.

My research focuses on general robot learning, particularly learning from human videos, real2sim2real transfer, and dexterous manipulation. I develop methods for scalable robot data generation, cross-embodiment world models, in-context imitation learning, and multimodal perception with touch.

Email  /  Google Scholar  /  CV

profile photo

Honors and Service

Selected Publications
RoboRender: Robot-Oriented Video Generation for Visual Sim-to-Real Transfer
Huang Huang*, Wensi Ai*, Ziyu Chen*, Youhui Wang*, Zijian Du, Yang Liu, Jiaolong Yang, Li Fei-Fei, Jiajun Wu *Equal contribution
arXiv 2026 · Website · Paper

We convert simulated robot trajectories into photorealistic RGB videos while preserving geometry, robot motion, and action labels, enabling visual sim-to-real transfer for robot policies.

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library
Youhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu†, Huang Huang† †Equal advising
arXiv 2026 · Website · Paper

DexAgent turns a human demonstration video into dexterous robot training data through simulation reconstruction, trajectory optimization, and verification, using a self-evolving library of reusable tools.

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks
Zhi Cao*, Howard Ji*, Kevin Zhang*, Kuangzhi Ge, Li Fei-Fei, Jiajun Wu†, Huang Huang† *Equal contribution; †Equal advising
arXiv 2026 · Website · Paper

TrAct uses visual tracks as a shared interface between robot control and video prediction. It proposes action-track pairs, predicts their visual outcomes, and selects actions with a vision-language reward model.

Publication overview Cross-Embodiment Robot Foundation World Models with Latent Actions
Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Tsung-Yen Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier
ICML 2026 · Website · Paper

We learn a shared latent action space for robot world models, enabling transfer across embodiments and improving downstream dexterous manipulation and other robot learning tasks.

Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
Justin Yu, Letian Fu, Huang Huang, Karim El-Refai, Rares Andrei Ambrus, Richard Cheng, Muhammad Zubair Irshad, Ken Goldberg
CoRL 2025 · Oral Presentation · Website · Paper

We generate robot training data from smartphone scans and human demonstration videos by reconstructing object geometry and motion, then rendering robot demonstrations.

OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang Huang*, Fangchen Liu*, Letian Fu*, Tingfan Wu, Mustafa Mukadam, Jitendra Malik, Ken Goldberg, Pieter Abbeel *Equal contribution
ICML 2025 · Website · Paper

OTTER extracts task-relevant visual features aligned with language instructions from frozen vision-language encoders, improving generalization to unseen robot manipulation tasks.

Publication overview DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning
Huang Huang, Balakumar Sundaralingam, Arsalan Mousavian, Adithyavairavan Murali, Ken Goldberg, Dieter Fox
CoRL 2024 · Website · Paper

We use diffusion models to generate diverse trajectory seeds for motion optimization, enabling rapid robot motion planning.

In-Context Imitation Learning via Next-Token Prediction
Letian Fu*, Huang Huang*, Gaurav Datta*, Lawrence Yunliang Chen, William Chung-Ho Panitch, Fangchen Liu, Hui Li, Ken Goldberg *Equal contribution
ICRA 2025, Website

We explore how to enhance next-token prediction models to perform in-context imitation learning on a real robot. We propose In-Context Robot Transformer (ICRT), a causal transformer that generalizes to unseen tasks conditioned on prompts of sensorimotor trajectories of the new task composing of image observations, actions and states tuples, collected through human teleoperation.
Touch, vision, and language dataset overview
A touch, vision, and language dataset for multimodal alignment
Letian Fu, Gaurav Datta*, Huang Huang*, William Chung-Ho Panitch*, Jaimyn Drake*, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg *Equal contribution
ICML 2024, Oral Presentation, Website

We introduce the Touch-Vision-Language (TVL) dataset, which combines paired tactile and visual observations with both human-annotated and VLM-generated tactile-semantic labels. We then leverage a contrastive learning approach to train a CLIP-aligned tactile encoder and finetune an open-source LLM for a tactile description task. Our results show that incorporating tactile information allows us to significantly outperform state-of-the-art VLMs (including the label generating model) on a tactile understanding task.
Manipulator as a Tail: Promoting Dynamic Stability for Legged Locomotion
Huang Huang, Antonio Loquercio, Ashish Kumar, Neerja Thakkar, Ken Goldberg, Jitendra Malik
ICRA 2024, Website

For locomotion, is an arm on a legged robot a liability or an asset for locomotion? Biological systems evolved additional limbs beyond legs that facilitates postural control. This work shows how a manipulator can be an asset for legged locomotion at high speeds or under external perturbations, where the arm serves beyond manipulation.
Learning Self-Supervised Representations from Vision and Touch for Active Sliding Perception of Deformable Surfaces
Justin Kerr*, Huang Huang*, Albert Wilcox, Ryan Hoque, Jeffrey Ichnowski, Roberto Calandra, and Ken Goldberg, *Equal contribution
RSS 2023, Paper

We learn self-supervised representations from vision and touch using contrastive learning. The representations support active perception of deformable surfaces without task-specific fine-tuning.
Evo-NeRF: Evolving NeRF for Sequential Robot Grasping
Justin Kerr, Letian Fu, Huang Huang, Yahav Avigal, Matthew Tancik, Jeffrey Ichnowski, Angjoo Kanazawa, Ken Goldberg
CoRL 2022, Oral Presentation, OpenReview

We propose Evo-NeRF with additional geometry regularizations improving performance in rapid capture settings to achieve real-time, updateable scene reconstruction for rapidly grasping table-top transparent objects. We train a NeRF-adapted grasping network that learns to ignore reconstruction artifacts.
Real2Sim2Real: Self-Supervised Learning of Physical Single-Step Dynamic Actions for Planar Robot Casting
Vincent Lim*, Huang Huang*, Lawrence Yunliang Chen, Jonathan Wang, Jeffrey Ichnowski, Daniel Seita, Michael Laskey, Ken Goldberg, *Equal contribution
ICRA 2022, paper

We collect planar robot casting data in real in a self-supervised way to tune the simulation in Isaac Gym. We then collect more data in the tuned simulator. Combined with upsampled real data, we learn a policy for planar robot casting to reach to a given target, attaining median error distance (as % of cable length) ranging from 8% to 14%.