COWBOY A Scalable Sim-to-Real Framework for Learning Contextual Whole-Body Manipulation

Jiahui Yang*, Jason Jingzhou Liu*, Muxin Liu*, Andrew E. Wang, Kenneth Shaw, Deepak Pathak†, Ruslan Salakhutdinov†

*Equal contribution†Equal advising

Conference on Robot Learning (CoRL) 2026

Abstract

Mobile manipulators must adapt not only to objects, but also to the spatial context in which they appear: where to position the base, how to manipulate, and where to look while navigating surrounding scene geometry. To address this, we present COWBOY, a scalable sim-to-real framework for learning such contextual whole-body manipulation. On whole-body dexterous grasping, COWBOY transfers zero-shot to hundreds of objects and diverse in-the-wild scenes, spanning cluttered tabletops, shelves, drawers, and other unstructured settings.

Rather than learning scene adaptation directly with RL in simulation, a vectorized environment-aware whole-body controller (WBC) handles reaching in thousands of scenes by coordinating base-arm motion, collision avoidance, and active vision. Then, local RL experts are trained in local scenes for object interaction and can be reused across thousands of full-scene layouts. We distill these composed WBC--RL teachers into a single vision policy that controls the mobile base, arm, dexterous hand, and actuated camera from point-cloud observations. The resulting policy exhibits contextual whole-body behavior, automatically adapting base placement, camera viewpoint, arm trajectory, and grasp strategy to observed scene and object geometry.

All videos below show the same policy in different scenarios (1x Speed)

COWBOY enables sim-to-real mobile manipulation across diverse objects and environments

A single vision-based policy coordinates all 32 degrees of freedom:

  • Mobile base:3 DOF
  • Arm:7 DOF
  • Dexterous hand:16 DOF
  • Actuated neck for active vision:6 DOF

The policy enables the robot to grasp hundreds of objects across diverse in-the-wild scenes.

Click to See Full Evaluations

COWBOY policies exhibit contextual manipulation

Observing only the scene geometry, object geometry, and target object, the policy automatically adapts its behavior accordingly:

  • Selects the appropriate grasp and approach strategy
  • Avoids collisions with the surrounding scene
  • Maintains gaze on the target object through active neck control

All without explicit conditioning on scene type, base pose, camera motion, or grasp strategy.

Motion depends on Object Geometry

Same Scene, Different Object → Different Grasps

Motion also depends on Environment Geometry

Same Object, Different Scenes → Different Grasps

COWBOY policies coordinate the whole body

Active Vision

Policy controls an actuated neck to maintain gaze on the target object.

Base Motion

Policy coordinates base movement if the object is out of reach.

COWBOY integrates with SAM3 for language-conditioned manipulation

SAM3 segments the prompted object, and the estimated object position guides the same policy to accomplish the task

Green Pear, Grape, Apple
Tennis Ball, Pumpkin

Real World Evaluations

We evaluate our whole-body grasping policy in 24 in-the-wild scenes across over 100 unique objects, achieving a 91% success rate.

Hover to play video

Scene 01
Scene 01
Scene 02
Scene 02
Scene 03
Scene 03
Scene 04
Scene 04
Scene 05
Scene 05
Scene 06
Scene 06
Scene 07
Scene 07
Scene 08
Scene 08
Scene 09
Scene 09
Scene 10
Scene 10
Scene 11
Scene 11
Scene 12
Scene 12
Scene 13
Scene 13
Scene 14
Scene 14
Scene 15
Scene 15
Scene 16
Scene 16
Scene 17
Scene 17
Scene 18
Scene 18
Scene 19
Scene 19
Scene 20
Scene 20
Scene 21
Scene 21
Scene 22
Scene 22
Scene 23
Scene 23
Scene 24
Scene 24

COWBOY framework has three components:

1. Vectorized Environment-Aware Whole-Body Controller

We introduce a vectorized whole-body controller (WBC) capable of generating environment aware whole-body motion across thousands of simulated environments on a single GPU.

End-Effector Reaching

Camera Gaze Control

Self Collision Avoidance

Environment Collision Avoidance

2. Contact-Rich Local RL Policies

Our WBC handles collision-free approach to the target object. This allows us to train local RL policies that only focus on object interactions, simplifying the learning task.

Top Down Grasp

Side Grasp

Constrained Side Grasp

3. Local-to-Full-Scenes Multi-Teacher Distillation

We embed the local scenes within diverse full scenes, where the WBC first brings the end effector near the object. The local policy can then complete the task, enabling a policy trained in one local scene to generate expert motion across diverse full-scenes.

Local scenes embedded within diverse full scenes for multi-teacher distillation

We distill both the WBC reaching behavior and the state-based local policies into a unified vision-based whole-body policy, trained across thousands of environments and hundreds of objects. The resulting policy zero-shot transfers to the real world.

Local-to-full-scenes multi-teacher distillation overview

Acknowledgements

We thank Hengkai Pan, Tony Tao, Peiqi Liu, Ritvik Singh, Arthur Allshire, Tal Daniel, and Dieter Fox for valuable discussions and feedback on this work. We also thank Adam Kan for help setting up the TidyBot++ base. This work was supported in part by XXX.

BibTeX

@inproceedings{yang2026cowboy,
  title     = {COWBOY: A Scalable Sim-to-Real Framework for Learning Contextual Whole-Body Manipulation},
  author    = {Yang, Jiahui and Liu, Jason Jingzhou and Liu, Muxin and
               Wang, Andrew E. and Shaw, Kenneth and
               Pathak, Deepak and Salakhutdinov, Ruslan},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}