OmniDex Scales Dexterous Hand Grasping to 2.6 Million Cluttered Scenes
On October 8, a paper titled “OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes” was uploaded to arXiv. Its authors are Naiyu Fang, Zhongjin Luo, Yuxin Mo, Siyuan Huang, Jianbo Liu, Yufei Liu, Zheyuan Zhou, Chenkai Jin, Xiaogang Wang, and Hongsheng Li. The work is listed under Computer Science > Robotics.
The paper addresses dexterous grasping, a foundational primitive for embodied AI. Training robust grasp models requires massive amounts of data, yet collecting real-world demonstrations is expensive and slow. Simulation has therefore become the dominant paradigm for scalable data generation.
However, cluttered scenes are the setting that best reflects real-world applications, and learning grasping in such scenes remains bottlenecked by a critical scarcity of large-scale data. OmniDex targets this gap by scaling dexterous hand grasping to 2.6 million cluttered scenes.
The reported approach expands grasp learning from isolated objects or simple tabletop setups to diverse, cluttered configurations. By generating large-scale simulated scenes with varied object arrangements, the framework aims to expose dexterous hands to the kind of spatial complexity and occlusion found in practical environments.
Because cluttered scenes combine many objects, occlusions, and contact constraints, they are especially difficult for data-driven grasp synthesis. OmniDex’s scale is meant to make such cases a first-class training target rather than a rare long-tail condition.
The resulting scale is intended to support training of more robust grasp policies that can generalize across messy scenes. For embodied AI, such a resource could help reduce reliance on costly real-world data collection and improve sim-to-real transfer.
OmniDex is positioned as a step toward general-purpose dexterous manipulation: a large-scale benchmark and data source for studying grasping under clutter. The authors’ contribution highlights both the promise of simulation and the persistent challenge of building diverse, realistic cluttered-scene datasets at scale.