MULTI-HAND DEXTEROUS MANIPULATION

DexJoCo-X

Benchmarking Action Representations
for Multi-Hand Dexterous Manipulation

Seven different hands. Shared manipulation tasks.
A common foundation for studying how dexterous skills transfer across morphologies.

7dexterous hands
6manipulation tasks
2,100demonstrations
42hand–task pairs
A shared benchmark connecting diverse hands, balanced demonstrations, and consistent evaluation.
01

THE RESEARCH QUESTION

How do we learn across different hands?

Robot hands differ in their kinematics, actuation, and control. DexJoCo-X makes it possible to compare action representations using the same tasks, demonstrations, scenes, and success criteria.

Our study examines how action coordinates interact with pretraining and model architecture. We compare Native coordinates, function-aligned slots (FAAS), and learned cross-hand latents (DexLatent), alongside per-hand π0.5 (Ego-Pi) and joint seven-hand Being-H0.5 policies.

A matched benchmark

Common scenes, success rules, and execution interfaces across all 42 hand–task pairs.

A balanced dataset

50 demonstrations per pair, with synchronized RGB, robot states, actions, and masks.

Reusable evaluation

Connect your own policy to the same simulator and compare representations under consistent conditions.

AllegroInspireLEAPLinkerHandSharpa WaveWujiXHand
02

SEE IT IN ACTION

One task, seven embodiments.

Watch the benchmark overview and recorded simulation demonstrations across all seven hands.

Recorded demonstrations in simulation · 2 min 40 sec

03

SIX COMPLEMENTARY TASKS

From tool use to bimanual coordination.

SINGLE ARM

Bucket lifting · Nail hammering · Tower of Hanoi

BIMANUAL

Microwave cooking · iPad unlocking · Photography

Task-specific success rules cover grasp stability, tool use, sequencing, and coordinated manipulation.
Consistent evaluation

Three independent runs of 50 rollouts per hand–task pair. Every pair receives equal weight in the full-grid macro-average.

04

DEMONSTRATION COLLECTION

From human motion to balanced data.

Seven glove-to-hand mappings support demonstration recording. Automated scene expansion and verification produce the balanced training dataset.

Reviewed source demonstrations are expanded, replayed, verified, and exported with synchronized observations and commands.
05

ACTION REPRESENTATIONS

Shared skills, morphology-specific control.

Native, FAAS, and DexLatent share arm control and execution conditions. Their hand-action representations are decoded into each embodiment's native commands.

Representation, pretraining, and architecture jointly shape multi-hand learning.
06

BUILD ON DEXJOCO-X

Paper, data, and evaluation.

READ

The paper

The benchmark design, action representations, and multi-hand learning study.

Read PDF
TRAIN

The dataset

2,100 demonstrations in six task packages, with native tensors, RGB videos, and checksums.

Release in preparation
EVALUATE

The code

Simulation environments, evaluation interfaces, data conversion, and a π0.5 inference example.

Release in preparation