Teams budget bimanual manipulation datasets by doubling their single-arm numbers. The real multiplier is higher, and the gap is not hardware. It is coordination, yield, and the specific ways two-armed collection fails that single-arm collection never does.
This guide explains why two arms cost more than twice one, distinguishes the three kinds of bimanual task, sets out what breaks first at scale, and gives practical guidance for scoping a program realistically.
Why two arms is not twice one arm
The intuition is that a bimanual dataset costs roughly double a single-arm one. Twice the hardware, twice the streams, twice the storage. In practice the multiplier is larger, and the reason is coordination.
A single-arm policy learns a mapping from observation to one action. A bimanual policy learns a mapping to two actions that must be correct relative to each other. Two arms in the right places at the wrong times is a failure, and it is a failure mode that does not exist with one arm.
Where the extra cost actually appears
| Cost area | Single arm | Bimanual |
|---|---|---|
| Hardware | Baseline | Roughly double |
| Sync requirements | Moderate | Strict; the arms must share one clock exactly |
| Camera count | Two to three | Three to five; both wrists need coverage |
| Operator skill | Learnable in days | Genuinely harder; two-handed coordination through a rig |
| Episode duration | Shorter | Longer; more sub-steps per task |
| Scene reset time | Quick | Slower; more objects, more staging |
| Failure rate | Lower | Higher; more ways to go wrong per episode |
| QA effort | Per-arm checks | Per-arm plus relative-timing checks |
The failure rate line is the one that surprises budget models. A higher share of collected episodes is unusable, so cost per usable trajectory rises faster than cost per collected trajectory.
The three kinds of bimanual task
Parallel independent
Both arms work simultaneously on unrelated sub-tasks, such as sorting into two bins. Coordination requirements are minimal. This is close to two single-arm datasets and costs accordingly.
Sequential dependent
One arm acts, then the other, with handovers between them. Passing an object from left to right is the canonical case. Timing matters at the transition points and nowhere else, so the coordination burden is concentrated and manageable.
Simultaneous coupled
Both arms act on the same object at the same time, and the forces interact. Folding cloth, opening a jar, stretching a bag. This is where the real cost lives, because every timestep carries a coordination constraint and both arms are in continuous contact.
Scope your program by category, not by arm count. A parallel-independent task set does not justify the instrumentation a simultaneous-coupled one demands.
What breaks first at scale
- Clock drift between arms. Two controllers with independent clocks will drift. A few milliseconds is invisible in review and destroys relative-timing information.
- Asymmetric operator skill. Most operators are noticeably better with their dominant hand, and that asymmetry appears as a systematic quality difference between arms.
- Occlusion by the other arm. The second arm blocks the first far more often than any fixed obstacle, which raises the value of wrist cameras on both sides. See wrist-cam vs head-cam.
- Collision-avoidance artifacts. Safety controllers that prevent arm-to-arm collision alter trajectories. If those interventions are not logged, they appear as inexplicable motion.
- Longer episodes, more compounding. Bimanual tasks are usually long-horizon tasks too, which brings the requirements in our long-horizon capture guide.
Practical guidance for a bimanual program
- Hardware-sync the arms. Software timestamps are not sufficient when relative timing is the signal you are capturing.
- Log per-arm and relative metrics. Relative pose and relative timing deserve their own quality checks.
- Wrist cameras on both arms, always. This is the one place where cutting camera count is a false economy.
- Balance operators across handedness and track per-arm quality per operator, feeding into operator quality scoring.
- Log every safety intervention as a structured event so those trajectories can be excluded or studied deliberately.
- Budget on usable episodes, not collected ones. Apply a realistic yield assumption from your pilot rather than an optimistic one.
On rig choice, our comparison of GELLO and ALOHA covers adapting arms you own versus buying a purpose-built bimanual cell.
How to measure coordination quality
Per-arm quality metrics miss the thing that makes bimanual data valuable. An episode where both arms individually scored well on smoothness can still be a failed demonstration, because the arms were smooth at the wrong moments relative to each other.
Score coordination as its own dimension, alongside per-arm quality.
- Relative pose error at handover. How closely the receiving arm matched the presenting arm at the transfer instant. This is the single most diagnostic number for sequential-dependent tasks.
- Synchrony of contact onset. For simultaneous-coupled tasks, the time difference between the two arms making contact. Large or inconsistent gaps mean the operator is working the arms sequentially rather than together.
- Opposing force balance. When both arms hold one object, the forces should oppose cleanly. Imbalance indicates one arm is fighting the other, which is invisible on video and obvious in the traces.
- Idle ratio per arm. The share of the episode each arm spends stationary. A high idle ratio on one side usually means the task was not really bimanual, or the operator defaulted to their dominant hand.
Track these per operator as well as per episode. Coordination is a trainable operator skill and it improves measurably with feedback, which is why it belongs in operator quality scoring rather than in a post-hoc review.
What household tasks reveal
In our everyday task capture for a humanoid robotics developer building general-purpose home and service robots, most genuinely useful tasks turn out to be bimanual, and most are simultaneous-coupled rather than parallel.
Folding a towel, opening a jar, holding a bag while filling it, stabilizing a plate while wiping it. In each case one hand stabilizes while the other acts, and the stabilizing hand is doing real work that a single-arm dataset would never record.
Two consequences follow for anyone scoping a home robot program. Bimanual capture is not an advanced phase to defer, it is the baseline requirement for most of the task list. And the stabilizing role needs explicit labeling, because a policy that learns only the acting arm will drop everything it should have been holding. The taxonomy work behind this is in activities of daily living.
Deformable items are almost always two-handed work. Our warehouse policy case study tracks success on them from 61 to 84 percent.
Frequently asked questions
Can we train a bimanual policy on single-arm data?
You can pretrain useful representations that way, and you cannot learn coordination from it. Coordination is precisely the information single-arm data does not contain.
Do both arms need identical hardware?
Not necessarily, and many real applications are asymmetric, such as a gripper paired with a specialized tool. Asymmetric setups need more careful action representation because the arms have different action spaces.
How much bimanual data do we need relative to single-arm?
Expect a substantial multiple for simultaneous-coupled tasks, roughly comparable for parallel-independent ones. Scope by task category rather than assuming one ratio for everything.
Is a humanoid a bimanual platform?
For manipulation purposes largely yes, with the added complication that the torso and head also move, which changes camera geometry between episodes. See our humanoid platform page.
Bimanual programs go over budget because they are scoped as two single-arm programs. If you are planning one and want a realistic yield model before you commit, tell us the task list.






