Every bimanual data program starts with the same question: which rig? The GELLO vs ALOHA comparison comes up first because both are open, both are documented, and both avoid the retargeting problem that makes joystick and motion-capture data hard to use.
They solve different problems. This guide covers what each system is, how they differ in practice, what actually changes in the data you get, and five questions that will settle the choice for your team.
The two systems in plain terms
GELLO
GELLO is a low-cost puppeteering device. It is a small, largely 3D-printed leader arm built from hobby servos, shaped as a scaled replica of the robot arm you already own. You move the little arm, the real arm follows. Its defining feature is adaptability: it is designed to be fitted to commercial arms you already have rather than shipping as a complete robot system.
ALOHA
ALOHA is a complete bimanual cell. Two leader arms, two follower arms, a fixed camera set, and a documented recording stack. It is a system rather than an accessory, and it assumes you are buying the follower arms as part of the build.
The distinction sounds academic. It drives almost every practical difference below.
Head to head
| Dimension | GELLO | ALOHA |
|---|---|---|
| What you buy | A leader device for arms you own | A full four-arm bimanual cell |
| Relative hardware cost | Lowest in class | Substantially higher |
| Arms supported | Adapts to several commercial arms | Built around a specific arm family |
| Bimanual out of the box | Requires two units and coordination | Yes, by design |
| Build effort | Printing, assembly, calibration | Mostly integration and tuning |
| Operator learning curve | Short | Short |
| Best at | Single-arm and exploratory work | Sustained bimanual production capture |
| Weakest at | High-duty continuous operation | Cost of a first experiment |
What actually changes in the data
Both produce joint-space action labels at high rate with no retargeting step, which is the main reason either beats a joystick. The differences show up at the edges.
- Bimanual coordination. Two-handed tasks require the two arms to be recorded on one clock with one operator’s intent behind both. A purpose-built bimanual cell gives you this by construction; two independent leader devices need careful sync work.
- Force fidelity. Lighter servo-based leaders give the operator less resistance feedback. On contact-rich tasks that means more crushed or dropped objects unless you instrument the follower properly with force-torque sensing.
- Session length. A lighter rig is more comfortable for short sessions and less stable across a full shift. Fatigue effects differ, which changes your session design.
- Camera consistency. A documented cell has a documented camera layout. A retrofit leaves camera placement to you, which is fine until you compare batches collected three months apart.
Neither system removes the need for structured recording. Both need the metadata set described in our guide to reading a teleoperation demonstration log.
Five questions that decide it for you
- Do you already own robot arms? If yes, an adaptable leader device is the cheaper path. If no, a complete cell removes integration risk.
- Is the target task genuinely bimanual? Folding, opening containers, and tool handovers are. Pick-and-place usually is not. Our bimanual arms page covers the distinction.
- Are you exploring or producing? Exploration rewards cheap and flexible. Production rewards documented and repeatable.
- What is your weekly trajectory target? Above a few hundred sustained per week, build quality and duty cycle matter more than purchase price.
- Who operates it? A research team tolerates a fiddly rig. A dedicated operations team running shifts needs something that survives daily handling.
The middle path most teams end up taking
The common pattern is to start cheap and standardize later. Teams prototype with a low-cost leader on arms they already own, use it to discover which tasks actually matter and where the failure modes sit, then commit to a purpose-built bimanual cell once the task list stops changing every week.
The trap is failing to plan the transition. Data collected on the exploratory rig often cannot be pooled with production data because camera layouts, control rates, and log schemas differ. Fix the schema before the first episode, not after the thousandth. If you are outsourcing, a hybrid delivery model lets you run exploration and production capture in parallel without duplicating the pipeline.
Rig choice matters far less than what happens around it. Our humanoid foundation model case study covers a programme that reached 3x sample efficiency.
Frequently asked questions
Can I use two GELLO units for bimanual capture?
Yes, and teams do. Budget engineering time for clock sync and for a control loop that treats both arms as one system. Otherwise you get two single-arm datasets recorded at the same time, which is not the same thing.
Which produces better data for imitation learning?
Neither has an inherent advantage on data quality. The rig determines what is practical to collect; operator selection, session design, and QA determine whether it is any good. See our operator quality guide.
Do these rigs work with humanoids?
Leader-follower control maps cleanly to a humanoid’s arms but not to its legs, torso, or head. Whole-body capture generally moves to headset-driven control, covered in VR headset teleoperation for everyday tasks.
Is open-source hardware a compliance problem for enterprise work?
Not usually. Questions concentrate on data handling, operator vetting, and residency rather than rig provenance. Our compliance page covers what enterprise buyers actually audit.
The rig is the cheapest decision in a teleoperation program and the one teams agonize over longest. Operator hours, scene resets, and QA will dominate your budget regardless of which you choose. If you want the operating plan sized before you commit to hardware, tell us what you are building.
Related reading
- ALOHA teleoperation data collection: setup, costs, mistakes
- Teleoperation data collection
- Teleoperation latency budgets explained
- Build in-house vs buy outsourced
- Bimanual arms platform
- Case study: Humanoid foundation model: 3x sample efficiency
External reference

Manish Jain · Chief Marketing Officer
Manish Jain is Chief Marketing Officer at Roborax, bringing over 20 years of experience in business strategy, digital transformation, and growth leadership to help enterprises build scalable, high-quality AI data operations.





