ULOHA:
An Underwater Bimanual Robot System for Robot Learning
The University of Osaka / Kobe University, Japan
Team
As of September 16, 2026
Masato KobayashiAssistant Professor · Project tech leadSoftware, hardware, experiments, paper, video editing, and website development.
Takeru TsunooriMaster’s student · 2nd yearHardware development, experiments, and paper.

Overview
Learning coordinated manipulation underwater
ULOHA connects custom underwater arms with a LeRobot-based workflow for teleoperation, demonstration recording, and autonomous execution. Operators use two dry leader arms to control two underwater followers and collect demonstrations of object transfer, shared-object manipulation, and interception of buoyant objects.
- One integrated platformCustom hardware, three camera views, and a shared learning workflow.
- Nine bimanual tasksTask-specific ACT policies, with Diffusion Policy and SmolVLA also deployed on selected tasks.
- Underwater evaluationBuoyancy, bubbles, action timing, and transfer between air and water.
Videos include audio. Playback speeds are shown in the footage and preserved from the source edit. Short excerpts appear alongside the corresponding experiments below.
ULOHA System
Custom Leader–Follower Hardware

Each arm has six arm joints and an actuated gripper. The two dry leaders use DYNAMIXEL XM430-W350 servos, while the underwater followers use waterproof XW430-T333 servos. Custom 3D-printed structures and grippers retain corresponding joint layouts across the leaders and followers. The two followers together provide a 14-dimensional joint state and action space, ordered left then right.

Waterproof cabling is a central design constraint: each follower motor requires an individual waterproof cable. The follower structures provide connector clearance at J6 and mechanically limit motion at J4 to prevent interference. These adaptations preserve joint correspondence while accommodating the followers’ different travel limits.
Teleoperation, Recording, and Policy Deployment
How are demonstrations of coordinated underwater motion collected?
Workflow Two dry leader arms control the underwater followers while joint states and three camera views are recorded at 30 Hz.
Demonstrations use unilateral joint-position teleoperation, with leader torque disabled. LeRobot calibration and normalization map the independently calibrated leader and follower joint ranges. At 30 Hz, the system records leader targets, follower states, and three camera streams: the left hand camera, top camera, and right hand camera. Language task descriptions are stored with demonstrations for compatible policies such as SmolVLA.
During autonomous execution, the policy receives the camera images and follower state and predicts a sequence of joint targets. In synchronous execution, a configurable action prefix is executed before the policy is queried again; the SmolVLA RTC experiment below instead overlaps inference with execution. ULOHA-specific hardware adapters and recording, training, and deployment configurations connect this workflow on the same platform.
Underwater Bimanual Learning
Which coordinated skills can the robot learn from demonstrations?
Finding ACT achieves 68 successes in 90 clear-water trials across nine separately trained task policies.
Each task uses a separately trained policy. ACT is evaluated in clear water over ten trials per task, with a 60-second limit and a baseline prediction/execution horizon of 100/100 actions. The first seven tasks use ten demonstrations each; surface hand-over and release and catch use 50 each. Success requires reaching the task-specific state within the time limit and retaining it at the end of the episode.
| Task | Coordination | Demonstrations | ACT success |
|---|---|---|---|
| Block hand-over | Pass a block from the right arm to the left without floor contact | 10 | 10/10 |
| Hand-over then place | Transfer without floor contact, then place the block fully inside the target region at rest | 10 | 6/10 |
| Block stacking | Each arm handles one block; the white block rests stably on the black block | 10 | 9/10 |
| Sequential transfer | Right arm completes the first leg; left arm completes the second after an intermediate placement | 10 | 10/10 |
| Bimanual lifting | Jointly grasp and hold a block clear of the floor | 10 | 10/10 |
| Cooperative insertion | Hold a cup upright and release a block inside it | 10 | 9/10 |
| Lid opening | Hold the container and remove its lid | 10 | 7/10 |
| Surface hand-over | Grasp a floating sponge and transfer it to the other arm | 50 | 4/10 |
| Release and catch | Release a submerged sponge and intercept it before it surfaces | 50 | 3/10 |
Block hand-over is also evaluated with Diffusion Policy and SmolVLA using the same ten demonstrations and robot initial pose. All three methods achieve 10/10 on this task, demonstrating their deployment on ULOHA.
Buoyancy-Driven Manipulation
Can the robot grasp and transfer objects that rise or float?
Finding In the baseline ACT evaluation, surface hand-over succeeds in 4/10 trials and release and catch in 3/10.

Surface hand-over requires grasping a sponge that can move horizontally on the water surface; the right hand camera also crosses from water into air during approach. Release and catch requires the right arm to intercept a sponge rising after the left arm releases it. Both tasks require a nearly fully open gripper, leaving little positional clearance.
Failure Examples
Where do the learned behaviors break down?
Finding The examples show loss of stack stability, missed surface grasps, and a rising sponge escaping the gripper.
Underwater Deployment Sensitivities
Bubble Disturbances
How do clear-water policies perform when bubbles disturb the scene?
Finding Without retraining, sequential transfer drops from 10/10 to 3/10; bimanual lifting remains at 10/10.
Two aerators introduce bubble streams into the workspace. Policies trained in clear water are evaluated with the same ACT checkpoints, robot initial pose, and success criteria, without retraining. The clear-water results reuse the nine-task baseline trials.
| Task | Clear water | Bubbles |
|---|---|---|
| Sequential transfer | 10/10 | 3/10 |
| Bimanual lifting | 10/10 | 10/10 |
Sensitivity differs between the two tasks. Bubbles are visible in both hand-camera examples, but these observations do not isolate the contribution of visual occlusion to the success rates.
ACT Execution Horizon
Does more frequent re-planning improve interception?
On release and catch, one ACT checkpoint predicts 100 actions per query. We vary how many of those actions are executed before re-planning, discarding the unused suffix without retraining.
| Executed actions | Nominal interval at 30 Hz | Success |
|---|---|---|
| 100 (baseline) | 3.33 s | 3/10 |
| 50 | 1.67 s | 4/10 |
| 30 | 1.00 s | 6/10 |
| 15 | 0.50 s | 3/10 |
The best observed result is 6/10 at 30 executed actions. Shortening the horizon further returns success to 3/10, so more frequent re-planning alone does not guarantee better interception for this checkpoint. The 100-action result reuses the nine-task baseline trials.
SmolVLA with Real-Time Chunking
Does overlapping inference and execution help with a rising object?
Finding RTC achieves 6/10 successes versus 4/10 for synchronous execution using the same SmolVLA checkpoint. Ten trials per condition do not establish statistical superiority.
Real-time chunking (RTC) generates a new action chunk while the previous one is executing, conditioning on its remaining plan and accounting for inference delay. The synchronous and RTC conditions use the same SmolVLA checkpoint trained on the same 50 release-and-catch demonstrations used for ACT.
| SmolVLA execution strategy | Success |
|---|---|
| Synchronous execution | 4/10 |
| Real-time chunking (RTC) | 6/10 |
RTC supports deployment on this buoyancy-driven interception task. Its observed success rate matches ACT with a 30-step execution horizon; ten trials per condition do not establish a ranking of the methods.
Single-Arm Air–Water Transfer
Does a policy trained in one medium transfer to the other?
Finding Single-medium training gives 0/10 success in the other medium. With five demonstrations in each medium, one policy achieves 10/10 in both.
A separate ACT study uses the right follower for pick-and-place, keeping the task, block, and camera placement identical while filling or emptying the tank. Success requires the block to be fully inside the target region at rest. Each policy uses ten demonstrations in total, with ten evaluation trials per condition and a 45-second limit. Single-arm demonstrations use two cameras.
| Training demonstrations | Tested underwater | Tested in air |
|---|---|---|
| 10 underwater | 10/10 | 0/10 |
| 10 in air | 0/10 | 10/10 |
| 5 underwater + 5 in air | 10/10 | 10/10 |
For this task and the tested conditions, single-medium policies succeed in their training medium but fail in the other. One mixed-medium policy succeeds in both air and water without increasing the total demonstration count. Filling the tank changes perception and physical interactions together, so this comparison measures their combined effect.
The mixed-medium policy also achieves 7/10 underwater on an unseen black rubber block. Because both appearance and material properties differ from the white polyurethane training block, this result does not isolate color generalization.
Discussion and Scope
ULOHA provides an integrated experimental platform for studying underwater bimanual robot learning with established policy models. The experiments demonstrate coordinated transfers, shared-object handling, and buoyancy-driven interception, while revealing sensitivity to bubbles, execution timing, and training-medium coverage. Results apply to the tested tank setup and checkpoints; broader evaluation is needed to assess generalization and separate visual changes from hydrodynamic effects.
The paper plans an open-source release of the hardware designs and software.
ULOHA
Bimanual learning, underwater.
Citation
@misc{kobayashi2026uloha,
title={ULOHA: An Underwater Bimanual Robot System for Robot Learning},
author={Masato Kobayashi and Takeru Tsunoori},
year={2026},
eprint={2609.19200},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.19200},
}
Contact
Masato Kobayashi
Corresponding author · Assistant Professor
The University of Osaka / Kobe University, Japan





