ULOHA:
An Underwater Bimanual Robot System for Robot Learning

The University of Osaka / Kobe University, Japan

Masato Kobayashi*, Takeru Tsunoori*
* Co-first authors equally contributed to this work.
ULOHA has been featured on IEEE Spectrum Video Friday, "Your weekly selection of awesome robot videos"! (IEEE Spectrum Video Friday)
ULOHA platform, teleoperation, and autonomous underwater manipulation shown side by side.
From demonstrations to autonomous underwater skills.Watch on YouTube

Team

As of September 16, 2026

Masato KobayashiAssistant Professor · Project tech leadSoftware, hardware, experiments, paper, video editing, and website development.

Takeru TsunooriMaster’s student · 2nd yearHardware development, experiments, and paper.

ULOHA team

Overview

Learning coordinated manipulation underwater

ULOHA connects custom underwater arms with a LeRobot-based workflow for teleoperation, demonstration recording, and autonomous execution. Operators use two dry leader arms to control two underwater followers and collect demonstrations of object transfer, shared-object manipulation, and interception of buoyant objects.

  • One integrated platformCustom hardware, three camera views, and a shared learning workflow.
  • Nine bimanual tasksTask-specific ACT policies, with Diffusion Policy and SmolVLA also deployed on selected tasks.
  • Underwater evaluationBuoyancy, bubbles, action timing, and transfer between air and water.
Project overview 2:59
A complete walkthrough of the platform, demonstrations, learned behaviors, and deployment experiments. Watch on YouTube

Videos include audio. Playback speeds are shown in the footage and preserved from the source edit. Short excerpts appear alongside the corresponding experiments below.

Platform overview figure
ULOHA leader–follower system, three camera views, teleoperated lifting, and learned underwater manipulation
ULOHA — Underwater Bimanual Robot Learning

ULOHA System

ULOHA · Teleoperation · Autonomous 0:03
One platform. From demonstrations to autonomous underwater skills. Watch on YouTube

Custom Leader–Follower Hardware

ULOHA hardware arrangement and leader and follower arm designs with joints J1–J7
Custom-Designed ULOHA Hardware

Each arm has six arm joints and an actuated gripper. The two dry leaders use DYNAMIXEL XM430-W350 servos, while the underwater followers use waterproof XW430-T333 servos. Custom 3D-printed structures and grippers retain corresponding joint layouts across the leaders and followers. The two followers together provide a 14-dimensional joint state and action space, ordered left then right.

Tank enclosure and follower structures providing clearance for individual waterproof cable connectors
Experimental Setup and Waterproof Connector Accommodation

Waterproof cabling is a central design constraint: each follower motor requires an individual waterproof cable. The follower structures provide connector clearance at J6 and mechanically limit motion at J4 to prevent interference. These adaptations preserve joint correspondence while accommodating the followers’ different travel limits.

Teleoperation, Recording, and Policy Deployment

How are demonstrations of coordinated underwater motion collected?

Learning from teleoperation 0:28
An operator demonstrates bimanual lifting, insertion, and surface hand-over using the dry leader arms. These are teleoperated demonstrations, not autonomous rollouts. Watch on YouTube

Workflow Two dry leader arms control the underwater followers while joint states and three camera views are recorded at 30 Hz.

Demonstrations use unilateral joint-position teleoperation, with leader torque disabled. LeRobot calibration and normalization map the independently calibrated leader and follower joint ranges. At 30 Hz, the system records leader targets, follower states, and three camera streams: the left hand camera, top camera, and right hand camera. Language task descriptions are stored with demonstrations for compatible policies such as SmolVLA.

During autonomous execution, the policy receives the camera images and follower state and predicts a sequence of joint targets. In synchronous execution, a configurable action prefix is executed before the policy is queried again; the SmolVLA RTC experiment below instead overlaps inference with execution. ULOHA-specific hardware adapters and recording, training, and deployment configurations connect this workflow on the same platform.

Underwater Bimanual Learning

Which coordinated skills can the robot learn from demonstrations?

Autonomous bimanual manipulation 0:57
Selected ACT rollouts: lifting, hand-over, placement, stacking, sequential transfer, cooperative insertion, and lid opening. Watch on YouTube

Finding ACT achieves 68 successes in 90 clear-water trials across nine separately trained task policies.

Nine-task overview figure
Underwater execution sequences for nine bimanual manipulation tasks
Nine Underwater Bimanual Tasks

Each task uses a separately trained policy. ACT is evaluated in clear water over ten trials per task, with a 60-second limit and a baseline prediction/execution horizon of 100/100 actions. The first seven tasks use ten demonstrations each; surface hand-over and release and catch use 50 each. Success requires reaching the task-specific state within the time limit and retaining it at the end of the episode.

TaskCoordinationDemonstrationsACT success
Block hand-overPass a block from the right arm to the left without floor contact1010/10
Hand-over then placeTransfer without floor contact, then place the block fully inside the target region at rest106/10
Block stackingEach arm handles one block; the white block rests stably on the black block109/10
Sequential transferRight arm completes the first leg; left arm completes the second after an intermediate placement1010/10
Bimanual liftingJointly grasp and hold a block clear of the floor1010/10
Cooperative insertionHold a cup upright and release a block inside it109/10
Lid openingHold the container and remove its lid107/10
Surface hand-overGrasp a floating sponge and transfer it to the other arm504/10
Release and catchRelease a submerged sponge and intercept it before it surfaces503/10

Block hand-over is also evaluated with Diffusion Policy and SmolVLA using the same ten demonstrations and robot initial pose. All three methods achieve 10/10 on this task, demonstrating their deployment on ULOHA.

Buoyancy-Driven Manipulation

Can the robot grasp and transfer objects that rise or float?

Surface hand-over and release and catch 0:15
Autonomous sponge manipulation: transfer from the water surface, followed by interception of a rising sponge. Watch on YouTube

Finding In the baseline ACT evaluation, surface hand-over succeeds in 4/10 trials and release and catch in 3/10.

Right hand camera views during surface hand-over and release-and-catch demonstrations
Demonstrations of Buoyant Sponge Manipulation

Surface hand-over requires grasping a sponge that can move horizontally on the water surface; the right hand camera also crosses from water into air during approach. Release and catch requires the right arm to intercept a sponge rising after the left arm releases it. Both tasks require a nearly fully open gripper, leaving little positional clearance.

Failure Examples

Where do the learned behaviors break down?

Failure examples 0:13
Unsuccessful stacking, surface hand-over, and release-and-catch examples from the source video. Watch on YouTube

Finding The examples show loss of stack stability, missed surface grasps, and a rising sponge escaping the gripper.

Sponge-grasp failure figure
Failed sponge grasps: the sponge remains at the surface or escapes upward after closure
Examples of Unsuccessful Sponge Grasps

Underwater Deployment Sensitivities

Bubble Disturbances

How do clear-water policies perform when bubbles disturb the scene?

Manipulation under bubble disturbances 0:18
Transfer success and failure examples, followed by bimanual lifting. The policies are reused without retraining. Watch on YouTube

Finding Without retraining, sequential transfer drops from 10/10 to 3/10; bimanual lifting remains at 10/10.

Bubble-disturbance figure
Bubble streams during sequential transfer and bimanual lifting, with external and left hand camera views
Evaluation under Bubble Disturbances

Two aerators introduce bubble streams into the workspace. Policies trained in clear water are evaluated with the same ACT checkpoints, robot initial pose, and success criteria, without retraining. The clear-water results reuse the nine-task baseline trials.

TaskClear waterBubbles
Sequential transfer10/103/10
Bimanual lifting10/1010/10

Sensitivity differs between the two tasks. Bubbles are visible in both hand-camera examples, but these observations do not isolate the contribution of visual occlusion to the success rates.

ACT Execution Horizon

Does more frequent re-planning improve interception?

On release and catch, one ACT checkpoint predicts 100 actions per query. We vary how many of those actions are executed before re-planning, discarding the unused suffix without retraining.

Executed actionsNominal interval at 30 HzSuccess
100 (baseline)3.33 s3/10
501.67 s4/10
301.00 s6/10
150.50 s3/10

The best observed result is 6/10 at 30 executed actions. Shortening the horizon further returns success to 3/10, so more frequent re-planning alone does not guarantee better interception for this checkpoint. The 100-action result reuses the nine-task baseline trials.

SmolVLA with Real-Time Chunking

Does overlapping inference and execution help with a rising object?

SmolVLA: synchronous execution and RTC 0:08
Side-by-side execution examples. The displayed 4/10 and 6/10 figures are the ten-trial evaluation results, not a claim of statistical superiority. Watch on YouTube

Finding RTC achieves 6/10 successes versus 4/10 for synchronous execution using the same SmolVLA checkpoint. Ten trials per condition do not establish statistical superiority.

Real-time chunking (RTC) generates a new action chunk while the previous one is executing, conditioning on its remaining plan and accounting for inference delay. The synchronous and RTC conditions use the same SmolVLA checkpoint trained on the same 50 release-and-catch demonstrations used for ACT.

SmolVLA execution strategySuccess
Synchronous execution4/10
Real-time chunking (RTC)6/10

RTC supports deployment on this buoyancy-driven interception task. Its observed success rate matches ACT with a 30-step execution horizon; ten trials per condition do not establish a ranking of the methods.

Single-Arm Air–Water Transfer

Does a policy trained in one medium transfer to the other?

Learning across air and water 0:15
Single-medium and mixed-medium training, including the unseen-object evaluation. Watch on YouTube

Finding Single-medium training gives 0/10 success in the other medium. With five demonstrations in each medium, one policy achieves 10/10 in both.

Air-water experimental setup figure
Single-arm pick-and-place in air and underwater from the hand camera and an external viewpoint
Single-Arm Pick-and-Place in Air and Water

A separate ACT study uses the right follower for pick-and-place, keeping the task, block, and camera placement identical while filling or emptying the tank. Success requires the block to be fully inside the target region at rest. Each policy uses ten demonstrations in total, with ten evaluation trials per condition and a 45-second limit. Single-arm demonstrations use two cameras.

Training demonstrationsTested underwaterTested in air
10 underwater10/100/10
10 in air0/1010/10
5 underwater + 5 in air10/1010/10

For this task and the tested conditions, single-medium policies succeed in their training medium but fail in the other. One mixed-medium policy succeeds in both air and water without increasing the total demonstration count. Filling the tank changes perception and physical interactions together, so this comparison measures their combined effect.

The mixed-medium policy also achieves 7/10 underwater on an unseen black rubber block. Because both appearance and material properties differ from the white polyurethane training block, this result does not isolate color generalization.

Discussion and Scope

ULOHA provides an integrated experimental platform for studying underwater bimanual robot learning with established policy models. The experiments demonstrate coordinated transfers, shared-object handling, and buoyancy-driven interception, while revealing sensitivity to bubbles, execution timing, and training-medium coverage. Results apply to the tested tank setup and checkpoints; broader evaluation is needed to assess generalization and separate visual changes from hydrodynamic effects.

The paper plans an open-source release of the hardware designs and software.

ULOHA

Bimanual learning, underwater.

ULOHA — Closing film 0:12
Coordinated underwater manipulation, from demonstrations to learned behavior. Watch on YouTube

Citation

@misc{kobayashi2026uloha,
      title={ULOHA: An Underwater Bimanual Robot System for Robot Learning},
      author={Masato Kobayashi and Takeru Tsunoori},
      year={2026},
      eprint={2609.19200},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.19200},
}

Contact

Masato Kobayashi
Corresponding author · Assistant Professor
The University of Osaka / Kobe University, Japan