MR-GLi: Mixed Reality-Based Gripper-Linked Overlays for Underwater Robot Arm Teleoperation via Bilateral Control
The University of Osaka / Kobe University, Japan

MR-GLi in Operation — the torque indicator and wrist-camera image follow the gripper
The operator teleoperates an underwater robot arm through bilateral control while wearing a Quest 3 headset. Inside the headset, the reaction torque indicator (RTI) and the wrist-camera image are registered to the gripper and follow the arm as it moves (inset), so the grasping torque can be read without looking away from the manipulation site. In the 2D Monitor condition, the same visual information is shown on a display beside the tank.

System Overview — MR-GLi and the 2D Monitor configuration
The submerged follower arm and the leader arm operated by the user are coupled by four-channel bilateral control at about 1 kHz, which returns haptic feedback and estimates the gripper’s reaction torque. A PC converts that torque into the RTI and combines it with the wrist-camera stream. In MR-GLi, both are placed on a panel anchored to the gripper and viewed through the headset’s video passthrough; in the 2D Monitor condition, they appear on a monitor beside the leader device. The controller, RTI encoding and camera image are identical in both conditions.
Method
Robot Platform and Bilateral Control

Leader and Follower Arms — three rotational joints (J1–J3) and a two-fingered gripper (J4), waterproof Dynamixel XW540-T260; only the follower is submerged
Four-channel leader–follower bilateral control enforces position tracking and the action–reaction law at approximately 1 kHz on all four joints. Joint torque is estimated sensorlessly with a reaction torque observer (RTOB), and the displayed quantity is the gripper-joint reaction torque — the same signal used inside the bilateral loop.
Reaction Torque Indicator

RTI Encoding — bar length for magnitude, hue for state
The RTI encodes torque with two cues: bar length for magnitude and hue for state (blue below, green inside, red above the optimal band). The optimal band is 0.20–0.40 N·m around a target of 0.30 N·m, with linear blends at the thresholds. The encoding is identical in both conditions.
Gripper-Linked Overlays and Calibration

Tracking and Calibration in MR-GLi
Panel poses come from the follower arm’s joint encoders: forward kinematics gives the gripper pose, and both panels sit at a fixed offset in the gripper frame, so the visual information follows the gripper as the arm moves. The canvas is turned toward the operator each frame so it stays legible at any gripper orientation. The robot-to-MR transform is set once per session by aligning a virtual arm model to the physical arm, then stored relative to the headset’s play-area boundary so it can be restored without repeating the alignment.
Experiments

Experimental Setup — operator beside the tank; the two configurations; the rigid block, compliant sponge and practice block

The Two Tasks — Lift (hold at a marked height for 5 s) and Pick and Place (grasp, carry, release into the basket)

Experimental Protocol under the Two Display Conditions
A counterbalanced within-subject study with 20 participants (9 female, 11 male; mean age 22.0, SD 2.7). Each condition comprised Lift and Pick&Place on a rigid polyurethane block and a compliant cellulose sponge — 160 trials in total (20 × 2 conditions × 4 task–object combinations), logged at approximately 1 kHz.
TABLE I. Comparison of the two display configurations.
| Component | 2D Monitor | MR-GLi |
|---|---|---|
| Four-channel bilateral control | ✔ | ✔ |
| RTI (identical encoding, thresholds) | ✔ | ✔ |
| Wrist-camera image of the gripper | ✔ | ✔ |
| Covarying with placement: | ||
| RTI and camera registered to the gripper | – | ✔ |
| Direct unmediated view of the tank | ✔ | passthrough |
| Gaze shift to a separate display | ✔ | – |
| Head-borne weight, narrowed FOV | – | ✔ |
Torque Regulation
TABLE II. Experimental results. *Trial time was evaluated for Pick&Place only.
| Optimal [%] | Low [%] | High [%] | MAE [N·m] | SD [N·m] | Trial time [s]* | |
|---|---|---|---|---|---|---|
| 2D Monitor | 86.36 | 6.72 | 6.92 | 0.056 | 0.057 | 16.74 |
| MR-GLi | 86.42 | 6.25 | 7.34 | 0.059 | 0.053 | 14.14 |
| Difference, MR-GLi − 2D Monitor | ||||||
| p | .546 | .498 | .475 | .409 | .409 | .053 |
| r | 0.14 | 0.16 | 0.17 | 0.19 | 0.19 | 0.43 |
No significant difference in time within the optimal torque range (p = .546, r = .14) or in any other torque measure (all p ≥ .409). Torque regulation is preserved. Trial time was shorter with MR-GLi but not significantly (p = .053); because it also depended on condition order (p = .021), the difference may partly reflect familiarization with the task.
Subjective Measures

Subjective Measures by Configuration (n = 20)
TABLE III. Subjective measures at participant level (n = 20), median [IQR]. Higher is better for SUS and the two five-point items, lower for NASA-TLX and its sub-scales.
| Measure | 2D Monitor | MR-GLi | r | p |
|---|---|---|---|---|
| SUS | 76.2 [21.2] | 77.5 [21.2] | 0.15 | .506 |
| NASA-TLX | 49.5 [23.4] | 41.7 [22.3] | 0.26 | .261 |
| Mental | 62.5 [41.2] | 37.5 [37.5] | 0.48 | .034 |
| Physical | 42.5 [32.5] | 40.0 [22.5] | 0.06 | .825 |
| Temporal | 22.5 [30.0] | 27.5 [26.2] | 0.04 | .888 |
| Performance | 32.5 [50.0] | 30.0 [60.0] | 0.00 | 1.000 |
| Effort | 45.0 [46.2] | 42.5 [36.2] | 0.22 | .346 |
| Frustration | 42.5 [50.0] | 35.0 [32.5] | 0.20 | .359 |
| Gaze-shift ease | 2.0 [1.2] | 5.0 [1.0] | 0.68 | .002 |
| Task focus | 3.0 [2.0] | 4.0 [2.0] | 0.39 | .081 |
Gaze-shift ease improved significantly (2.0 → 5.0, p = .002, r = .68), and mental demand also favored MR-GLi (p = .034, r = .48). Usability and overall workload showed no significant difference but numerically favored MR-GLi (SUS 77.5 vs 76.2; NASA-TLX 41.7 vs 49.5), and 14 of 20 participants found MR-GLi easier to use.
Summary
- MR-GLi registers the reaction torque indicator and wrist-camera image to the gripper, so visual feedback travels with the manipulation site instead of sitting on a separate monitor.
- Against a strong 2D-monitor baseline with identical bilateral control, RTI encoding and camera content, torque regulation stayed at a comparable level. Workload and usability showed no significant differences but numerically favored MR-GLi (SUS 77.5 vs 76.2; NASA-TLX 41.7 vs 49.5), indicating a tendency toward higher usability and lower perceived workload.
- The gain is in information access: perceived gaze-shift burden dropped significantly, suggesting that gripper-linked MR overlays can improve access to visual feedback without degrading torque-regulation performance.
Citation
@misc{sasago2026mrglimixedrealitybasedgripperlinked,
title={MR-GLi: Mixed Reality-Based Gripper-Linked Overlays for Underwater Robot Arm Teleoperation via Bilateral Control},
author={Masashi Sasago and Masato Kobayashi and Yuki Uranishi},
year={2026},
eprint={2609.16041},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.16041},
}
Contact
Masato Kobayashi (Assistant Professor, The University of Osaka, Kobe University, Japan)
- X (Twitter)
- English : https://twitter.com/MeRTcookingEN
- Japanese : https://twitter.com/MeRTcooking
- Linkedin https://www.linkedin.com/in/kobayashi-masato-robot/