Bi-MoDe: Bilateral Control-based Imitation Learning via Modifier-Conditioned Decoding for Modulation of Execution Speed and Contact Intensity

The University of Osaka / Kobe University, Japan

Takumi Kobayashi*, Masato Kobayashi*, Yuki Uranishi
* Co-first authors equally contributed to this work.
Overview

Abstract:
Bilateral control-based imitation learning captures both position and force information, making it well suited to contact-rich manipulation. However, existing approaches provide limited means for an operator to specify how a learned task should be executed at inference time, such as slowly or quickly, gently or firmly. We propose Bi-MoDe, a modifier-conditioned decoding framework that injects a constrained latent into every layer of the Transformer action decoder via adaLN-Zero, allowing behavioral directives to directly influence action-chunk generation. We evaluate the method on a real-world whiteboard wiping task with combinations of temporal and physical modifiers. Bi-MoDe improves physical directive following over the action-chunking baseline while maintaining comparable temporal control. An ablation further shows that decoder conditioning and latent-space composition interact, and that their combination is important for accurate physical directive following.

Concept of Bi-MoDe
Concept of Bi-MoDe — the same wiping motion at a commanded speed and a commanded contact force

The same task can be executed in different ways: a wipe may need to be quick or slow, gentle or firm. With Bi-MoDe, the operator specifies how the learned task is executed at inference time — a temporal modifier sets the motion speed (slow / moderate / fast) and a physical modifier sets the contact force (weak / moderate / strong).

Overview of Bi-MoDe
Overview of Bi-MoDe — the constrained latent is appended as an encoder token and, through adaLN-Zero, modulates every layer of the action decoder

Bi-MoDe is a Transformer-based CVAE that takes the follower robot’s joint angles, velocities and torques, and outputs action chunks of the leader robot’s joint angles, velocities and torques. The constrained latent z_c, aligned with the modifier labels, enters through two paths: as a token appended to the encoder input, and — the key addition — through adaLN-Zero, modulating every layer of the action decoder, so the directive directly shapes each generated action chunk.

Method

Modifier Directives

Data is collected with four-channel bilateral control. Before each demonstration the operator assigns a pair of scalar labels:

  • Temporal modifier — motion speed: slow (0.0), moderate (0.5), fast (1.0)
  • Physical modifier — contact force: weak (0.0), moderate (0.5), strong (1.0)

The label is assigned before the trial rather than inferred afterwards, so the operator adjusts the motion to the intended level rather than describing what was done.

Modifier-Conditioned Decoding

One encoder layer and one decoder layer
One Encoder Layer and One Decoder Layer — a zero-initialized MLP regresses the scale, shift and gate of all three decoder sub-layers

Bi-MoDe builds on a Transformer-driven CVAE that splits the latent into a constrained part z_c and an unconstrained part z_u. The key contribution is that z_c enters through two conditioning paths: as a token appended to the encoder input, and — new here — modulating every decoder layer via adaLN-Zero, acting on all three sub-layers (self-attention, cross-attention, feedforward).

The scale, shift and gate are regressed from z_c by a single zero-initialized MLP, so no direct adaLN modulation is applied at the start of training while encoder-side conditioning stays active. Residual branches are gated by (1 + α) rather than α alone, so observation-dependent information still flows through the DETR-style decoder from initialization.

Inference

The operator selects a directive level on each axis, and each level is realized by a fixed z_c value determined in advance: the trained encoder is applied to every demonstration, and the median constrained latent within each label level becomes the command for that level. A given directive level is therefore realized by the same latent command on every rollout.

Experiment

The follower robot
The Follower Robot — Joint2 and Joint3 bear the load when the cleaner is pressed against the board

Two ROBOTIS OpenManipulator-X arms serve as leader and follower. The follower holds a whiteboard cleaner against the board and sweeps end to end and back, repeated three times per trial. Joint angles, velocities and torques were recorded from both robots at 1000 Hz; the models take proprioceptive state alone, so the comparison isolates the conditioning mechanism from visual factors.

Demonstrations cover all 9 combinations of the two modifiers, 5 demonstrations each (45 demonstrations). At inference each configuration was evaluated over the same 9 conditions with 5 rollouts each (45 rollouts).

Effect of the Modifiers in the Demonstrations

Effect of the temporal modifier
Temporal Modifier — fast, moderate and slow; over six seconds the three conditions complete three, one and a half, and one stroke
Effect of the physical modifier
Physical Modifier — weak, moderate and strong; the arm follows the same path while the Joint2 and Joint3 torques separate by commanded force

Ablation Design

A 2 × 2 ablation over two factors. Factor A (adaLN-Zero): whether z_c additionally modulates the LayerNorm parameters of every decoder sub-layer. Factor B (z_u): whether the unconstrained latent is included in the latent token supplied to the encoder. All four configurations append the modifier token to the encoder input and share the same latent structure and training objective — adaLN-Zero off with z_u on reproduces the action-chunking baseline, and Bi-MoDe turns adaLN-Zero on and removes z_u.

Directive-following fidelity is measured by the Modifier Directive Error (MDE), the distance in coefficient space between the line fitted to the generated trials and the reference line fitted to the demonstrations (lower is better). The slope ratio isolates the slope component: 1.0 means the generated motion spans the same dynamic range as the demonstrations.

Results

TABLE II. Task success rate and directive-following performance. MDE ↓; slope ratio → 1.0.

ConfigurationDesignTSRPhysicalTemporal
Decoder cond.zuMDEslopeMDEslope
Baseline (ACT-based)–✔45/450.1230.8960.1611.104
No zu––45/450.3000.7540.0931.069
Decoder conditioned✔✔45/450.1390.8540.0390.959
Bi-MoDe (Ours)✔–45/450.0360.9680.1161.096
Directive following results
Cycle Duration vs Commanded Speed (top) and Contact Torque vs Commanded Force (bottom) — columns are the four configurations, each with the reference line fitted to the demonstrations

The two design factors interact. Reading the physical slope ratio, the baseline sits at 0.896; removing z_u alone moves it to 0.754 and adding adaLN-Zero while keeping z_u moves it to 0.854 — both further from unity. Only Bi-MoDe, which combines the two, reaches 0.968. MDE follows the same ordering (0.123 → 0.300 / 0.139 under either single change, 0.036 under both). Neither change helps alone because each removes one of the two effects without the other.

All four configurations follow the temporal directive closely (slope ratios 0.959–1.104, MDE 0.039–0.161). The temporal modifier acts on stroke duration, which the policy regulates through the commanded joint trajectory alone, whereas the physical modifier must survive contact with the board — which is exactly where the conditioning design matters.

Joint2 and Joint3 profiles
Joint2 and Joint3 Profiles — demonstrations (left) and Bi-MoDe (right); angles for the three commanded speed levels, torques for the three commanded force levels

In both halves the angle traces separate in duration while following the same path, and the torque traces separate in magnitude during the contact phase — each directive acts on the quantity it is meant to control.

Summary

  • Bi-MoDe propagates a constrained latent aligned with scalar modifier directives through every layer of the Transformer action decoder via adaLN-Zero, making execution speed and contact intensity specifiable at inference time.
  • On a real contact-rich wiping task it improved physical directive following (MDE 0.036, slope ratio 0.968) while maintaining comparable temporal modulation and 45/45 task success.
  • A 2 × 2 ablation showed that decoder conditioning and latent-space composition jointly contribute — neither change helps on its own.

Citation

@misc{kobayashi2026bimodebilateralcontrolbasedimitation,
      title={Bi-MoDe: Bilateral Control-based Imitation Learning via Modifier-Conditioned Decoding for Modulation of Execution Speed and Contact Intensity}, 
      author={Takumi Kobayashi and Masato Kobayashi and Yuki Uranishi},
      year={2026},
      eprint={2609.16040},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.16040}, 
}

Contact

Masato Kobayashi (Assistant Professor, The University of Osaka, Kobe University, Japan)

* Corresponding author: Masato Kobayashi