D3DWA: Adaptive Weight and Prediction-Horizon for Dynamic Window Approach via Dueling Double Deep Q-Network
The University of Osaka / Kobe University

Concept of D3DWA
Conventional DWA fixes its evaluation weights and prediction horizon before navigation, but the right values change with the free space along a route: long horizons suit open areas, while short horizons keep motions feasible in narrow or cluttered regions. D3DWA lets a Dueling Double Deep Q-Network (D3QN) choose both at every control step, while DWA itself still generates the trajectories, checks collisions and issues the velocity command.

System Overview
At every control step (10 Hz), a body-frame observation is built from the lidar scan and wheel odometry — goal distance and bearing, the robot’s velocities, and the nearest obstacle range in 8 angular sectors. The D3QN agent maps this state to one of 100 discrete actions, each decoding to three evaluation weights (progress, clearance, speed) and a prediction horizon of 2–5 s. The DWA planner then scores its candidate trajectories with those parameters and executes the best velocity command.
Why the Horizon Matters
A short horizon (2–3 s) stays reactive and keeps motions feasible in narrow, cluttered areas, while a long horizon (5 s) commits to direct, efficient paths in open space. A single route can contain both, so a horizon fixed before navigation can be a compromise.
Simulation

Evaluation Environments — TS1–TS4 trained, US1 and H1–H3 unseen
TABLE I. Simulation results. TL and PD denote trajectory length and movement posture displacement. T denotes timeout. Best values in each environment and metric are shown in bold.
| Method | TS1 | TS2 | TS3 | TS4 | US1 | H1 | H2 | H3 | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Time | TL | PD | Time | TL | PD | Time | TL | PD | Time | TL | PD | Time | TL | PD | Time | TL | PD | Time | TL | PD | Time | TL | PD | |
| DWA I (1,1,2) | 18.53 | 3.58 | 0.93 | T | – | – | T | – | – | T | – | – | T | – | – | 22.11 | 7.30 | 5.25 | T | – | – | T | – | – |
| DWA II (1,2,1) | 40.42 | 6.90 | 14.76 | 29.65 | 7.09 | 3.61 | 29.20 | 4.51 | 6.41 | T | – | – | T | – | – | 51.19 | 8.04 | 5.23 | 24.32 | 6.84 | 4.07 | T | – | – |
| DWA III (2,1,1) | 17.08 | 3.58 | 0.90 | T | – | – | 29.48 | 4.40 | 5.58 | T | – | – | 47.63 | 5.29 | 7.22 | 24.45 | 6.01 | 1.79 | 22.03 | 5.13 | 1.73 | 51.15 | 11.83 | 13.42 |
| DQDWA | 16.59 | 3.58 | 0.96 | T | – | – | 49.43 | 9.54 | 17.23 | T | – | – | 46.09 | 5.38 | 6.95 | 58.41 | 9.86 | 13.53 | 21.56 | 5.14 | 1.81 | T | – | – |
| D3DWA (w/o horizon) | 18.05 | 3.59 | 0.92 | T | – | – | 28.56 | 4.42 | 5.90 | T | – | – | 37.39 | 5.81 | 7.35 | 36.45 | 10.04 | 10.26 | 25.26 | 6.67 | 4.70 | 37.82 | 9.95 | 6.26 |
| D3DWA (proposed) | 17.01 | 3.59 | 0.88 | 29.22 | 7.02 | 3.62 | 24.11 | 4.01 | 4.46 | 44.12 | 9.73 | 7.72 | 35.73 | 4.93 | 9.46 | 19.27 | 6.26 | 2.81 | 25.85 | 5.64 | 7.53 | 32.84 | 10.04 | 5.95 |
Trajectories
Executed trajectories in all eight environments, coloured by elapsed time. The robot starts at the origin and the goal is marked in red.

DWA I (1,1,2) — fixed parameters

DWA II (1,2,1) — fixed parameters

DWA III (2,1,1) — fixed parameters

DQDWA — tabular baseline

D3DWA (w/o horizon) — weights-only ablation

D3DWA (proposed) — top: elapsed time, bottom: selected prediction horizon
The agent picks long horizons in open stretches and short ones in tight spots — TS2 switches from 5 s to 2 s at the turn into the narrow goal region, exactly where the weights-only variant, fixed at 4 s, times out.
Real World Experiments

Real-Robot Setup (Kachaka) — policies transferred from simulation without fine-tuning
TABLE II. Real-robot results. One trial per condition. T denotes timeout. Best values for each case and metric are shown in bold.
| Method | Case R1 | Case R2 | Case R3 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Time | TL | PD | Time | TL | PD | Time | TL | PD | |
| DWA I (1,1,2) | 12.45 | 3.14 | 1.35 | T | – | – | 19.38 | 5.89 | 4.07 |
| DWA II (1,2,1) | T | – | – | T | – | – | 24.76 | 4.45 | 5.07 |
| DWA III (2,1,1) | 19.10 | 3.43 | 2.93 | T | – | – | 22.62 | 4.33 | 5.19 |
| DQDWA | 13.99 | 3.08 | 1.41 | 20.22 | 3.36 | 5.10 | 24.27 | 6.02 | 6.18 |
| D3DWA (w/o horizon) | 14.91 | 3.22 | 1.72 | T | – | – | 38.70 | 5.51 | 11.64 |
| D3DWA (proposed) | 13.70 | 3.04 | 1.20 | 15.70 | 2.98 | 2.77 | 17.52 | 4.61 | 5.52 |
In R2, only DQDWA and D3DWA reached the goal — the weights-only ablation timed out. In R3, D3DWA had the shortest navigation time — 27.8 % shorter than DQDWA and 54.7 % shorter than the ablation.

Case R1

Case R2

Case R3

Prediction Horizon Selected by the Agent
Summary
- D3DWA jointly adapts the DWA weights and prediction horizon with a Dueling Double DQN, leaving DWA’s trajectory generation and collision checking intact.
- It reached every goal in all 8 simulated environments and was fastest in 6.
- On a real robot it completed all 3 cases, including the narrow R2 where the weights-only ablation failed — horizon adaptation is what the weights alone cannot provide.
Citation
coming soon
Contact
Masato Kobayashi (Assistant Professor, The University of Osaka, Kobe University, Japan)
- X (Twitter)
- English : https://twitter.com/MeRTcookingEN
- Japanese : https://twitter.com/MeRTcooking
- Linkedin https://www.linkedin.com/in/kobayashi-masato-robot/