Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

University of Southern California
* Equal contribution

We introduce Reward-DAgger, a robot-gated interactive imitation learning framework that detects progress drops and plateaus using estimates from a general-purpose reward model.

framework
Short-window gate detects rapid decreases in progress, capturing immediate failures such as dropping an object.
Long-window gate detects progress plateaus corresponding to stalled behavior, such as repeatedly failed grasps.
Reward-DAgger operates across policy checkpoints and architectures, including black-box policies, without retraining or recalibrating the gate between DAgger rounds.

Method

Reward-DAgger monitors the temporal evolution of task progress predicted by a general-purpose reward model. Given the observation history \(o_{\leq t}\) and task instruction \(\ell\), the reward model outputs a scalar progress estimate:

\[ \phi_t = f_{\psi}(o_{\leq t}, \ell). \]

We then examine the recent progress history at two timescales: a short window of length \(L_s\) for detecting abrupt progress drops, and a long window of length \(L_l\) for detecting sustained progress plateaus, where \(L_l > L_s\). In each window, we compute the Spearman rank correlation between predicted progress and time:

\[ \rho_t^s = \rho\!\left(\phi_{t-L_s+1:t}\right), \qquad \rho_t^l = \rho\!\left(\phi_{t-L_l+1:t}\right). \]

Since correlation captures the direction of a trend but not its magnitude, the short-window gate also checks the maximum observed progress drop:

\[ \Delta_t = \max_{j \in \{t-L_s+1,\ldots,t\}} \left(\phi_j - \phi_t\right). \]

The two gates are combined as:

\[ g_t^s = \mathbb{I}\!\left[ (\rho_t^s < \tau_s) \wedge (\Delta_t \geq \delta_{\min}) \right], \qquad g_t^l = \mathbb{I}\!\left[ \rho_t^l < \tau_l \right], \] \[ g_t = g_t^s \vee g_t^l . \]

If either gate fires, Reward-DAgger requests expert takeover. The resulting corrections are aggregated with the existing dataset and used to finetune the policy, while the reward model and gate configuration remain fixed across DAgger rounds.

Example Progress Drops and Plateaus

Example progress drops and plateaus from the real robot evaluation tasks. All videos are played at 2× speed.

Put Bread has no progress drop example because the policy never visibly undoes progress on this task, such as dropping the bread once it's lifted.


Failure Detection Experiments

Failure detection frontier: balanced accuracy vs. normalized trigger time for Reward-DAgger and baselines
Reward-DAgger achieves the highest maximum balanced accuracy on every task, reaching at least 0.96. Its frontier also lies above or on par with the baselines.

Transfer of Gating Configurations

Transfer results at a balanced accuracy threshold $\epsilon=0.02$ for three-to-one transfer and held-out Tasks 4 and 6. $\Delta\mathrm{BA}$ is the change in balanced accuracy, $T_{\mathrm{det}}$ is mean normalized trigger time and Pen is matched-accuracy latency penalty.
3-to-1 on T4 on T6
Held-out $\Delta$BA $T_\mathrm{det}$ Pen $\Delta$BA $T_\mathrm{det}$ Pen $\Delta$BA $T_\mathrm{det}$ Pen
Task 0 -1.20.3549.0 -4.40.2860.8 -1.70.3848.7
Task 1 -0.90.3414.3 -11.00.26011.7 -0.80.3793.6
Task 5 1.40.33611.2 -1.30.3521.9 0.00.4006.7
Task 8 0.50.4854.5 -1.30.3521.9 0.00.4006.7
None (all 4) ––– -1.30.3521.9 0.40.4123.2
By calibrating across a broad set of source tasks, reward gate configurations transfer reliably to target tasks, preserving balanced accuracy and inducing a small, controllable latency penalty.
transfer_gating_configurations
Frontier transfer of Reward-DAgger gate configurations. The solid orange curve denotes the native frontier obtained through task-specific tuning, and each dashed curve is obtained by applying configurations from another task's frontier without retuning. Numbers indicate the hypervolume ratios (HVR) for each transferred curve. The dashed grey line indicates the mean normalized length of successful trajectories.

Downstream Policy Improvement

The same reward model and gate configuration is used without additional tuning across eight simulated and real-world tasks, and different policy architectures.

Simulation

transfer_gating_configurations
Reward-DAgger improves autonomous success rates across the evaluated simulation tasks, while achieving the highest or near-highest ROHE throughout data collection.

Real World

Task Method Success rate (%) ROHE
R0 R1 R2 R0 R1 R2
Stack Reward-DAgger 657580 .667.731.643
ThriftyDAgger 657055 .606.556.513
Diff-DAgger 657570 .588.783.600
Pick Reward-DAgger 203040 .541.533.588
ThriftyDAgger 202525 .526.474.424
Diff-DAgger 203540 .548.435.478
Put Reward-DAgger 354045 .594.606.645
ThriftyDAgger 353550 .559.500.526
Diff-DAgger 352030 .500.346.528
Remove Reward-DAgger 758590 .588.714.769
ThriftyDAgger 757080 .526.486.563
Diff-DAgger 757080 .545.613.773
Reward-DAgger consistently improves autonomous success across DAgger rounds, obtaining the highest ROHE in most settings and staying comparable in the others.

Evaluation Videos

Before Finetuning

After Finetuning

Policy rollouts before and after finetuning with Reward-DAgger. All videos are played at 2× speed.


Ablations

Design Choices

Gate design ablation: earliest normalized trigger time for each balanced-accuracy threshold.
Trigger time at balanced accuracy ≥
Task Detector 0.80 0.85 0.90 0.95
Task 0 Full gate .184.211.224.389
Short only .194.228.255–
Long only .219.219.234.411
Absolute thresh. .205.266.331.500
Task 1 Full gate .199.252.261.303
Short only .233.241.299.324
Long only .243.260.279.319
Absolute thresh. .270.270.323.385
Task 4 Full gate .206.212.279.378
Short only .342.342––
Long only .242.312.337.373
Absolute thresh. .313.313.313.357
Task 6 Full gate .215.265.347.448
Short only .215.249.358–
Long only .319.319.329.551
Absolute thresh. .287.314.355.445

Different Reward Model Backbone

Reward model backbone ablation
Success rate (%) ROHE
Task Reward Model R0 R1 R2 R0 R1 R2
Stack Robometer-4B 657580 .667.731.643
RynnValue-4B 656570 .606.645.613
Pick Robometer-4B 203040 .541.533.588
RynnValue-4B 202525 .515.556.571