Reward-DAgger monitors the temporal evolution of task progress predicted by a
general-purpose reward model. Given the observation history \(o_{\leq t}\) and task instruction
\(\ell\), the reward model outputs a scalar progress estimate:
\[
\phi_t = f_{\psi}(o_{\leq t}, \ell).
\]
We then examine the recent progress history at two timescales: a short window
of length \(L_s\) for detecting abrupt progress drops, and a long window of
length \(L_l\) for detecting sustained progress plateaus, where \(L_l > L_s\).
In each window, we compute the Spearman rank correlation between predicted progress
and time:
\[
\rho_t^s = \rho\!\left(\phi_{t-L_s+1:t}\right),
\qquad
\rho_t^l = \rho\!\left(\phi_{t-L_l+1:t}\right).
\]
Since correlation captures the direction of a trend but not its magnitude, the short-window
gate also checks the maximum observed progress drop:
\[
\Delta_t =
\max_{j \in \{t-L_s+1,\ldots,t\}}
\left(\phi_j - \phi_t\right).
\]
The two gates are combined as:
\[
g_t^s =
\mathbb{I}\!\left[
(\rho_t^s < \tau_s)
\wedge
(\Delta_t \geq \delta_{\min})
\right],
\qquad
g_t^l =
\mathbb{I}\!\left[
\rho_t^l < \tau_l
\right],
\]
\[
g_t = g_t^s \vee g_t^l .
\]
If either gate fires, Reward-DAgger requests expert
takeover. The resulting corrections are aggregated with the existing dataset and used to finetune
the policy, while the reward model and gate configuration remain fixed across DAgger rounds.
Example progress drops and plateaus from the real robot evaluation tasks. All videos are played at 2× speed.
Reward-DAgger improves autonomous success rates across the evaluated simulation tasks, while achieving the highest or near-highest ROHE throughout data collection.
Policy rollouts before and after finetuning with Reward-DAgger. All videos are played at 2× speed.