CUSP: CUSUM-Governed Survival Hazard Alarms at the Perception Onset for Off-Road Navigation

Inuk Kang, Seung-Woo Seo
Seoul National University, Republic of Korea
arXiv preprint (2026)
The challenge this paper addresses: the autonomy model keeps proposing paths and raises no alarm; the human perceives the safe-to-danger transition (the onset) and takes over; CUSP learns this perception and its accumulated statistic crosses the threshold shortly after the onset.
(Top) The autonomy model keeps proposing candidate paths as the robot approaches a deteriorating slope, and nothing in the stack raises an alarm. (Middle) The human perceives the safe-to-dangerous transition, marked as the onset oi (red box), and takes over before failure. (Bottom) CUSP learns this perception. It accumulates evidence of the transition over time and raises an alarm once that evidence exceeds the threshold h, here shortly after the onset.

Abstract

Off-road navigation exposes a robot to potentially hazardous terrain en route. Although learning-based navigation uses safety supervision to choose which path to drive, it provides no runtime alarm when the robot following that path is heading into danger. Such an alarm must be learned from field logs, where human intervention preempts the failure and the failure itself is therefore never observed. The human judges driving unsafe early but typically intervenes only once failure is clearly near, so the intervention marks that judgment late. That earlier judgment is what a runtime alarm must detect, yet no prior intervention-supervised method has targeted it. To address this problem, we introduce CUSP (CUSUM-governed Survival model of the Perception onset), a model-agnostic runtime hazard alarm that learns this moment from intervention-terminated logs. "Cusp" is a word for the point at which one state is about to turn into another, and the moment we target is exactly such a cusp: the point at which safe driving turns unsafe in a human's judgment. We call this point the perception onset and annotate it separately from the intervention. A visual hazard head is trained on the annotated onset with a discrete-time survival objective so that driving with and without an onset both supervise the head, and a CUSUM accumulates the predicted onset risk into alarms. We evaluate CUSP at five unseen sites, two autonomous and three teleoperated, with 142 events and every method tuned to the same rate of ten false alarms per hour. CUSP detected 85 events compared to 26 for the best of nine adapted baselines, and the margin comes from hazards to which every signal in the navigation model is blind.

Method

CUSP overview: (a) training with episode annotation, hazard-prediction head on frozen DINOv3 features with FiLM conditioning, survival targets and masks, and evidence-graded weights; (b) runtime hazard score, standardization, one-sided CUSUM and site commissioning.
(a) Training. Each episode is annotated with the perception onset oi and the end just before intervention. Frozen DINOv3 features are pooled over the committed-path corridor with temporal differences and global context and FiLM-conditioned on path curvature and speed. Survival targets and masks are anchored on oi, and per-event weights wi are graded by IMU tilt around the onset (training only). (b) Runtime. The onset-risk score is robust-standardized and accumulated by a one-sided CUSUM under a per-site false-alarm budget.

Results

Method Detections / 142
(Sites I / II / III / IV / V)
Margin > 0.5 / 1 / 2 s C(+3) Delay, median [IQR] (s)
Obstacle proximity
NavDP3 (0/3/0/0/0)3 / 2 / 20.014+2.04 [1.8–5.7]
CARE10 (0/5/1/1/3)10 / 10 / 100.021+4.59 [1.7–5.7]
Plan / policy-conditioned
LaND12 (0/0/3/3/6)11 / 11 / 100.049+2.58 [1.4–3.8]
FIPER21 (1/10/6/0/4)19 / 16 / 120.028+5.24 [3.2–8.4]
PAAD26 (0/14/3/1/8)25 / 22 / 190.056+5.13 [2.8–7.7]
Traversability
STERLING7 (1/5/0/1/0)7 / 6 / 50.000+5.21 [4.4–6.0]
CAHSOR9 (2/6/1/0/0)8 / 7 / 60.014+5.60 [3.4–9.4]
SALON24 (0/20/2/2/0)19 / 17 / 130.028+4.69 [3.2–8.4]
WVN26 (1/16/2/3/4)21 / 20 / 170.063+5.00 [2.5–7.4]
Ours
CUSP (intervention label)34 (3/17/8/4/2)34 / 32 / 240.127+2.91 [1.6–5.4]
CUSP (ours)85 (4/40/16/4/21)77 / 72 / 640.218+3.72 [1.9–6.5]

Five-site results under a common 10/h false-alarm budget (142 events: 11 / 50 / 28 / 13 / 40 at Sites I–V). Detections: events with an alarm inside the perception-to-action window [oi, end]. Margin: detections whose alarm precedes the intervention by more than 0.5 / 1 / 2 s. C(+3): fraction of all events alarmed within 3 s of the onset. Delay: alarm time minus onset over detected events. Each baseline is reported at its best configuration (CUSUM with k ∈ {1, 2, 3} or the instantaneous rule). Table II of the paper.

Example detections at Sites II, III and V: frames sampled around the onset and plots of the per-frame score and CUSUM statistic against time from onset, with the threshold, the perception-to-action window and the alarm time marked.
Example detections at Sites II, III, and V. (a) Four vegetation intrusion events at Site II, where the path is visibly impassable and the baselines also detect the events. (b) Mounting a right cross-slope (Site V). (c) Sliding into a downward right cross-slope (Site III). (d) Entering a left cross-slope overgrown with dense bushes (Site III). In (b)–(d) nothing in the view marks the path as impassable, and CUSP is the only method that alarms. Each plot shows the score st (grey) and the CUSUM statistic Ct (blue) against time from onset, with the threshold h as a dashed line, the perception-to-action window [oi, end] shaded in green, and the alarm time marked by the orange vertical line. Frames above each plot are sampled at the indicated times from onset.
Detection rate against realized false-alarm rate on the teleoperated Sites III–V as the threshold varies; CUSP dominates every baseline at every budget, and the dashed line marks the 10/h operating point.
Detection rate against realized false-alarm rate on Sites III–V as the threshold h varies. The dashed line marks the 10/h operating point. The lead holds at every budget: the detection rate the best baseline reaches at 10/h, CUSP reaches at under 3/h, and the 0.30 the best baseline reaches at 30/h, CUSP reaches at 6/h.

Real-world Experiments

Alarm examples from the held-out sites, grouped by site. Each clip shows the left ZED 2i image with the committed path, the per-frame score st and the CUSUM statistic Ct against the threshold h, the onset and the intervention, and the alarm times of CUSP and the three strongest baselines (PAAD, WVN, SALON). Site II is driven autonomously by NoMaD with CUSP attached as an alarm; Sites III–V are teleoperated.

Site II (autonomous). A fallen tree lying across the path, a steep downward slope on the left, an upward slope that turns into a hill, and a longitudinal pit with a slope on the right.
Site III (teleoperated). A right cross-slope with a tree directly ahead (shown at 2×), and a left cross-slope that ends at a tree.
Site IV (teleoperated). Leaving a stone-paved path into vegetation and rocks, and tall grass closing in on the right while following the paved path.
Site V (teleoperated). A pile of logs and a tree on the left, a hill with a slope and vegetation on the right, a mound on the right, and three downward slopes on the left, one of them with a log.
Missed cases (Site III). Two events where Ct never reaches h: an approach into dense bushes that the head scores no differently from traversable ground, and a traverse along the left upper slope that looks like nominal driving.

BibTeX

@article{kang2026cusp,
  title   = {CUSP: CUSUM-Governed Survival Hazard Alarms at the Perception Onset for Off-Road Navigation},
  author  = {Kang, Inuk and Seo, Seung-Woo},
  journal = {arXiv preprint arXiv:2610.07882},
  year    = {2026},
  eprint  = {2610.07882},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url     = {https://arxiv.org/abs/2610.07882}
}