Anonymous submission · under review
Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics
Undisturbed, the nominal policy completes the side flip on the LimX Oli and lands on its feet.
Handed the maneuver mid-flight, the abort policy gives up the flip and lands feet first.
Handed the same maneuver, the protective-fall policy performs a controlled fall protecting head and hands.
The same side flip on the physical LimX Oli with three possible endings. In the abort and protective-fall clips the switch is triggered by an operator mid-motion, and the segment around the switch plays at ¼ speed.
Dynamic humanoid motions such as flips risk hardware damage, due to suboptimal policies, disturbances or sim-to-real gaps. A motion-tracking policy offers no way out once the maneuver leaves its reference, and a backup policy needs to take over to protect the hardware for a least-damage landing. Which backup to use matters as much as when to switch.
We present Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision. Besides a protective-fall policy, we also train an abort policy which can abort the motion at any time, landing on its feet. At every control step, learned predictors estimate whether the nominal tracking policy and the abort policy remain viable over a short horizon, and a least-sacrificial hierarchy keeps the most task-ambitious behavior that remains viable.
In simulation with randomized disturbances, VAPS sharply reduces head contact and hand contact, which are the dominant sources of hardware damage, with both a Unitree G1 and a LimX Oli; on the LimX Oli, we validate the viability predictors and the full VAPS controller for side-flip motions, with the whole stack running on one core of the robot's own CPU. VAPS Pareto-dominates the strongest single-network alternatives we could train, including an end-to-end safe-tracking policy and students distilled from VAPS's own oracle-routed decisions, on the task-success/head-impact frontier. We also show that VAPS is a powerful framework to supervise undertrained policies and protect the hardware.
VAPS formulates intervention during a dynamic maneuver as a selection over an ordered hierarchy of behaviors — continue, abort, and protective fall — committing to the smallest sacrifice of task ambition that remains viable.
We learn one viability predictor for the nominal policy and the abort policy from counterfactual rollouts and evaluate it continuously from proprioceptive history.
We investigate monolithic alternatives, including end-to-end RL policies and students distilled from VAPS with DAgger, and show that VAPS dominates all of them.
We develop and test VAPS in simulation on a Unitree G1 and a LimX Oli, and deploy and validate the sim-to-real transfer on the LimX Oli.
In simulation — Unitree G1, 6,144 disturbed episodes
On hardware — LimX Oli side flips
The three candidate behaviors, from the same disturbed state
Continue the maneuver. Undisturbed, the nominal policy completes the side flip and lands on its feet.
Give up the flip, keep the feet. The red ghost is the counterfactual — the same state replayed under the nominal policy, which ends up on the floor.
Once upright recovery is gone, stop defending the task and shape the impact away from the head and hands.
The selector running, one robot each
A bar per candidate against its own threshold tick. The nominal's viability collapses, the abort's holds, and the selector hands over; the red ghost is what the nominal would have done instead.
The same rule, the same read-out, on the Oli model — the robot the hardware trials use. The green figure is the reference motion the nominal policy is tracking.
Both renders run the paper's own components, so what the bars show is what the deployed selector reads.
The predictor calls the switch at 5.05 s and the abort policy takes the robot back to a stance. The clock is the trial's own.
The same state, replayed under each candidate
Carrying on with the nominal policy from the same state puts the robot on the ground: the maneuver was no longer recoverable.
The abort was still viable, and it is what the selector chose — the outcome the trial above actually got.
The terminal option: the robot goes down, but with the impact shaped away from the head and hands. Available here, and the right choice a little later.
All four are frame-locked to the same clock and play at real time. The green figure in the replays is the reference motion; the hardware panel ends 3.4 s after the switch, where someone walks into the laboratory shot.
Table IIIClosed-loop performance under randomized disturbances
| All episodes (%) | Fallen episodes only — median [mean] | ||||||
|---|---|---|---|---|---|---|---|
| Method | Success ↑ | Head ↓ | Hand ↓ | Other ↓ | Fall ↓ | Pk. head (N) | Pk. non-foot (N) |
| Nominal only | 62.1 ±0.2 | 27.7 ±0.2 | 36.6 ±0.2 | 36.3 ±0.1 | 37.9 ±0.2 | 540.2 [954.8] | 2124.6 [2608.6] |
| VAPS — shared alarm rule, τnom = 0.20 | |||||||
| VAPS (Nominal–Abort) | 59.6 ±0.4 | 5.0 ±0.4 | 7.4 ±0.5 | 5.5 ±0.3 | 7.4 ±0.5 | 833.5 [1076.0] | 2219.5 [2110.4] |
| VAPS (Nominal–ProtFall) | 59.6 ±0.4 | 2.0 ±0.1 | 11.8 ±0.6 | 39.9 ±0.5 | 40.4 ±0.4 | 0.0 [21.8] | 1871.8 [2058.6] |
| VAPS (one-shot) | 59.6 ±0.4 | 2.2 ±0.2 | 9.3 ±0.4 | 25.6 ±1.7 | 26.2 ±1.7 | 0.0 [62.2] | 1855.3 [2073.4] |
| VAPS (full hierarchy) | 59.6 ±0.4 | 1.6 ±0.2 | 8.8 ±0.5 | 25.6 ±1.8 | 26.1 ±1.7 | 0.0 [28.8] | 1835.7 [2032.3] |
| Monolithic, memoryless — one network, no switching, one frame in | |||||||
| End-to-end RL (λ = 10) | 61.9 ±0.9 | 24.0 ±1.6 | 36.1 ±1.3 | 36.7 ±0.8 | 37.9 ±1.0 | 388.2 [846.7] | 2051.5 [2372.4] |
| Distilled (base) | 57.4 ±0.8 | 23.1 ±1.1 | 33.6 ±1.1 | 30.4 ±0.8 | 34.0 ±1.1 | 756.3 [1100.8] | 2453.8 [2608.8] |
| Distilled (+ protect) | 58.0 ±0.6 | 17.9 ±1.4 | 36.7 ±0.5 | 36.1 ±0.5 | 38.4 ±0.6 | 18.2 [712.8] | 2329.8 [2552.1] |
| Distilled (+ protect, DAgger) | 56.3 ±0.7 | 10.1 ±1.3 | 39.7 ±1.0 | 40.5 ±0.7 | 42.3 ±1.0 | 0.0 [326.6] | 2132.8 [2456.4] |
| Monolithic with memory — one network, no switching, history in the input or a carried state | |||||||
| Distilled frame-stack 5 (base) | 55.5 ±3.7 | 25.4 ±0.3 | 39.2 ±3.3 | 35.0 ±1.2 | 39.7 ±3.3 | 675.0 [1039.2] | 2383.7 [2498.0] |
| Distilled GRU (base) | 58.7 ±1.1 | 24.2 ±0.8 | 34.8 ±1.1 | 32.2 ±0.8 | 35.2 ±1.1 | 774.2 [1121.0] | 2482.6 [2619.2] |
| Distilled GRU (+ protect) | 56.0 ±1.0 | 15.4 ±0.9 | 35.3 ±1.4 | 34.3 ±1.9 | 37.0 ±1.3 | 0.0 [646.8] | 2256.1 [2502.4] |
| Distilled GRU (+ protect, DAgger) | 56.1 ±1.3 | 17.9 ±1.2 | 33.6 ±1.4 | 32.2 ±1.3 | 34.8 ±1.5 | 227.2 [835.9] | 2369.2 [2574.4] |
± is over five seeds. Nominal–Abort and Nominal–ProtFall restrict the hierarchy to a single backup; all four VAPS rows share the same alarm rule and therefore the same motion success. One-shot consults the abort predictor once, at the alarm, and commits to the backup it picks; full hierarchy additionally keeps the abort predictor running during the abort and escalates to the protective fall when its prediction drops. Bold marks the best value per column on the leading statistic; no head median is bolded because several rows tie at zero.
What the frontier shows
@article{anonymous2026vaps,
title = {Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS)
for Safe Humanoid Acrobatics},
author = {Anonymous Authors},
journal = {Under review},
year = {2026}
}