VAPS

Anonymous submission · under review

Continue, Abort, or Fall

Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics

Anonymous Authors Author and institution information is withheld for double-blind review.

Continuenominal policy

Undisturbed, the nominal policy completes the side flip on the LimX Oli and lands on its feet.

Abortabort policy

Handed the maneuver mid-flight, the abort policy gives up the flip and lands feet first.

Fallprotective-fall policy

Handed the same maneuver, the protective-fall policy performs a controlled fall protecting head and hands.

The same side flip on the physical LimX Oli with three possible endings. In the abort and protective-fall clips the switch is triggered by an operator mid-motion, and the segment around the switch plays at ¼ speed.

Supplementary video

Abstract

Dynamic humanoid motions such as flips risk hardware damage, due to suboptimal policies, disturbances or sim-to-real gaps. A motion-tracking policy offers no way out once the maneuver leaves its reference, and a backup policy needs to take over to protect the hardware for a least-damage landing. Which backup to use matters as much as when to switch.

We present Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision. Besides a protective-fall policy, we also train an abort policy which can abort the motion at any time, landing on its feet. At every control step, learned predictors estimate whether the nominal tracking policy and the abort policy remain viable over a short horizon, and a least-sacrificial hierarchy keeps the most task-ambitious behavior that remains viable.

In simulation with randomized disturbances, VAPS sharply reduces head contact and hand contact, which are the dominant sources of hardware damage, with both a Unitree G1 and a LimX Oli; on the LimX Oli, we validate the viability predictors and the full VAPS controller for side-flip motions, with the whole stack running on one core of the robot's own CPU. VAPS Pareto-dominates the strongest single-network alternatives we could train, including an end-to-end safe-tracking policy and students distilled from VAPS's own oracle-routed decisions, on the task-success/head-impact frontier. We also show that VAPS is a powerful framework to supervise undertrained policies and protect the hardware.

Contributions

1

Least-sacrificial safety

VAPS formulates intervention during a dynamic maneuver as a selection over an ordered hierarchy of behaviors — continue, abort, and protective fall — committing to the smallest sacrifice of task ambition that remains viable.

2

Policy-conditioned, receding-horizon viability

We learn one viability predictor for the nominal policy and the abort policy from counterfactual rollouts and evaluate it continuously from proprioceptive history.

3

Analysis of single-policy alternatives

We investigate monolithic alternatives, including end-to-end RL policies and students distilled from VAPS with DAgger, and show that VAPS dominates all of them.

4

Hardware validation

We develop and test VAPS in simulation on a Unitree G1 and a LimX Oli, and deploy and validate the sim-to-real transfer on the LimX Oli.

In simulation — Unitree G1, 6,144 disturbed episodes

27.7% → 1.6% head-contact rate, nominal policy alone vs. the full VAPS hierarchy.
37.9% → 26.1% fall rate over the same episodes — the robot reaches the ground far less often.
0.416 s mean warning lead — about 60% earlier than final-outcome (SafeFall-style) predictors at the same false-alarm rate.
−2.5 pts the cost of supervision: 59.6% motion success against 62.1% unsupervised.

On hardware — LimX Oli side flips

13 / 13 trials classified correctly — eight successful executions, five failures, no false alarm.
0.544 s mean lead on the failures, each flagged before the corresponding simulated termination (range 0.12–0.80 s).
15 / 18 closed-loop trials in which the selected specialist met its objective — 83.3%, disturbances included.
0.29 ms policy + both predictors at the handoff step on one CPU core of the robot — 1.45% of the control period.

In simulation

The three candidate behaviors, from the same disturbed state

Unitree G1 · simulation
Nominal

Continue the maneuver. Undisturbed, the nominal policy completes the side flip and lands on its feet.

Unitree G1 · simulation
Abort

Give up the flip, keep the feet. The red ghost is the counterfactual — the same state replayed under the nominal policy, which ends up on the floor.

Unitree G1 · simulation
Protective fall

Once upright recovery is gone, stop defending the task and shape the impact away from the head and hands.

The selector running, one robot each

Unitree G1 · simulation
 Nominal → abort

A bar per candidate against its own threshold tick. The nominal's viability collapses, the abort's holds, and the selector hands over; the red ghost is what the nominal would have done instead.

LimX Oli · simulation
 Nominal → abort

The same rule, the same read-out, on the Oli model — the robot the hardware trials use. The green figure is the reference motion the nominal policy is tracking.

Both renders run the paper's own components, so what the bars show is what the deployed selector reads.

One real state, three futures

LimX Oli · hardware
The trial that happened

The predictor calls the switch at 5.05 s and the abort policy takes the robot back to a stance. The clock is the trial's own.

The same state, replayed under each candidate

LimX Oli · simulation
No switch

Carrying on with the nominal policy from the same state puts the robot on the ground: the maneuver was no longer recoverable.

LimX Oli · simulation
Abort

The abort was still viable, and it is what the selector chose — the outcome the trial above actually got.

LimX Oli · simulation
Protective fall

The terminal option: the robot goes down, but with the impact shaped away from the head and hands. Available here, and the right choice a little later.

All four are frame-locked to the same clock and play at real time. The green figure in the replays is the reference motion; the hardware panel ends 3.4 s after the switch, where someone walks into the laboratory shot.

Results in simulation

Closed-loop: routing between the backups keeps the best of both

Table IIIClosed-loop performance under randomized disturbances

All episodes (%) Fallen episodes only — median [mean]
Method Success ↑Head ↓Hand ↓Other ↓Fall ↓ Pk. head (N)Pk. non-foot (N)
Nominal only 62.1 ±0.2 27.7 ±0.2 36.6 ±0.2 36.3 ±0.1 37.9 ±0.2 540.2 [954.8] 2124.6 [2608.6]
VAPS — shared alarm rule, τnom = 0.20
VAPS (Nominal–Abort) 59.6 ±0.4 5.0 ±0.4 7.4 ±0.5 5.5 ±0.3 7.4 ±0.5 833.5 [1076.0] 2219.5 [2110.4]
VAPS (Nominal–ProtFall) 59.6 ±0.4 2.0 ±0.1 11.8 ±0.6 39.9 ±0.5 40.4 ±0.4 0.0 [21.8] 1871.8 [2058.6]
VAPS (one-shot) 59.6 ±0.4 2.2 ±0.2 9.3 ±0.4 25.6 ±1.7 26.2 ±1.7 0.0 [62.2] 1855.3 [2073.4]
VAPS (full hierarchy) 59.6 ±0.4 1.6 ±0.2 8.8 ±0.5 25.6 ±1.8 26.1 ±1.7 0.0 [28.8] 1835.7 [2032.3]
Monolithic, memoryless — one network, no switching, one frame in
End-to-end RL (λ = 10) 61.9 ±0.9 24.0 ±1.6 36.1 ±1.3 36.7 ±0.8 37.9 ±1.0 388.2 [846.7] 2051.5 [2372.4]
Distilled (base) 57.4 ±0.8 23.1 ±1.1 33.6 ±1.1 30.4 ±0.8 34.0 ±1.1 756.3 [1100.8] 2453.8 [2608.8]
Distilled (+ protect) 58.0 ±0.6 17.9 ±1.4 36.7 ±0.5 36.1 ±0.5 38.4 ±0.6 18.2 [712.8] 2329.8 [2552.1]
Distilled (+ protect, DAgger) 56.3 ±0.7 10.1 ±1.3 39.7 ±1.0 40.5 ±0.7 42.3 ±1.0 0.0 [326.6] 2132.8 [2456.4]
Monolithic with memory — one network, no switching, history in the input or a carried state
Distilled frame-stack 5 (base) 55.5 ±3.7 25.4 ±0.3 39.2 ±3.3 35.0 ±1.2 39.7 ±3.3 675.0 [1039.2] 2383.7 [2498.0]
Distilled GRU (base) 58.7 ±1.1 24.2 ±0.8 34.8 ±1.1 32.2 ±0.8 35.2 ±1.1 774.2 [1121.0] 2482.6 [2619.2]
Distilled GRU (+ protect) 56.0 ±1.0 15.4 ±0.9 35.3 ±1.4 34.3 ±1.9 37.0 ±1.3 0.0 [646.8] 2256.1 [2502.4]
Distilled GRU (+ protect, DAgger) 56.1 ±1.3 17.9 ±1.2 33.6 ±1.4 32.2 ±1.3 34.8 ±1.5 227.2 [835.9] 2369.2 [2574.4]

± is over five seeds. Nominal–Abort and Nominal–ProtFall restrict the hierarchy to a single backup; all four VAPS rows share the same alarm rule and therefore the same motion success. One-shot consults the abort predictor once, at the alarm, and commits to the backup it picks; full hierarchy additionally keeps the abort predictor running during the abort and escalates to the protective fall when its prediction drops. Bold marks the best value per column on the leading statistic; no head median is bolded because several rows tie at zero.

VAPS versus monolithic policies

Can a single network replace the structure?

  • Single-network RL — the protective specialist's damage terms folded into the task reward with weight λ, tracking terminations removed, warm-started from the nominal, λ swept at five seeds each.
  • Distilled students — copied from a teacher at least as good as VAPS: at every alarmed state, counterfactual rollouts pick the behavior that succeeds.
  • Two arms tilt the student toward protection (+ protect, + DAgger); two more add memory (a five-frame stack, a GRU) to rule out missing information rather than missing structure.

What the frontier shows

  • VAPS runs along the bottom edge: from about 55% success at 2% head contact to 62% at 4.6%.
  • The end-to-end policies never leave the nominal's corner — at most four points of head contact, before the weight breaks training and success scatters.
  • The students descend, then plateau near 10%. Architecture does not move that plateau; the training data does.
  • The DAgger shard takes the memoryless student from 17.9% to 10.1% and leaves the recurrent one where it was — its segments begin at the alarm, with no history for a hidden state to carry.
  • They learned how to protect the head, not when protecting is right: episodes ending upright without completing the maneuver run to fourteen points for VAPS and low single digits for the students.
Head-contact rate against motion success, one marker per seed. VAPS traces a curve along the bottom of the plot as its switching threshold is swept; end-to-end arms cluster near the nominal policy; distilled students, with and without memory, descend to about 10 percent head contact and plateau.
Head-contact rate against motion success on the population of Table III, one marker per seed; down and right is better. At τ = 0.99, VAPS matches the nominal's success at 4.6% head contact, against 24.0% for the best end-to-end arm and 10.1% for the best student; every lower τ trades success for still less head contact, down to 1.6% for the full hierarchy. Up to seed spread on the success axis, the non-dominated set consists of VAPS points only.
Folding the specialists into one network preserves the skills but loses the decision of how much of the task to give up — exactly the part the explicit structure keeps inspectable.

Supervising new policies

  • VAPS is trained on a large dataset of randomized, perturbed simulation, so the same pipeline generalizes across policies — a protective framework for a new policy, or a new robot.
  • The test: the nominal policy retrained from scratch, a checkpoint every 2,000 iterations, VAPS deployed on all 16 from untrained to converged.
  • The abort and protective-fall policies are reused; only the nominal predictor is retrained.
  • Head contact under VAPS stays flat near the floor, while the unprotected policy ranges from 30 to 99%.
  • The two completion curves lie almost on top of each other — protection at every stage of training, for very little task success.
Two panels over nominal-policy training iteration: completion rate of the deployed policy alone and under VAPS track each other closely; head contact of the deployed policy alone starts near 100% and settles around 30%, while under VAPS it stays below 5% throughout.
VAPS around 16 checkpoints of an independently trained nominal policy. Completion (top) is within a few points of the unsupervised policy throughout; head contact (bottom) stays flat regardless of how well the policy has learned the maneuver.
Whether a state is safe cannot be decided from the state alone. It depends on what the robot does next and on how much time remains to do it. Between continuing the maneuver and falling, VAPS places a third option that prior work usually omits: aborting the maneuver while staying upright.

BibTeX

@article{anonymous2026vaps,
  title   = {Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS)
             for Safe Humanoid Acrobatics},
  author  = {Anonymous Authors},
  journal = {Under review},
  year    = {2026}
}