Real Franka / DROID hardware rollouts. For each task a router chooses between a
base planner
(no-think · smaller · no-memory) and a
thinking planner
(larger · memory), and
escalates only when the base planner would fail
— on the multi-step task it decides per subtask.
Planning latencies are the real VLM call times (vlm.latency_ms). The on-video planning bars are
sped up ~× faster than real time to stay watchable, but the counter ticks up
to the true model latency. Some clips are trimmed and background bystanders are blurred.