Explain the next action
Connect the current state to a cause, a possible consequence, and an intended effect. Specify the next action, what to avoid, and whether the task is complete.
ARCMore than 5× higher success on reasoning-heavy tasks, with no change to the model architecture.
ARC teaches existing robot foundation models to reason, without new robot demonstrations or foundation-scale training.
RoboLab-120 · Default instructions · success rate Cosmos3-Nano-Policy 36.8 → 48.8%, π0.5 28.0 → 45.3%
Zero-shot evaluation on RoboLab-120 and MolmoSpaces.
Success rates in percent; gains in percentage points.
ARC improves pretrained robot policies through action-grounded reasoning, without new robot demonstrations or foundation-scale training. Each trace explains why the next action is appropriate and what it should accomplish. The policy learns to use that explanation to choose its actions.
The same recipe improves both π0.5 and Cosmos3-Nano-Policy: new state-of-the-art results on RoboLab-120 and MolmoSpaces, gains of up to 50 percentage points on RoboLab-Reasoning-50, and 91.7% real-robot success for π0.5 + ARC, up from 9.5%.
ARC here is π0.5 fine-tuned with the ARC recipe.
No additional hardware-specific fine-tuning.
The videos start together. Each policy retains its recorded timing; shorter runs hold on their final frame.
Two additional ARC rollouts in dynamic environments.
75K demonstrations. 1.2M annotated frames.
Action-grounded reasoning from existing robot data.
Connect the current state to a cause, a possible consequence, and an intended effect. Specify the next action, what to avoid, and whether the task is complete.
Automatically recover reasoning from DROID demonstrations: approximately 75K episodes and 1.2M action-aligned frames, with no new robot trajectories.
Fine-tune the pretrained model to predict demonstrated actions conditioned on causal traces, observations, and instructions. Tailor the recipe to each architecture.
Jointly fine-tune the vision-language backbone, trace encoder, and action expert. Factual examples use flow matching; contradictory traces receive a counterfactual margin loss.
ARC policies are less sensitive to instruction detail: performance stays stronger across vague, default, and specific instructions, while base policies deteriorate as instructions become less explicit.
π0.5 + ARC comes within 0.4 percentage points of its ten-step base policy using one Euler step. Cosmos3-Nano-Policy + ARC exceeds its four-step base policy with two UniPC steps.
Solver-step reductions concern action generation; reasoning and context encoding remain separate costs.
On RoboLab-120 (specific), Cosmos3-Nano-Policy + ARC reaches the instruction-only baseline's 10K-iteration performance with approximately 4.3× fewer training updates, then continues to improve.
Ablations vary the external reasoner, trace content, and fine-tuning strategy.
A background VLM refreshes reasoning from the current camera views. The policy uses the latest validated trace while executing 15-action chunks at 15 Hz.
Reasoning refresh and robot control are separate rates. The example rollout in the paper refreshes reasoning at approximately 1 Hz.
Models, code, and ARC-Trace-DROID will be released after the review period.