ARC lightbulb logoARC

A Reasoning Recipe for
Robot Foundation Models

Anonymous authors

More than 5× higher success on reasoning-heavy tasks, with no change to the model architecture.
ARC teaches existing robot foundation models to reason, without new robot demonstrations or foundation-scale training.

CodeComing soon ModelComing soon DatasetComing soon RoboLab-Reasoning-50 benchmarkComing soon
Benchmark results

+12.0 points on Cosmos3-Nano-Policy+17.3 points on π0.5

RoboLab-120 · Default instructions · success rate Cosmos3-Nano-Policy 36.8 → 48.8%, π0.5 28.0 → 45.3%

Zero-shot evaluation on RoboLab-120 and MolmoSpaces.
Success rates in percent; gains in percentage points.

91.7%Real-robot task successπ0.5 + ARC · +82.2 pp
61.0%RoboLab-Reasoning-50Cosmos3-Nano-Policy + ARC · +50.0 pp
57.6%MolmoSpaces task successCosmos3-Nano-Policy + ARC · +18.6 pp
ARC labels existing demonstrations with action-grounded traces and fine-tunes pretrained robot policies.
The idea

Reasoning grounded in the next action.

ARC improves pretrained robot policies through action-grounded reasoning, without new robot demonstrations or foundation-scale training. Each trace explains why the next action is appropriate and what it should accomplish. The policy learns to use that explanation to choose its actions.

The same recipe improves both π0.5 and Cosmos3-Nano-Policy: new state-of-the-art results on RoboLab-120 and MolmoSpaces, gains of up to 50 percentage points on RoboLab-Reasoning-50, and 91.7% real-robot success for π0.5 + ARC, up from 9.5%.

Read the full abstract
On real robots

Same task. Three policies.

ARC here is π0.5 fine-tuned with the ARC recipe.
No additional hardware-specific fine-tuning.

3× recorded speed
0:00 / 0:00

The videos start together. Each policy retains its recorded timing; shorter runs hold on their final frame.

When the scene changes

Two additional ARC rollouts in dynamic environments.

The dataset

ARC-TRACE-DROID

75K demonstrations. 1.2M annotated frames.
Action-grounded reasoning from existing robot data.

Robot demonstrations paired with state, cause, consequence, effect, action, avoidance, and completion.
The recipe

How ARC trains robot policies.

01

Explain the next action

Connect the current state to a cause, a possible consequence, and an intended effect. Specify the next action, what to avoid, and whether the task is complete.

02

Label existing demonstrations

Automatically recover reasoning from DROID demonstrations: approximately 75K episodes and 1.2M action-aligned frames, with no new robot trajectories.

03

Fine-tune the policy

Fine-tune the pretrained model to predict demonstrated actions conditioned on causal traces, observations, and instructions. Tailor the recipe to each architecture.

ARC-Trace
[State][Cause][Consequence][Effect][Action][Avoid][Completion]
The ARC-Trace labeling pipeline

π0.5-DROID + ARC

Jointly fine-tune the vision-language backbone, trace encoder, and action expert. Factual examples use flow matching; contradictory traces receive a counterfactual margin loss.

  • Continuous action prediction
  • Trace-conditioned action expert
  • Counterfactual supervision
What changes with reasoning

Robustness, efficiency, and grounding.

01

Less dependence on instruction detail

ARC policies are less sensitive to instruction detail: performance stays stronger across vague, default, and specific instructions, while base policies deteriorate as instructions become less explicit.

02

Fewer action-generation steps

π0.5 + ARC comes within 0.4 percentage points of its ten-step base policy using one Euler step. Cosmos3-Nano-Policy + ARC exceeds its four-step base policy with two UniPC steps.

Solver-step reductions concern action generation; reasoning and context encoding remain separate costs.

03

Reach target success with fewer updates

On RoboLab-120 (specific), Cosmos3-Nano-Policy + ARC reaches the instruction-only baseline's 10K-iteration performance with approximately 4.3× fewer training updates, then continues to improve.

Which choices matter?

Ablations vary the external reasoner, trace content, and fine-tuning strategy.

Complete causal context and full fine-tuning improve how effectively the controller uses the trace.
How the action head attends to reasoning
Attention provides evidence of trace use; the behavioral evaluations test its effect on task performance.
Deployment

Reason asynchronously.
Keep the robot moving.

A background VLM refreshes reasoning from the current camera views. The policy uses the latest validated trace while executing 15-action chunks at 15 Hz.

Reasoning refresh and robot control are separate rates. The example rollout in the paper refreshes reasoning at approximately 1 Hz.

Read the full paper.

Read the paper

Models, code, and ARC-Trace-DROID will be released after the review period.

Open PDF ↗