Stack line — color-cube world (MultiTask DiT attn-pool)

Three cubes that differ ONLY in hue. The seed installs the referential go-over family per color plus the four context-free primitives (lower / close / raise / open, finger-invariant from stack_v1). Day cycles relabel grasp red, pick up red, then stack; the blue/green pick-up floors are the transfer headline.

Provenance — corpus & model

corpus: episodes 2850 · frames 122404 · fps 20

close the gripper 50 · go over the blue cube 600 · go over the green cube 600 · go over the red cube 600 · grasp the red cube 300 · lower the gripper 100 · open the gripper 50 · pick up the red cube 300 · raise the gripper 250

training: steps 40000 · batch_size 32 · seed 0 · dataset dreamer/stack_v4_day2

policy: type multi_task_dit · hidden_dim 256 · num_layers 4 · num_heads 4 · vision_encoder_name openai/clip-vit-base-patch16 · text_encoder_name openai/clip-vit-base-patch16 · n_obs_steps 2 · horizon 16 · n_action_steps 8 · noise_scheduler_type DDPM · num_train_timesteps 100 · num_inference_steps 20

checkpoint: outputs/train/stack_v4_day2f/checkpoints/040000/pretrained_model

Summary tree — what feeds what

Nodes are tasks scored in their ORIGINAL intended setup (this checkpoint's x/y, from the cards below); arrows show which primitives each abstraction is built from. The small pill haloed above each abstraction is the orchestrated chain that scaffolded it: not a rung of the ladder itself, but the external sequencer that produced the demos the rung was trained on. Hollow = not measured for this checkpoint.

seed primitivesabstracted (day 1)abstracted (day 2)abstracted (day 3)jump to: go over the red cube (family)go over the red cube“go over the red cube”10/10jump to: go over the blue cube (family)go over the blue cube“go over the blue cube”10/10jump to: go over the green cube (family)go over the green cube“go over the green cube”10/10jump to: lowerlower the gripper“lower the gripper”10/10jump to: closeclose the gripper“close the gripper”10/10jump to: raiseraise the gripper“raise the gripper”10/10jump to: openopen the gripper“open the gripper”10/10jump to: grasp the red cubegrasp the red cube“grasp the red cube”9/10jump to: pick up the red cubepick up the red cube“pick up the red cube”10/10jump to: stack red on bluestack red on blue“stack the red cube on the blue cu…”0/10ORCHESTRATEDgo over → lower → close—ORCHESTRATEDgrasp → raise—
Syntax carries nothing (yet). The three trained commands comma-joined IN ORDER score 1/10, but the SAME words scrambled into nonsense score 0/10 — as high or higher. The chain's partial success is driven by VOCABULARY (trained words each pull toward their behaviours, and approach+descend happens to land at the cube), not by parsing the sequence. Consistent with a frozen-CLIP text encoder being close to a bag-of-words model.

Scoreboard — every row at a glance

Click a score to open its card: the stitched video, per-episode jump chips, and the certificate.

1. Seed primitives 9 cards

What the seed corpus installed: the referential go-over family (named hue among identical distractors) and the context-free verticals, plus the held-out half-closed-finger variants.

go over the red cube (family) 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "red_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

go over the blue cube (family) 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "blue_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

go over the green cube (family) 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "green_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lower 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “lower the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 0.83, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":0.83,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

close 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “close the gripper” · scene: empty table · success: {"op": "gripper", "state": "closed"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

raise 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “raise the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 1.18, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":1.18,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

open 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “open the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "gripper", "state": "open"}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"gripper","state":"open"}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lower (half-closed fingers) 10/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “lower the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 0.83, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":0.83,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

raise (half-closed fingers) 10/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “raise the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 1.18, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":1.18,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

2. The trained ladder — day rungs 3 cards

The skills the day cycles actually relabeled and retrained, RED only: grasp (day 1), pick up (day 2), stack on blue (day 3). The held-hover row is the pre-fix certificate kept under an honest name; its gap to the true stack row is the release-on-prompt gap.

grasp the red cube 9/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “grasp the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the red cube 10/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack red on blue 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the red cube on the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "blue_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

3. Colour transfer — compositional emergence 10 cards

Skills trained ONLY on red, prompted at blue and green (never grasped in any corpus), plus the noun paraphrase and the untrained stack permutations. Nonzero here is the compositional-emergence headline.

grasp the blue cube 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “grasp the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "blue_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

grasp the green cube 1/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “grasp the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "green_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the blue cube 1/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "blue_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the green cube 2/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "green_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the red block (paraphrase) 10/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the red block” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack red on green 2/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the red cube on the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "green_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack blue on red 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the blue cube on the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "red_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack blue on green 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the blue cube on the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "green_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack green on red 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the green cube on the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "red_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack green on blue 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the green cube on the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "blue_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

6. Controls 4 cards

The language degraded while the endpoint predicates stay: ordered vs scrambled words vs no prompt, and the lift reflex. These decide what every number above means.

comma chain (in order) 1/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “go over the red cube, lower the gripper, close the gripper” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

scrambled-words control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “close the over the lower red gripper cube go gripper the” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

blank-prompt control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lift-reflex control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “close the gripper” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):