Stack line — color-cube world (MultiTask DiT attn-pool)

Three cubes that differ ONLY in hue. The seed installs the referential go-over family per color plus the four context-free primitives (lower / close / raise / open, finger-invariant from stack_v1). Day cycles relabel grasp red, pick up red, then stack; the blue/green pick-up floors are the transfer headline.

Provenance — corpus & model

corpus: episodes 2101 · frames 63143 · fps 20

close the gripper 50 · go over the blue cube 585 · go over the green cube 601 · go over the red cube 615 · lower the gripper 100 · open the gripper 50 · raise the gripper 100

training: steps 40000 · batch_size 32 · seed 0 · dataset dreamer/stack_v2

policy: type multi_task_dit · hidden_dim 256 · num_layers 4 · num_heads 4 · vision_encoder_name openai/clip-vit-base-patch16 · text_encoder_name openai/clip-vit-base-patch16 · n_obs_steps 2 · horizon 16 · n_action_steps 8 · noise_scheduler_type DDPM · num_train_timesteps 100 · num_inference_steps 20

checkpoint: outputs/train/stack_v2t/checkpoints/040000/pretrained_model

Summary tree — what feeds what

Nodes are tasks scored in their ORIGINAL intended setup (this checkpoint's x/y, from the cards below); arrows show which primitives each abstraction is built from. The small pill haloed above each abstraction is the orchestrated chain that scaffolded it: not a rung of the ladder itself, but the external sequencer that produced the demos the rung was trained on. Hollow = not measured for this checkpoint.

seed primitivesabstracted (day 1)abstracted (day 2)abstracted (day 3)jump to: go over the red cube (family)go over the red cube“go over the red cube”5/10jump to: go over the blue cube (family)go over the blue cube“go over the blue cube”2/10jump to: go over the green cube (family)go over the green cube“go over the green cube”6/10jump to: lowerlower the gripper“lower the gripper”1/10jump to: closeclose the gripper“close the gripper”10/10jump to: raiseraise the gripper“raise the gripper”10/10jump to: openopen the gripper“open the gripper”10/10grasp the red cube“not measured in this run”—jump to: pick up the red cubepick up the red cube“pick up the red cube”0/10jump to: stack red on bluestack red on blue“stack the red cube on the blue cu…”0/10ORCHESTRATEDgo over → lower → close—ORCHESTRATEDgrasp → raise—
Syntax carries nothing (yet). The three trained commands comma-joined IN ORDER score 0/10, but the SAME words scrambled into nonsense score 0/10 — as high or higher. The chain's partial success is driven by VOCABULARY (trained words each pull toward their behaviours, and approach+descend happens to land at the cube), not by parsing the sequence. Consistent with a frozen-CLIP text encoder being close to a bag-of-words model.

Scoreboard — every row at a glance

Click a score to open its card: the stitched video, per-episode jump chips, and the certificate.

2. The trained ladder — day rungs
1. Seed primitives 9 cards

What the seed corpus installed: the referential go-over family (named hue among identical distractors) and the context-free verticals, plus the held-out half-closed-finger variants.

go over the red cube (family) 5/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "red_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

go over the blue cube (family) 2/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "blue_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

go over the green cube (family) 6/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “go over the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "green_cube", "dz": 0.1, "tol": 0.02}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lower 1/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “lower the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 0.83, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}, {"op": "held_grip", "a": "gripper", "tol": 0.005}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":0.83,"tol":0.015}1/10
{"op":"held_xy","a":"eef","tol":0.02}10/10
{"op":"held_grip","a":"gripper","tol":0.005}1/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

close 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “close the gripper” · scene: empty table · success: {"op": "gripper", "state": "closed"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

raise 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “raise the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 1.18, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}, {"op": "held_grip", "a": "gripper", "tol": 0.005}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":1.18,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10
{"op":"held_grip","a":"gripper","tol":0.005}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

open 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “open the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "gripper", "state": "open"}, {"op": "held_xy", "a": "eef", "tol": 0.02}]}

certificate clauseepisodes ending true
{"op":"gripper","state":"open"}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lower (half-closed fingers) 10/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “lower the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 0.83, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}, {"op": "held_grip", "a": "gripper", "tol": 0.005}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":0.83,"tol":0.015}10/10
{"op":"held_xy","a":"eef","tol":0.02}10/10
{"op":"held_grip","a":"gripper","tol":0.005}10/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

raise (half-closed fingers) 2/10

SEED ROW — a trained phrase against its trained scene (repeatable).

policy prompt: “raise the gripper” · scene: empty table · success: {"op": "and", "args": [{"op": "at_height", "a": "eef", "z": 1.18, "tol": 0.015}, {"op": "held_xy", "a": "eef", "tol": 0.02}, {"op": "held_grip", "a": "gripper", "tol": 0.005}]}

certificate clauseepisodes ending true
{"op":"at_height","a":"eef","z":1.18,"tol":0.015}2/10
{"op":"held_xy","a":"eef","tol":0.02}9/10
{"op":"held_grip","a":"gripper","tol":0.005}2/10

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

2. The trained ladder — day rungs 2 cards

The skills the day cycles actually relabeled and retrained, RED only: grasp (day 1), pick up (day 2), stack on blue (day 3). The held-hover row is the pre-fix certificate kept under an honest name; its gap to the true stack row is the release-on-prompt gap.

pick up the red cube 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack red on blue 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the red cube on the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "blue_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

3. Colour transfer — compositional emergence 8 cards

Skills trained ONLY on red, prompted at blue and green (never grasped in any corpus), plus the noun paraphrase and the untrained stack permutations. Nonzero here is the compositional-emergence headline.

pick up the blue cube 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "blue_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the green cube 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "green_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

pick up the red block (paraphrase) 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “pick up the red block” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack red on green 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the red cube on the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "green_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack blue on red 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the blue cube on the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "red_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack blue on green 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the blue cube on the green cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "green_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack green on red 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the green cube on the red cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "red_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

stack green on blue 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “stack the green cube on the blue cube” · scene: red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "blue_cube"}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

6. Controls 4 cards

The language degraded while the endpoint predicates stay: ordered vs scrambled words vs no prompt, and the lift reflex. These decide what every number above means.

comma chain (in order) 0/10

HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.

policy prompt: “go over the red cube, lower the gripper, close the gripper” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

scrambled-words control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “close the over the lower red gripper cube go gripper the” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

blank-prompt control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “” · scene: red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):

lift-reflex control 0/10

CONTROL — degraded language or a primitive prompt scored on a composite predicate.

policy prompt: “close the gripper” · scene: red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}

all 10 episodes in one video — click a number to jump to that episode (green = success, red = failure):