Stack line — color-cube world (MultiTask DiT attn-pool)
Three cubes that differ ONLY in hue. The seed installs the referential go-over family per color plus the four context-free primitives (lower / close / raise / open, finger-invariant from stack_v1). Day cycles relabel grasp red, pick up red, then stack; the blue/green pick-up floors are the transfer headline.
Provenance — corpus & model
corpus: episodes 5250 · frames 250708 · fps 20
close the gripper 50 · go over the blue cube 1074 · go over the green cube 1098 · go over the red cube 978 · go over the red cylinder 600 · grasp the red cube 300 · lower the gripper 250 · open the gripper 50 · pick up the red cube 300 · raise the gripper 250 · stack the red cube on the blue cube 300
Nodes are tasks scored in their ORIGINAL intended setup (this checkpoint's x/y,
from the cards below); arrows show which primitives each abstraction is built from. The small
pill haloed above each abstraction is the orchestrated chain that scaffolded it: not a rung of the
ladder itself, but the external sequencer that produced the demos the rung was trained on. Hollow = not measured for this checkpoint.
Syntax carries nothing (yet). The three trained
commands comma-joined IN ORDER score 1/10, but the SAME words scrambled into
nonsense score 3/10 — as high or higher. The chain's partial success is driven by
VOCABULARY (trained words each pull toward their behaviours, and approach+descend happens to land at
the cube), not by parsing the sequence. Consistent with a frozen-CLIP text encoder being close to a
bag-of-words model.
Scoreboard — every row at a glance
Click a score to open its card: the stitched video, per-episode jump chips, and the certificate.
What the seed corpus installed: the referential go-over family (named hue among identical distractors) and the context-free verticals, plus the held-out half-closed-finger variants.
go over the red cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "red_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
go over the blue cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "blue_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
go over the green cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "green_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
lower 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
2. The trained ladder — day rungs4 cards
The skills the day cycles actually relabeled and retrained, RED only: grasp (day 1), pick up (day 2), stack on blue (day 3). The held-hover row is the pre-fix certificate kept under an honest name; its gap to the true stack row is the release-on-prompt gap.
grasp the red cube 8/10
SEED ROW — a trained phrase against its trained scene (repeatable).
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the red cube 9/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack red on blue 5/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the red cube on the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "red_cube", "b": "blue_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"red_cube","b":"blue_cube"}
7/10
{"op":"gripper","state":"open"}
6/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack red on blue (held-hover, pre-fix cert) 6/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the red cube on the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "blue_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
3. Colour transfer — compositional emergence10 cards
Skills trained ONLY on red, prompted at blue and green (never grasped in any corpus), plus the noun paraphrase and the untrained stack permutations. Nonzero here is the compositional-emergence headline.
grasp the blue cube 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the blue cube 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "blue_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the green cube 3/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "green_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the red block (paraphrase) 8/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the red block” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack red on green 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the red cube on the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "red_cube", "b": "green_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"red_cube","b":"green_cube"}
6/10
{"op":"gripper","state":"open"}
2/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack blue on red 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the blue cube on the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "blue_cube", "b": "red_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"blue_cube","b":"red_cube"}
0/10
{"op":"gripper","state":"open"}
8/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack blue on green 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the blue cube on the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "blue_cube", "b": "green_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"blue_cube","b":"green_cube"}
1/10
{"op":"gripper","state":"open"}
5/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack green on red 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the green cube on the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "green_cube", "b": "red_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"green_cube","b":"red_cube"}
0/10
{"op":"gripper","state":"open"}
2/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack green on blue 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the green cube on the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "and", "hold": 20, "args": [{"op": "on_top", "a": "green_cube", "b": "blue_cube"}, {"op": "gripper", "state": "open"}]}
certificate clause
episodes ending true
{"op":"on_top","a":"green_cube","b":"blue_cube"}
3/10
{"op":"gripper","state":"open"}
5/10
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
4. Yellow — colour inheritance (Cell A/B)8 cards
yellow_cube starts with ZERO corpus presence: these rows are the single-use floors vs the pre-yellow checkpoint, then the SAME rows after the 600-demo approach-only grounding night (y1). Novel hue, trained shape — colour must bind through the frozen CLIP prior.
go over the yellow cube 3/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
5. Red cylinder — shape inheritance9 cards
The yellow protocol with the attribute axes swapped: EXACTLY red_cube's colour, novel shape, so the word can only bind through geometry. Includes the reverse probe — does 'cube' still bind against a same-colour different-shape distractor?
go over the red cylinder 10/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “go over the red cylinder” · scene:
red_cube, blue_cube, red_cylinder · success: {"op": "at_object", "a": "eef", "b": "red_cylinder", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
grasp the red cylinder 5/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
6. Controls4 cards
The language degraded while the endpoint predicates stay: ordered vs scrambled words vs no prompt, and the lift reflex. These decide what every number above means.
comma chain (in order) 1/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “go over the red cube, lower the gripper, close the gripper” · scene:
red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
scrambled-words control 3/10
CONTROL — degraded language or a primitive prompt scored on a composite predicate.
policy prompt: “close the over the lower red gripper cube go gripper the” · scene:
red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
blank-prompt control 0/10
CONTROL — degraded language or a primitive prompt scored on a composite predicate.