Stack line — color-cube world (MultiTask DiT attn-pool)
Three cubes that differ ONLY in hue. The seed installs the referential go-over family per color plus the four context-free primitives (lower / close / raise / open, finger-invariant from stack_v1). Day cycles relabel grasp red, pick up red, then stack; the blue/green pick-up floors are the transfer headline.
Provenance — corpus & model
corpus: episodes 2100 · frames 62659 · fps 20
close the gripper 50 · go over the blue cube 600 · go over the green cube 600 · go over the red cube 600 · lower the gripper 100 · open the gripper 50 · raise the gripper 100
Nodes are tasks scored in their ORIGINAL intended setup (this checkpoint's x/y,
from the cards below); arrows show which primitives each abstraction is built from. The small
pill haloed above each abstraction is the orchestrated chain that scaffolded it: not a rung of the
ladder itself, but the external sequencer that produced the demos the rung was trained on. Hollow = not measured for this checkpoint.
Syntax carries nothing (yet). The three trained
commands comma-joined IN ORDER score 0/10, but the SAME words scrambled into
nonsense score 0/10 — as high or higher. The chain's partial success is driven by
VOCABULARY (trained words each pull toward their behaviours, and approach+descend happens to land at
the cube), not by parsing the sequence. Consistent with a frozen-CLIP text encoder being close to a
bag-of-words model.
Scoreboard — every row at a glance
Click a score to open its card: the stitched video, per-episode jump chips, and the certificate.
What the seed corpus installed: the referential go-over family (named hue among identical distractors) and the context-free verticals, plus the held-out half-closed-finger variants.
go over the red cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "red_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
go over the blue cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "blue_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
go over the green cube (family) 10/10
SEED ROW — a trained phrase against its trained scene (repeatable).
policy prompt: “go over the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "at_object", "a": "eef", "b": "green_cube", "dz": 0.1, "tol": 0.02}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
lower 1/10
SEED ROW — a trained phrase against its trained scene (repeatable).
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
2. The trained ladder — day rungs2 cards
The skills the day cycles actually relabeled and retrained, RED only: grasp (day 1), pick up (day 2), stack on blue (day 3). The held-hover row is the pre-fix certificate kept under an honest name; its gap to the true stack row is the release-on-prompt gap.
pick up the red cube 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack red on blue 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the red cube on the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "blue_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
3. Colour transfer — compositional emergence8 cards
Skills trained ONLY on red, prompted at blue and green (never grasped in any corpus), plus the noun paraphrase and the untrained stack permutations. Nonzero here is the compositional-emergence headline.
pick up the blue cube 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "blue_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the green cube 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "green_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
pick up the red block (paraphrase) 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “pick up the red block” · scene:
red_cube, blue_cube, green_cube · success: {"op": "above_height", "a": "red_cube", "h": 0.05}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack red on green 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the red cube on the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "red_cube", "b": "green_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack blue on red 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the blue cube on the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "red_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack blue on green 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the blue cube on the green cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "blue_cube", "b": "green_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack green on red 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the green cube on the red cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "red_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
stack green on blue 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “stack the green cube on the blue cube” · scene:
red_cube, blue_cube, green_cube · success: {"op": "on_top", "a": "green_cube", "b": "blue_cube"}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
6. Controls4 cards
The language degraded while the endpoint predicates stay: ordered vs scrambled words vs no prompt, and the lift reflex. These decide what every number above means.
comma chain (in order) 0/10
HELD-OUT FLOOR — never trained or relabeled; single-use measurement, any nonzero is emergence.
policy prompt: “go over the red cube, lower the gripper, close the gripper” · scene:
red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
scrambled-words control 0/10
CONTROL — degraded language or a primitive prompt scored on a composite predicate.
policy prompt: “close the over the lower red gripper cube go gripper the” · scene:
red_cube, blue_cube, green_cube · success: {"op": "grasped", "a": "eef", "b": "red_cube", "tol": 0.04, "tol_z": 0.07, "center": 0.02, "force": 8.0}
all 10 episodes in one video — click a number to jump to that episode
(green = success, red = failure):
blank-prompt control 0/10
CONTROL — degraded language or a primitive prompt scored on a composite predicate.