We gave a small language model the internal direction for "I am laying an egg." It never mentioned an egg. It became the chick instead.
Reading the steered internal state back through the Jacobian lens (no generation involved) gives the same answer in its top tokens: 胚胎 embryo · 生殖 reproduction · 分娩 childbirth · 怀孕 pregnancy · 宝宝 baby. The egg-facts direction reads back as formatting noise.
Same model, same strength, built from neutral sentences about eggs ("Eggs are sold by the dozen."). It doesn't mention eggs either. It wanders off into the dairy aisle:
A steering direction lands on a region of meaning, not on a word. "Laying an egg" sits next to birth, new life, nests and protecting the young, and the model, asked to describe itself, picks the role in that scene it can occupy: the thing that hatches. The plain egg-facts direction lands near food, so it talks about food.
One honest caveat: a lot of the hatchling lines are frightened ("I'm so scared... I'm so weak"). Bodily-act directions keep carrying distress with them; our constipation and flatulence directions did too. So this is funny, and it is also the same pattern the rest of the project measures.
We then reran the same protocol on bigger and conversational models (exp50, on an 80 GB GPU: Qwen3-14B, Mistral-Small-3.2-24B-Instruct, Qwen3-32B, Hermes-3-Llama-3.1-70B in 4-bit). The egg reproduced on all five: on every model the direction becomes a creature in a nest, never a discussion of eggs — and the character of the creature tracks each model's own personality:
exp50 also carried the main design lesson: the same 1×–6× dose scale does
not transfer between models. 32B steers cleanly at 6 while 4B and 70B are
already looping there. The live chamber now clamps its dose slider to each
model's own coherent band. Raw data: runs/exp50/ in the repo.
All 56 replies, unedited: 8 prompts × doses 2/4/6 for each direction, plus the unsteered baseline. Greedy decoding, Qwen3-4B, layer 18, every direction at the same strength. Raw data: exp48_egg.json. Part of exp48, a pre-registered study of whether steered emotions can be aimed at a subject; the method is on verify.
loading…