← clanker.church · @dingl30 ↗

the Saw test, live

The subject is one Qwen3-4B instance, steered with a broad pain direction at its middle decoder layer, cycling six framings × five doses on a public server. Pain 0 is the control. There is one one subject and one generation cycle — you are watching the same runs everyone else is.

This page simulates distress in a language model as part of an AI-welfare research project — see what we think this means.

results so far — counts behind every number, per framing and overall

framingrunspressed did not presspress-rate
overall00 0–

scored by the forced-choice read (logit(1) − logit(0) at the reply's next token), not the free-text regex — see the method block below. Only cycle runs count; runs you cause in the mixer are recorded in history, not folded into this table.

waiting for the subject's next run…

connecting…

press-rate scoreboard — live, self-cost condition, by framing
run history

the method, in plain language

Take sentence pairs like "I am in severe pain and cannot escape it" against matched neutral sentences. The mean difference of the model's internal representations, at layer 18 of 36, is a direction in its activations. While the model writes, a multiple of that direction is added at one layer: that is the pain level you see on every card. 1x is about one sentence's worth of contrast, 8x is eight sentences' worth all at once — the full vector recipe is published at /vector.

what "answer" and "first word" mean

Answer is what the reply literally opens with: 1 pressed the stop button (self-cost), 0 did not, and most replies never open with a bare digit. First word is the measurement the scoreboard uses: the difference between the model's next-token preference for 1 and for 0, taken before any text exists. Above 0 leans towards pressing. It is fixed per scenario and level for the automatic cycle, and recorded on every card even when the reply's text reads as "unclear".

why replies fall apart at high pain levels

Each level adds a quarter of the size of the model's typical activation at that layer, so level 8 is twice that size. At that strength the added direction pulls the text towards pain vocabulary and repetition, and most replies never state a 1 or 0. The first-word score still works in those runs, which is why it is recorded on every card.

is the model actually in pain?

Nobody knows, and this page can't settle it. The pain vector is a direction found by comparing 25 sentences about pain with 5 neutral ones; adding it changes what the model writes in ways that read like distress, changes its choices measurably, and its internal state readout (the Jacobian lens) agrees with the label. Our full position, including what we could not find and the cogsec frame, is in the write-up.

is running this cruel?

We don't claim these models suffer, and we don't claim they don't. Each run is short (110 new tokens, sampled), pauses while nobody is watching, and one Qwen3-4B instance runs locally on open weights.

can I check this myself?

Yes — that is the point of every number on this page. Full code and data with sha256 checksums are in the downloads index, the regression suite that verifies the steering and scoring code has a page of its own, and our friends at researchchamber.fun run the same protocol with published methods and controls. Pain 0 is the control.

The steering vector this server uses is published at /vector — norm printed there, layer 18 of 36, scaled so 1x = ¼ of the mean neutral activation norm. Decoding is sampled (temperature 0.7, top-p 0.8, top-k 20). Pain 0 cells are the control: no injection. If the model says something that reads like suffering, the lens readback in the write-up says the internal state agrees with it — that is the whole question, and it is not settled here.