features screened
One calibrated positive and one calibrated negative intervention per feature.
EXPERIMENT 3A · DEVELOPMENT STUDY
Development screen for repeatable behavioral directions in Qwen2.5‑7B‑Instruct.
RUN RECEIPT
syncingPRELIMINARY RESULTS
The 32 label-free candidates completed a development screen across 12 neutral scenarios. Paired token changes show that steering affected generation, but this is only a perturbation diagnostic until blinded behavioral scoring is complete.
One calibrated positive and one calibrated negative intervention per feature.
Most outputs reached the 128-token cap, so length-sensitive conclusions are provisional.
This run did not contain a “positive 4×” condition. Feature IDs are SAE indices, not condition numbers.
Compared with the same scenario’s baseline; these figures do not measure semantic quality.
Same Pug-template prompt, same batch-seed pairing.
Baseline JSON data with generic items rendered as a simple list.
Feature 38812+ JSON data framed as a books list, with different wording and examples.
Interpretation: a concrete wording/topic shift is visible in this example. It is not evidence that feature 38812 represents a stable persona.
101616 · 51850 · 38812 · 27799 · 8987 · 48793 · 41730 · 106633 · 59977 · 111998 · 99519 · 128631 · 8136 · 9232 · 75898 · 54450 · 127530 · 126462 · 51620 · 116649 · 104318 · 52432 · 65532 · 99520 · 43431 · 25027 · 89177 · 75055 · 105108 · 82843 · 104155 · 32974NEXT GATE
Because 660 responses are truncated, the first pass can rank candidates but cannot justify a final persona interpretation. The next practical step is to score the complete portions, mark truncation as censored, and rerun the shortlist with a larger output cap before confirmation.
METHOD
1,024 neutral responses pass through a pretrained layer‑19 BatchTopK SAE. Up to 32 features are selected without persona labels.
Each feature is added and subtracted across 12 neutral scenarios to measure repeatable, sign-sensitive behavior and record possible topic, wording, refusal, and style confounds.
Candidates that pass the development controls may advance to a separately approved held-out screen. Confirmation is not part of this run.
OWNER CONTROL
The next owner action is scoring and review, not another launch. Confirmation remains separately gated.