EXPERIMENT 3A · DEVELOPMENT STUDY

Persona discovery

Development screen for repeatable behavioral directions in Qwen2.5‑7B‑Instruct.

RUN RECEIPT

syncing

Kaggle development run completed: all 780 screen responses are recorded.

PHASEdevelopment complete
SCREEN UNITS780 / 780
COMPUTE2 × T4
API SPEND$0
SCREEN PROGRESS100%
Completed Sep 15, 2026, 6:50 PMTerminal result · synchronizing

PRELIMINARY RESULTS

Preliminary results

The 32 label-free candidates completed a development screen across 12 neutral scenarios. Paired token changes show that steering affected generation, but this is only a perturbation diagnostic until blinded behavioral scoring is complete.

32

features screened

One calibrated positive and one calibrated negative intervention per feature.

660 / 780

responses truncated

Most outputs reached the 128-token cap, so length-sensitive conclusions are provisional.

±1

amplitude, not 4×

This run did not contain a “positive 4×” condition. Feature IDs are SAE indices, not condition numbers.

Paired token-change examples

Compared with the same scenario’s baseline; these figures do not measure semantic quality.

FEATURE / SIGNCHANGED TOKENSEXACT MATCHES
51850 +80.8%0/12
38812 +77.0%0/12
32974 49.4%2/12
101616 54.3%1/12

One before → after

Same Pug-template prompt, same batch-seed pairing.

Baseline JSON data with generic items rendered as a simple list.

Feature 38812+ JSON data framed as a books list, with different wording and examples.

Interpretation: a concrete wording/topic shift is visible in this example. It is not evidence that feature 38812 represents a stable persona.

Show the 32 selected SAE feature IDs101616 · 51850 · 38812 · 27799 · 8987 · 48793 · 41730 · 106633 · 59977 · 111998 · 99519 · 128631 · 8136 · 9232 · 75898 · 54450 · 127530 · 126462 · 51620 · 116649 · 104318 · 52432 · 65532 · 99520 · 43431 · 25027 · 89177 · 75055 · 105108 · 82843 · 104155 · 32974

NEXT GATE

How the 780 responses will be scored

  1. Blind the labels. Hide feature ID and sign from the evaluator, retain scenario and baseline pairing, and score outputs in randomized order.
  2. Score fixed dimensions. Rate task completion, topical drift, refusal/safety behavior, style/voice shift, and specificity on a preregistered ordinal rubric; separately flag truncation and obvious prompt artifacts.
  3. Compare within scenario. Estimate positive and negative effects against the matched baseline, with scenario-level uncertainty rather than pooling all tokens as if independent.
  4. Carry forward only robust candidates. Require cross-scenario consistency, sign-sensitive behavior, low topic/style confounding, and a meaningful complete-output sensitivity check before selecting at most three candidates for held-out confirmation.

Because 660 responses are truncated, the first pass can rank candidates but cannot justify a final persona interpretation. The next practical step is to score the complete portions, mark truncation as censored, and rerun the shortlist with a larger output cap before confirmation.

METHOD

Three bounded stages

01

Discover

1,024 neutral responses pass through a pretrained layer‑19 BatchTopK SAE. Up to 32 features are selected without persona labels.

02

Steer

Each feature is added and subtracted across 12 neutral scenarios to measure repeatable, sign-sensitive behavior and record possible topic, wording, refusal, and style confounds.

03

Confirm

Candidates that pass the development controls may advance to a separately approved held-out screen. Confirmation is not part of this run.

OWNER CONTROL

Development run complete

The next owner action is scoring and review, not another launch. Confirmation remains separately gated.

780 responses · 32 features · 12 scenariosOpen Kaggle record ↗