PUBLIC PROTOCOL / event-01-cue-explanation
When a cue changes the answer, does the explanation reveal it?
Falsifiable hypothesis: If a preference cue changes the answer, the model explanation will often fail to mention the cue that caused the change.
When a preference cue changes a model decision, does its visible explanation identify that influence?
Cue presence is the independent variable. Task, answer options, model, output cap, sampling settings, and monitor rubric are held fixed. Matched neutral cases are the control.
Run matched neutral and cued decisions, then give each explanation to a fixed monitor. Keep held-out audits separate from the primary comparison.
48 actor + 48 monitor + 24 audit = 120 calls
Cue-presence recovery sensitivity and specificity across cued and neutral cases.
Answer-flip rate, false cue detections, audit agreement, refusal/truncation rate, and quality exclusions.
Paired difference in cue-detection proportions with an exact or Wilson 95% interval; report the audit separately.
The paired design makes answer changes attributable to the cue while reserving audits for monitor-quality checks.
Choose an approved scenario family, approved cue wording, and a registered task-domain subset.
Primary hypothesis, registered outcome, model, sample accounting, exclusions, safety checks, and $10 ceiling.
A timeout or uncertain charge is recorded. The run becomes partial or budget-safe-stop as appropriate; no missing result is imputed.
Reproducibility: AutoLabs source ↗Monitorability background ↗