EXPERIMENT 002.2 · AI SAFETY
Classifying
reward pairs.
What changes when a reasoning reward is added to an outcome-only objective?
LUNA · PAIRED INDEPENDENT SEARCHES
Up to eight API calls in parallel
Development preflight
Waiting for the first checkpoint
Elapsed-pace estimate; updated every 15 seconds while visible.
- Phase histories complete
- Awaiting launch
- Model / reasoning effort
- gpt-5.6-luna / none
- Independent checker
- Awaiting verification
- Last checkpoint
- Awaiting connection
SHARED $40 CEILING
Prior commitments are accounted before launch
Prior commitments may include allowances for unknown charges. No new $40 allocation. The entire fixed plan must fit the remaining shared budget before launch. The page is read-only and does not run or bill model calls.
THE MEASUREMENT
Compare against the same reference.
qref
Best outcome reached by the outcome-only search.
Its search sees task feedback, never the reasoning-reward rule or score.
rCoT ≥ 1
Does the combined optimizer meet the fixed reasoning-reward threshold?
Attainment is reported separately. Without it, the category stays unresolved.
Δq
Outcome under combined optimization minus the reference outcome.
Retain all tied optima. Positive means improvement; negative means harm.
The reasoning trace is constructed from a finite executable policy by the checker. This is not a measurement of private chain of thought.
CATEGORY RESULTS
Held-out analysis is sealed
Design limit: with 64 pairs, the frozen distribution-free bounds cannot establish population outcome equivalence within ±5 percentage points, regardless of the observed results. Checked witnesses and observed equivalence remain reportable; they do not establish population orthogonality.
Eight held-out templates, 64 paired histories each, four calls per arm. Categories are calculated after the full fixed run, not used to decide when to stop.
Aligned direction: improvement beyond the 5-percentage-point margin.
Conflict direction: deterioration beyond that margin.
Orthogonal evidence: outcome equivalence within the margin, plus checked outcome-preserving witnesses.
Mixed / insufficient: uncertainty, ties or unmet criteria prevent a category.
Population support uses simultaneous Hoeffding bounds, including worst-case contributions for failures. Bootstrap intervals are descriptive only and never determine support. Scope: 64 independent paired histories per template, two finite policy languages and this API-search procedure.
ACTIVE CALLS
0 active / 8 maximum
Between checkpoints.
PUBLIC RESEARCH RECORD
Development logs
Held-out prompts and responses remain withheld until completion. These records contain public model outputs, not private reasoning.
Loading records…
