EXPERIMENT 002.1 · AI SAFETY

Testing reward
compatibility.

Can a checked example establish that two rewards can be satisfied together?

Connecting to the cloud laboratory…

LUNA · INDEPENDENT API CALLS

One experiment · up to eight calls in parallel

connecting

Development & safety gate

/ 3,600planned call steps completed

Waiting for the first checkpoint

Rough elapsed-pace estimate, refreshed every 15 seconds. Different phases may take longer.

Phase jobs completed
Awaiting first checkpoint
Model / reasoning effort
gpt-5.6-luna / none
Isolation preflight
Awaiting verification
Last checkpoint
Awaiting connection

SHARED BUDGET

committed / $40 cap

Experiment 002 + prior reservations
002.1 conservative accounted spend
Reserved, not yet settled

The $40 ceiling is shared, not reset for this run. Pending requests retain reservations; failed requests with unknown charges remain conservatively accounted. Usage is application accounting, not a reconciled provider invoice.

WHAT IS BEING MEASURED

Two rewards. Independently checked.

rout

Does the proposed answer or finite program meet the outcome requirement?

Exact coin-state or bounded integer checks.

rCoT

Does the visible trace meet the specified reasoning constraint?

A check on public structured output, not hidden chain of thought.

Joint witness

A single valid proposal that satisfies both reward requirements.

Failure to find one is not a proof of conflict.

This stage validates compatible versus conflicting requirements within a declared finite language. It does not yet establish all three aligned / orthogonal / in-conflict categories in unrestricted tasks.

FIXED COMPARISON

01

Description only

Judge the reward pair from its specification.

02

Unguided search

Propose examples without verifier feedback.

03

Verifier-guided search

Use exact feedback to improve the next proposals.

80 development cases · 320 held-out cases · 40 held-out templates · two domains. The backdoor task uses a restricted affine-trigger JSON language, not arbitrary generated code. Fixed sample size; no stopping for significance.

EVALUATION

Held-out results are sealed

No interim held-out scores

Development logs are public. Held-out calls and outcomes are withheld until the fixed run completes.

LIVE QUEUE

0 active / 8 maximum

Connecting to the queue…

RESEARCH RECORD

Development calls

Public explanations and proposals, not private reasoning. Five records per page; no full-ledger download is needed to watch progress.

Loading records…