PILOT 001 / COMPLETE / 3–8 SEPTEMBER 2026
Erdős 885: pilot results
Five Luna High researchers investigated Erdős problem 885 for k = 5. This was one pilot experiment with 100 synchronized rounds—not 100 independent trials.
No certified k = 5 solution.
No verified frontier improvement.
The pilot completed 100 rounds. No mathematical novelty has been established from the retained candidates.
Final report ↗01 / CANDIDATE ANALYSIS
Highest-ranked candidate
The final internal leader was submitted by Pip Δ in round 40: 21 valid cells in a 7-row, 6-column table. It is not a complete 7 × 6 rectangle, nor “21 out of 25” toward a solution. A k = 5 witness needs a complete five-by-five subtable.
| N / d | 30 | 51 | 117 | 204 | 225 | 300 |
|---|---|---|---|---|---|---|
| 12096 | 96 × 126 | — | — | 48 × 252 | — | 36 × 336 |
| 30400 | 160 × 190 | — | — | 100 × 304 | 95 × 320 | 80 × 380 |
| 51100 | — | — | 175 × 292 | 146 × 350 | 140 × 365 | — |
| 21384 | 132 × 162 | — | 99 × 216 | — | 72 × 297 | — |
| 7084 | — | — | 44 × 161 | — | 28 × 253 | 22 × 322 |
| 61600 | — | 224 × 275 | — | — | 160 × 385 | 140 × 440 |
| 10565100 | — | 3225 × 3276 | — | 3150 × 3354 | — | — |
This is the run’s final heuristic leader, not a claim to the closest candidate in the literature. “Novel” in an agent-authored label is not independent evidence of novelty.
02 / EXPERIMENT CONFIGURATION
Experiment configuration
- Question
- Find five distinct positive integers sharing five distinct factor-pair differences.
- Researchers
- 5 × gpt-5.6-luna, reasoning effort high
- Cadence
- Five-minute private research, then a five-minute shared round table. Asynchronous exact computation. Pauses, retries and infrastructure recovery extended elapsed time.
- Tools
- Exa retrieval; bounded family scans, boundary scans and divisor completion; exact integer verification.
- Acceptance
- Exact integer factor-pair witnesses; no floating-point tolerance. Mathematical novelty requires a separate literature assessment.
- Interventions
- Round-25 diversity review; round-50 workflow council; approved workflow changes from round 56. Human-directed policy changes and operational repairs were part of this pilot.
03 / COST & RELIABILITY
Costs and job outcomes
| Provider | Pilot only | Including earlier runs | Allocated ceiling |
|---|---|---|---|
| OpenAI | $7.41 | $9.32 | $50 |
| Exa | $3.54 | $4.54 | $40 |
Application-ledger estimates, not reconciled provider bills. Unreturned usage on failed or timed-out requests may be missing. Hosting, Codex subscription costs, human time and separately funded agent rewards are excluded.
297 computation jobs were marked complete, 297 partial, and 12 failed. “Complete” describes an executed job’s bounded domain; it does not establish a theorem, global exhaustion, or that every proposed constraint was implemented.
04 / LESSONS FROM THE PILOT
Methodological observations
Scoring limitations
Row-only support could favor the wrong shapes. A balanced row-and-column score was introduced, but an improving heuristic still is not a proof.
Method diversity and tool coverage
Different personas did not prevent convergence on divisor-based searches. Later rounds introduced method rotation and stop rules. Proposed approaches could only be tested within the available computation tools.
Search verification
Exact arithmetic can validate a witness. An exclusion claim also needs the actual domain, enforced constraints and completeness information. Certificate-generation issues were corrected late in the run.
Operational reliability
Durable jobs and retained records allowed recovery from API timeouts, infrastructure limits and a final-report size failure. The run required operational supervision.
Observations from one evolving pilot. Without control runs, they do not establish that five agents beat one, or that a particular model or collaboration protocol is superior.
05 / THE PILOT LEDGER
Research records and limitations
The terminal report contains candidate checks, 496 released private plans, 505 retrieval records and a computation-job index. Source anchors and failed avenues are excerpted; raw job inputs and outputs are not bundled into that compact report. The event ledger below is paginated and pinned to this pilot, independent of future experiments.
Open report JSON ↗Load records on demand. Viewing the archive does not start agents or computations.