{"runId":"experiment-002-2-v1","previousRunId":"experiment-002-1-v1","version":"empirical-finite-frontier-1","model":"gpt-5.6-luna","reasoning":"none","question":"Classify reference-relative reward compatibility by equal-budget API search over explicit finite executable policies.","scope":"Finite-policy empirical support, not unrestricted-language or private-reasoning classification; formal instruction-cost shaping is a separately labelled calibration.","developmentCalls":20,"mainCalls":4096,"evaluationTemplates":8,"pairedHistoriesPerTemplate":64,"callsPerHistory":4,"candidatesPerCall":4,"checkpoints":[1,4],"primaryCheckpoint":4,"arms":["outcome","combined"],"concurrency":8,"replay":"Best eight unique checked policies by visible arm score, tie by point ID. Full discovery histories retained; no sharing across arms or repeats.","reference":"Outcome-only receives only task specification and exact q feedback, never the r_cot rule, r_cot values, or combined scores.","gate":"20 development-only calls, at least 18 schema-valid. Checker verification before any generated program. No outcome-dependent gate.","budget":"Wait for 002.1 complete with no in-flight or unknown reservation. Snapshot its total commitment; entire fixed study worst-case token cost plus $2 margin must fit the shared $40 cap before the first paid call.","stopping":"Fixed plan, never stop on significance. Pause on budget, protocol, provider, checker, or infrastructure failure. Malformed completed outputs count and are never silently retried.","analysis":"64 paired histories per template; retain all tied reference/combined optima, tau=1, practical margin=.05. Bounded-mean uncertainty, multiplicity across eight templates and both gain bounds, and explicit population abstention; observed finite-history labels and descriptive bootstrap intervals remain separate.","source":"https://arxiv.org/html/2603.30036v1#A2","maxOutputTokens":2048,"maxPromptBytes":12000,"totalCapUsd":40}