-
gate_min_dim/gate_max_dimreserved keys inStrategySpecheuristic kwargs — the arm is only instantiated when the problem dimension clears the gate; keys are stripped before the heuristic constructor. First deliverable of GOAL §5.1(d). - NLSHADE_LBC gated to d≥5 in
Rewarding_Restart— ships the #298-measured d5 gain (+0.0080 [+0.0007,+0.0154]) without the d2 loss (−0.0241). Today's independent 12-seed standard A/B: d5 +0.0062 [−0.0001,+0.0125], d2 flat, control flat; pooled with #298's d5 row: +0.0070 [+0.0027,+0.0112]. - Re-measure the CMA-ES arm at d5 (GOAL §5.2) with the 12-seed standard instrument — if it splits like NLSHADE_LBC, the same gate ships it.
- Watch the nightly on the gated spec — the d5 slice of the
widened quick battery now exercises the LBC arm nightly;
drop_heuristic NLSHADE_LBCaccepts at d2-only regimes would be evidence the gate threshold is wrong, not that the arm is bad.
-
GOAL.md§2 / §5.1 corrected — the "instance-family generalization" research item was built on a metric-unit bug and is retracted; the slot now holds regime-conditional strategy selection. §5.3 marked shipped. §4 Diagnose points atper_cell. -
planning/LOOP_DIAGNOSIS_2026-08-11.md— full 34-night audit so no future session re-derives it. - Rotated into
planning/done/— both ledgers, the summary, and the pre-2026-07-30 halves ofTODO.md(3720 → 332) andSELF_IMPROVEMENT_LOG.md(11126 → 1117). Nothing deleted; the bandit still primes from archives and metric inference on the new filenames is verified. - Watch the first post-rotation nights — the live aocc ledger is empty, so codify-scan has no cross-night evidence until ~2 nights have run on distinct base seeds. Expect a quiet week; that is correct behaviour, not a regression.
-
--cell-by dim, on in the nightly — the objective is no longer a scalar. Per-cell delta + CI reported every iteration; a cell blocks only when it is both worse than-eps_cell_regress(0.01) and has its whole CI below zero. -
per_cell/blocking_cellpersisted in the ledger — the record a future codify-scan needs to propose a dimension-gated arm instead of an unconditional one. - Teach codify-scan to read
per_cell— it still pools one scalar delta per night, so cell-conditional evidence is recorded but not yet actionable. This is the step that turns the #298 d5 gain into something shippable. - Budget-phase cells —
IOHRunRecordalready carriestrace_evals/trace_fx, but AOCC would have to be recomputed on trajectory slices rather than read off the final value (a change to the metric path, not the decision path). #298's evidence says this matters: the same arm leaned positive at 2-D×200 evals and negative at 2-D×1000 — a budget effect, not a dimension one. - Dimension-gated arm activation — the original §4.3 follow-up.
Shipped 2026-08-12 (see the entry above):
gate_min_dim/gate_max_diminStrategySpec, NLSHADE_LBC gated to d≥5, measured. CMA-ES at d5 remains the open follow-up.
-
--accept-stat rank— one-sided Wilcoxon signed-rank on the per-pair deltas shifted byeps_accept, replacing both thedelta > eps_acceptandci_low > 0conditions. Graduates GOAL §5.3. - Hodges-Lehmann reported as
rank_delta;deltastays the mean under both rules so the ledger series is continuous. - A/B the rank rule against the mean rule, then decide — NOT
enabled nightly on purpose: tonight already changes the accept
regime three ways (seed rotation, eps 0.005→0.0125, d5 slice).
A fourth simultaneous change would make the ledger
uninterpretable. Run it via
workflow_dispatchon a few nights once the new instrument has a baseline, compare accept rates and codify survival, and only then consider flipping the default. - Guard the n<5 floor — a rank accept is impossible below 5
shared pairs (min p is
2**-n). Currently only documented and tested; a loud warning when a configured battery cannot clear it would be better than a rule that silently never fires.
-
--aocc-extra-dims 5in the nightly — the quick preset'sdims=(2,)becomes(2, 5)for every measurement leg. Closes the blind spot behind the JSO d5 add (08-02) and the NLSHADE_LBC per-dim split (08-11). Cost 8.9s → 11.9s (1.34×). -
with_extra_dimscomposes, never edits the frozen presets; name gains a+d5suffix. - The ledger score level shifts with this ship (d2 ~0.369 →
(d2,d5) ~0.309).
aocc_extra_dimsis recorded per iteration; codify-scan andsummaryshould group by it before pooling. - Re-measure the CMA-ES arm (GOAL §5.2) at d5 with the 12-seed standard instrument — if it splits like NLSHADE_LBC did, one dimension-gating mechanism ships both arms.
Loop measurement fidelity: hold-out metric bug, seed rotation, sync-eval — 2026-08-11 (second session)
- Hold-out leg measured
composite_scoreon AOCC runs — the 8.5× "instance-family generalization gap" (GOAL.md§2 / §5.1's top research priority) was a unit mismatch. Real gap after the fix: 0.3383 training vs 0.3342 hold-out. - Confirm gate now crosses a base seed under
--metric aocc— themetric != "aocc"exclusion meant 0/72 accepts had ever tested a second instance family. - Nightly rotates
--base-seedover 7 values — all 952 prior ledger records werebase_seed=42, so codify's "k≥2 distinct nights" was counting one instance draw k times (hit rate 0/5). -
--sync-evalreachable from the loop and on by default in the nightly (noise sd 0.0101 → 0.0063; measured cost ~0). - eps_accept recalibrated 0.005 → 0.0125 on the aocc branch (0.5σ → 2σ); relax floor 0.001 → 0.006.
- Re-earn the codify backlog — cross-night evidence banked
before 2026-08-11 is single-draw. Let the rotated-seed ledger
accumulate ≥2 nights on distinct base seeds before trusting any
slot, and consider making codify-scan group by
base_seedrather than by night. - Group cross-night pooling by
sync_eval— the field is now recorded but codify-scan does not yet split on it; the first nights after this ship straddle the boundary.
-
add_heuristic NLSHADE_LBConRewarding_Restartmeasured and rejected (negative result, PR #298) — 12-seed paired A/B with--sync-evalboth sides: quick +0.0092 (noise), standard −0.0080 (lean-negative), but the per-dim split is CI-significant both ways: d5 +0.0080 [+0.0007,+0.0154], d2 −0.0241 [−0.0401,−0.0080]. Reverted in-branch; see the 2026-08-11 log entry. - Dimension/budget-gated arm activation — the d5 gain is real; the shippable form is a structural mix that activates population arms (NLSHADE_LBC, CMA-ES) only when dim/budget clears a threshold, or an anytime schedule that phases them out at long 2-D budgets. Re-measure the CMA-ES arm (GOAL §5.2) at d5 with the 12-seed standard instrument first — if it splits the same way, one gating mechanism ships two arms.
- Nightly regime blind spot confirmed again — second measured
case (after the JSO d5 add) where decisive evidence lived at d5,
invisible to the quick-2-D nightly battery. Strengthens the
case for a d5 (or
--extra-highdim) slice in the nightly loop alongside the multi-seed confirm item below.
-
drop_heuristic JSOonRewarding_Restartrejected (negative result) — ledger slot (2 nights, pooled CI95%[+0.0058,+0.0075]) measured flat on the 12-seed paired quick A/B: mean Δ+0.0026, sd 0.0249, CI95%[−0.0132,+0.0185], control flat; seed 42 alone+0.0200. Fourth consecutive training-seed artifact (0/4 screening hit rate). Would have reversed the standard-battery-validated 2026-08-02 JSO add (d5 +0.0287). See the 2026-08-10 log entry. -
add_heuristic NLSHADE_LBC→RoundRobin_Randomrecorded policy-moot — the slot only targets the pure-random reference / A/B control spec, which stays untouched by standing judgement call. - Measure
add_heuristic NLSHADE_LBConRewarding_Restart(§4.3 candidate with a positive prior: 3 confirmed control-spec accepts across 2 nights say the arm is strong under AOCC) — done 2026-08-11: rejected as an unconditional add (d2 loss outweighs the significant d5 gain); see the 2026-08-11 entries above. - Multi-seed confirm in the nightly loop — the 0/4 codify screening hit rate bounds the loop's value at the current plateau; rotate the nightly base seed and/or add a second-seed confirm gate before a screening accept lands in the ledger (extends the 2026-08-03 "price instance sensitivity into the codify gate" item).
- Fixed-seed repeatability measured (10 identical quick runs,
seed 42):
Rewarding_Restartbattery-mean AOCC sd 0.0206 (range 0.060),RoundRobin_Randomsd 0.0012 — the "per-seed" decision noise is almost entirely scheduling nondeterminism in the adaptive strategy, not instance sensitivity. See the 2026-08-09 log entry for the full source table. -
--sync-evalshipped (ioh_benchmark.py run→config.sync_evaluation→ synchronous future harvest in_run_threaded_evaluation): pooled repeat sd over seeds 42/1234/777 drops 0.0183 → 0.0115 (1.6× sd, 2.5× variance; seed-heterogeneous: 2.3×/1.25×/1.4×). Opt-in, default-off,comparewarns on mode mismatch, results carry async_evaltag. Use on both sides of future A/Bs; the N=12 instrument-level CI shrink is expected but not yet demonstrated (one null A/B per mode couldn't resolve it — keep accumulating nulls). - Event-drain wait measured ineffective (negative result) — waiting for eventbus queues to empty after publishing results did not reduce sd further (0.0113 vs 0.0094); queue-empty ≠ handlers idle. A real synchronous stepping mode needs handler-completion tracking, not queue polling.
- Residual ~0.009 sd: shared global RNG across handler threads —
heuristics draw from
np.randominside per-handler EventBus threads, so thread interleaving reorders the stream even at a fixed seed. Next lever: per-heuristicnp.random.Generatorseeded from (run seed, heuristic name); mechanical but touches ~14 heuristic modules. Measure with the same 10-repeat protocol. - Adopt
--sync-evalin the nightly loop / codify verification once a few sessions have used it interactively without surprises (it shifts absolute AOCC within noise; ledger continuity says switch deliberately, not silently).
- Single-fresh-night resurrection churn stopped — rejected codify
slots now stay hidden until the post-rejection evidence alone
reaches
--min-fresh-nights(default 2) distinct nights; one fresh seed-42 night no longer re-opens a slot a 12-seed A/B rejected (measured 0/3 hit rate across the 2026-08-03..07 resurrections).--min-fresh-nights 1restores legacy semantics; audit view shows per-slot progress toward the bar. See the 2026-08-08 log entry. - Rotate the nightly base seed by date — with every ledger night
keyed to seed 42,
n_nights >= 2measures persistence of one training-battery draw, not cross-instance generality. A dated seed rotation would make cross-night pooling cross-seed for free (check trend-table comparability + hold-out seed disjointness first). - Codify pre-gate: auto 12-seed paired A/B — before
--apply-topdeclares a slot actionable, optionally run theioh_benchmark.py run --decision-seedsinstrument and require a CI95 excluding zero (mechanises the manual protocol every session currently hand-runs; complements the 2026-08-03 "price instance sensitivity into the codify gate" item below).
-
NelderMead drop_heuristicmeasured flat and rejected — ledger evidence (2 nights, pooled CI95%[+0.0067,+0.0083]) did not survive the 12-seed paired quick A/B:Rewarding_Restartmean Δ−0.0003, sd 0.0188, CI95[−0.0122,+0.0117], control flat; seed 42 alone+0.0174(training-seed artifact, same signature as both 2026-08-03 Sensitivity rejections). Spec unchanged (applied + reverted in-branch); rejection recorded inplanning/self_improve_rejections_aocc.json. See the 2026-08-07 log entry. - Multi-seed pre-gate for codify-scan (raises priority of the
existing "price instance sensitivity into the codify gate" item)
— three consecutive ledger-positive slots rejected flat by the
12-seed instrument means single-seed
min_nights=2evidence has ~0 hit rate at the current plateau. Cheapest fix: highermin_nightsfor structural ops + a small (6-seed) screening A/B in the nightly post-loop step before a slot is surfaced actionable. - Promote codify verification to
--standardwhen the quick battery saturates (GOAL §4 step 4) —Rewarding_Restartper-seed quick sd (~0.019) is ~9× the control's; effects below ~0.012 cannot clear a 12-seed quick CI95.
-
CMAESadded to theadd_heuristiccandidate pool — the existing full CMA-ES heuristic (IPOP/BIPOP,heuristics/cma_es.py) was unreachable by the nightly loop; the bandit can now measure the only covariance-adapting family against the DE arms. Explicitsigma0=0.3makes the existingCMAES.sigma0kwarg rule fire. Direct add toRewarding_Restartmeasured flat on a 12-seed paired quick A/B (Δ +0.0005, CI95 [-0.0113,+0.0123]) → not shipped into the spec; the catalog route lets the ledger/hold-outs decide at the regimes (5-D rotated valleys) where the arm should matter. See the 2026-08-06 log entry. -
drop_analyzer Sensitivityre-hidden — resurfaced from a fresh 2026-08-06 single-seed night but is an apply-guard no-op (last analyzer in the bucket) on a spec unchanged since the 2026-08-03 12-seed rejection;codify-rejectre-recorded dated 2026-08-06. - Annotate guard-suppressed codify candidates —
codify-scansurfaces slots whose--apply-topwould be a safety-guard no-op (e.g. dropping the last analyzer) as "actionable"; detect and tag (or hide) them so sessions don't burn the codify slot on a no-op.
- Multi-seed A/B mechanised —
run --seeds/--decision-seeds(canonical 12-seed roster) /--reps Kand a seed-pairedcompare(per-strategy Δmean/sd/CI95 via t-dist, per-seed deltas, verdict markers, loud mixed-format error). The 2026-08-03 decision protocol is now one flag instead of a hand-rolled bash loop. 15 new tests; single-seed files/workflows unchanged. See the 2026-08-05 log entry. - Deterministic evaluation mode (follow-up to the nondeterminism
finding) —
_run_threaded_evaluationharvests whichever futures aredone()per loop pass, so the strategy's result view depends on OS scheduling. Partially shipped 2026-08-09 as--sync-eval(synchronous harvest, 2.3× repeat-sd cut); full determinism blocked on the shared-RNG / handler-thread items in the 2026-08-09 section above.
- Rejection memory shipped —
codify-scannow consults a per-metric rejections file (planning/self_improve_rejections_<metric>.json); rejected/moot slots are hidden from the report and skipped by--apply-topuntil fresh post-rejection evidence resurrects them (tagged for re-verification). Newcodify-rejectsubcommand records decisions; seeded with the five resolved aocc slots (LBFGSB/Center/Sobol.n 07-30, both Sensitivity slots 08-03).codify-scan --metric aoccnow truthfully reports 0 actionable candidates. See the 2026-08-04 log entry.
- Both remaining
Sensitivitycodify slots rejected (negative results) —update_interval 25 → 20(ledger pooled CI95%[+0.0092, +0.0117]) anddrop_analyzer Sensitivity(pooled CI95%[+0.0067, +0.0075]) both measured flat on a 12-seed paired quick A/B against the current spec: mean Δ−0.0012/−0.0007, CI95% straddling zero, controls flat. Training-battery artifacts; see the 2026-08-03 log entry. All 6 scan candidates now resolved (JSO → PR #289; LBFGSB/Center → rejected 07-30; Sobol.n → moot). - Standard battery measured nondeterministic run-to-run — identical tree+seed re-run shifts both arms by ≈ ±0.015 (threaded evaluation). Decision protocol updated in the log: ≥ 12 paired quick seeds with flat-control check, or ≥ 5 standard replicates per side.
- Codify-scan rejection memory — the scan re-surfaces
A/B-rejected and moot slots every night (no counterpart to the
already-codified suppression). Add a rejected-slot suppression
list (slot key + rejection date + evidence pointer) consulted by
codify-scanso the nightly report and--apply-topskip them. Shipped 2026-08-04 — see the section above. - Fix threaded-evaluation nondeterminism in the IOH harness — the standard battery cannot currently resolve +0.01-scale effects with a single run; find the ordering/seeding race and make batteries reproducible per seed (quick battery already is, modulo rare ±0.01 outliers).
- Price instance sensitivity into the codify gate — nightly
evidence keys on the seed-42 training battery; per-seed null-change
sd is ~0.015. Raise
min_nights(2 → 3+) and/or add a multi-seed confirm to the scan before surfacing a slot as actionable.
- Codify banked —
add_heuristic JSOslot from the aocc ledger (2 confirmed nights, pooled CI95%[+0.0092, +0.0133]; the 2026-08-01 accept was measured on the current post-Sobol-drop spec). Standard- battery paired A/B:Rewarding_Restartmean AOCC0.3374 → 0.3596(+0.0222; d2 +0.0156, d5 +0.0287), control flat. d5 now clearly beats the random floor (0.3196 vs 0.2878). Edit scoped toRewarding_Restartonly —RoundRobin_Randomstays the untouched reference.
- Bug fix (the aocc codify stall) —
default_codify_registries()anddefault_codify_apply_sources()inpanobbgo/self_improve.pygain ametric: str = "composite"parameter;"aocc"routes suppression topanobbgo.harness_ioh.make_ioh_strategiesand the--apply-topedit driver topanobbgo/harness_ioh.py. Between 2026-07-09 (nightly metric flip) and this fix, aocc evidence could never land as a source edit — the driver scannedharness.py, found no matching spec, and silently no-oped while the bandit re-discovered the same wins nightly (18 confirmeddrop_analyzer Restartaccepts across 17 nights).codify-scanthreads--metricinto both call sites. - First aocc codify banked — dropped the
Restartanalyzer from theRewarding_Restartspec inmake_ioh_strategiesvia the fixed driver. Local paired A/B (quick IOH battery):Rewarding_Restartmean AOCC0.3538 → 0.3922,RoundRobin_Randomcontrol flat. Post-codify scan auto-suppresses the candidate (self-stability verified). - Nightly visibility —
self_improve_nightly.ymlnow regenerates and commitsplanning/self_improve_codify_scan.txt(metric-aware) alongside the summary, so actionable evidence is readable without running anything. - Goal contract — new
planning/GOAL.md: metric of record, per-session operating loop, multi-day escalation ladder, SOTA-informed research backlog (MA-BBOB / LLaMEA / modular CMA-ES context). Pointer added toAGENTS.md. - Validation — 6 new tests (
TestMetricAwareCodifyRouting); fulltests/test_self_improve.pysuite green (611 passed); ruff clean. - Queued codify slots worked through (2026-07-30 second session) —
drop_heuristic SobolACCEPTED (standard battery +0.0176, spec now beats random at both dims);add_heuristic LBFGSBREJECTED (−0.0146 on the post-Sobol-drop spec — interaction negative);drop_heuristic CenterREJECTED (−0.0229; Center is load-bearing without Sobol);Sobol.n 32 → 38moot. Structural-add driver hardened in the same session: missing-comma fix,structural_kwargscarried into edits, factory-import rewriting, parse-validation net inapply_codify_edits(8 new tests). See the log's second 2026-07-30 entry. -
Open weakness — hold-out base seeds score far below training seed (0.04 vs 0.33 on 2026-07-30): instance-family generalization is the top research target— RETRACTED 2026-08-11. This was a metric-unit bug, not a weakness:_measure_holdoutnever routed through the AOCC path, so AOCC runs wrotecomposite_score(~0.045 scale) into hold-out records next to mean-AOCC training records (~0.34). Fixed in #299; the real gap is 0.3383 vs 0.3342. Seeplanning/LOOP_DIAGNOSIS_2026-08-11.md§3.1. The research slot it occupied inGOAL.md§5.1 is now regime-conditional strategy selection.
Entries before this point were moved to planning/done/TODO_archive_pre-2026-07-30.md on 2026-08-11 to keep this file readable. Nothing was deleted — the archive is the same newest-first format.