Skip to content

perf: back off the fnnls warm-start memo on scattered evaluation streams #613

Description

@Jammy2211

Overview

The fnnls warm-start memo (on by default since #498) seeds each positive-only solve from the previous evaluation's final passive set. On a scattered evaluation stream (iid draws, e.g. Nautilus early live points) the seed is bad, the post-solve fallback guard drops it, the next solve restarts dense and re-seeds -- so every other solve pays for a bad seed. autolens_profiling#332 measured the alma interferometer Delaunay solve 2.17x slower memo-on vs memo-off on an iid stream (rect 1.08x), while a local walk gains 0.68x / 0.13x. This adds a per-key back-off so the memo stops re-seeding a key that keeps falling back, and re-probes it on an exponential schedule so a stream that turns local recovers the gain.

Source task: organs/PyAutoPulse/tasks/interferometer_nnls_memo_scattered_stream_guard.md (related autolens_profiling#332).

Plan

  • Reproduce locally with a small CPU witness (scattered vs local-walk stream, solver-only timing) before changing anything.
  • Track, per memo key, how many seeded solves in a row fell back.
  • After 2 consecutive fallbacks, skip the memo seed (solve dense, still refreshing the entry) for an exponentially growing number of solves (1, 2, 4, ... capped at 32), then probe the seed again; an accepted seed resets the streak.
  • A local-walk stream never reaches two consecutive fallbacks, so its behaviour (and every reconstruction) is unchanged; no default tolerance or config key changes.
  • Unit tests for both streams over several seeds; re-time the witness.

Tier: glance — merge mode: human /prm (scope of this session ends at a local commit; Heart RED, shipping is the human's call)

Detailed implementation plan

Affected Repositories

  • PyAutoArray (primary)

Branch Survey

Repository Current Branch Dirty?
./PyAutoArray main clean

Suggested branch: feature/nnls-memo-scattered-backoff
Worktree root: ~/Code/PyAutoLabs-wt/nnls-memo-scattered-backoff/

Implementation Steps

  1. autoarray/inversion/inversion/nnls_memo.py: add a bounded per-key back-off table (fallback streak, skip remaining), with backoff_should_skip(key), backoff_record_fallback(key), backoff_record_accept(key); constants _NNLS_BACKOFF_AFTER_FALLBACKS = 2, _NNLS_BACKOFF_MAX_SKIP = 32.
  2. autoarray/inversion/inversion/inversion_util.py::reconstruction_positive_only_from: consult the back-off before reading the memo entry; record fallback / accept after the existing ratio guard; dense solves still refresh the entry.
  3. test_autoarray/inversion/inversion/test_nnls_memo.py: state-machine tests; iid stream (several seeds) — seeded solves bounded, reconstructions equal memo-off to round-off; local walk (several seeds) — back-off never engages and reconstructions are bit-identical with the back-off disabled.

Key Files

  • autoarray/inversion/inversion/nnls_memo.py
  • autoarray/inversion/inversion/inversion_util.py
  • test_autoarray/inversion/inversion/test_nnls_memo.py

Original Prompt

Click to expand

See the Pulse task file above; Witness: on a recorded real-sampler evaluation sequence the memo-on fnnls solve is no slower than memo-off on every phase, keeping >= 80 % of the local-walk gain; figure of merit unchanged.

Activity

  1. Jammy2211 commented on Oct 4, 2026

    @Jammy2211
    CollaboratorAuthor

    Status: implemented locally, ship pending human (Heart RED)

    Local commit 0d9bbecd on feature/nnls-memo-scattered-backoff (worktree ~/Code/PyAutoLabs-wt/nnls-memo-scattered-backoff). Not pushed, no PR.

    Guard. Per-key back-off in nnls_memo: after 2 consecutive seeded-solve fallbacks, the next 1, 2, 4, ... (cap 32) would-be-seeded solves start dense and still refresh the entry. Each backed-off dense solve also checks for free whether the seed it skipped would have passed the guard, using the symmetric difference with the final passive set, which equals warm_start_errors. If it would have passed, the remaining skips are cancelled. An accepted real seed resets the streak. A local walk never falls back twice in a row, so it is untouched. There is a new stats key warm_start_backoff, and a new helper memo_clear() that clears entries and back-off state together. No config or tolerance changes.

    Witness (local, synthetic SIS-like rect n=576, solver-only, min of 3 replays, 3 seeds, 64 solves, 8 threads CPU):

    stream main on/off branch on/off branch ms/solve off -> on
    local walk (step 0.002) 0.18x 0.18x 16.9 -> 2.9
    iid (central 20 %) 1.49x 1.17x 17.2 -> 20.2
    iid phase then walk phase (walk part) 0.20x 0.20x 16.6 -> 3.4

    Seeded solves on iid fall from 32/64 to 7-13/64. The remaining iid overhead is the exponential ramp (probes at solves 1, 3, 6, 10, 16, 26, 44, ...). Over long streams it tends to about 1 bad probe per 33 solves. Reconstructions agree with memo-off to <= 3.4e-15 in every arm.

    Tests: test_nnls_memo.py has 4 scattered + 4 walk + 3 scattered-then-walk seeds, plus state-machine tests. On the walk stream the reconstructions are bit-identical with the back-off disabled. The full test_autoarray suite gives 1975 passed, 4 xfailed.

    Not done / follow-ups: the Pulse-task witness (Nautilus replay on alma Delaunay-1500 / rect 39², plus the HST imaging Delaunay control) needs the autolens_profiling harness and real data. autolens_profiling harnesses that reset the memo with nnls_memo._nnls_passive_set_memo.clear() (fixed_light_numba.py, delaunay_numba_nnls_iterations.py, fixed_light_numpy_solvers.py, fixed_light_s4b_checks.py) should switch to nnls_memo.memo_clear(), or back-off state can carry from one arm into the next.

  2. Jammy2211 commented on Oct 7, 2026

    @Jammy2211
    CollaboratorAuthor

    Library PR Created

    PR: #615 (pending-release, Closes #613). Branch rebased onto main; full PyAutoArray suite 1975 passed, 4 xfailed.

    Heart readiness was YELLOW, acknowledged by the human for this ship:

    • autogalaxy_workspace: open PR 7d old
    • autolens_workspace: open PR 7d old
    • euclid_strong_lens_modeling_pipeline: open PR 7d old
    • release validation stale: source moved since rehearsal (PyAutoNerves)

    Workspace impact: none (internal solver behaviour; new warm_start_backoff stats key, new nnls_memo.memo_clear()).

    Follow-ups (not merge gates)

    • Real Nautilus-replay witness still open. The shipped evidence is the local solver-only witness (n=576, iid on/off 1.49x -> 1.17x, walk 0.18x unchanged, iid->walk recovery 0.69x). A replay of a real Nautilus evaluation stream (e.g. the autolens_profiling#332 ALMA Delaunay case) is still to be run; by human decision it does not gate the merge.
    • autolens_profiling harnesses should call nnls_memo.memo_clear() instead of clearing nnls_memo._nnls_passive_set_memo directly (fixed_light_numba.py, delaunay_numba_nnls_iterations.py, fixed_light_numpy_solvers.py, fixed_light_s4b_checks.py), otherwise back-off state leaks across rows. _production_config.py also cites nnls_memo.py:63, which this PR shifts (prose only).

    Next: CI on #615, then /prm (human).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions