Skip to content

feat(interferometer): fits, save_attributes and aggregator reload on array-free datasets (streaming phase 2) - #639

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/streaming-p2-fit-save-reload
Sep 30, 2026
Merged

Jammy2211 merged 1 commit into
mainfrom
feature/streaming-p2-fit-save-reload

Conversation

@Jammy2211

@Jammy2211 Jammy2211 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Streaming phase 2 (Mind epic streaming-visibilities; https://github.com/orgs/PyAutoLabs/discussions/13). Closes #638. Requires PyAutoLabs/PyAutoArray#593 (phase 1, on main, pending release). The PyAutoLens half is PyAutoLabs/PyAutoLens#758.

Phase 1 gave aa.Interferometer.from_stream / from_sparse_terms: a dataset with no data, noise_map, uv_wavelengths or transformer. An AnalysisInterferometer search still could not run on it — save_attributes wrote the visibility arrays and the transformer class, the aggregator read dataset.fits HDUs positionally and never re-attached a sparse operator, and FitInterferometer.profile_visibilities built Visibilities.zeros(N_vis) through the transformer. This PR:

  • Fit: profile_visibilities returns None on a transformer-less dataset with no ordinary light profile and raises a typed DatasetException when there is one (non-linear light profiles arrive in phase 4; linear-light and pixelization fits work); profile_subtracted_visibilities evaluates it first (so a light-profile fit raises rather than silently fitting unsubtracted terms) then returns None without data; inversion_with_data returns the inversion itself when there is no data. In-memory fits are unchanged.
  • Save: new module-level interferometer_hdu_list_from(dataset) (used by save_attributes here and in PyAutoLens): in-memory datasets write the same four arrays as before (EXTNAMEs MASK, DATA, NOISE_MAP, UV_WAVELENGTHS); array-free datasets write the mask, NUFFT_PRECISION_OPERATOR (W̃), DIRTY_IMAGE and DIRTY_BEAM (float64) plus a lossless SPARSE_TERMS_SCALARS float64 HDU (sum_weights, data_term, noise_normalization, n_vis, eps or NaN — order in SPARSE_TERMS_SCALARS_ORDER); the primary header carries readable copies (SUMW, DATATERM, NOISENRM, NVIS, EPS/TRNSFRMR when recorded). FITS header cards hold at most 20 characters, which truncates exponent-form float64, so the scalars are read from the HDU and only fall back to the cards for older files. transformer_class.json is written only when a transformer exists.
  • Reload: _interferometer_from looks HDUs up by EXTNAME (case-insensitive) with a positional fallback for pre-existing outputs; when the terms HDU is present it rebuilds aa.SparseTerms and returns aa.Interferometer.from_sparse_terms(terms, real_space_mask) — array-free, operator re-attached. PyAutoLens imports this loader, so it needs no mirror.

Witness (1e5 NUFFT visibilities, 616-pixel mask on a 40×40 grid streamed in 5 chunks, 20×20 rectangular pixelization, af.m.MockSearch with save_for_aggregator=True, visualization disabled — dataset/fit plots need arrays until phase 3):

lib dataset.fits log_evidence in-memory sparse array-free reloaded rel reloaded/array-free rel array-free/in-memory rel jit/numpy
ag 80,640 B −316529.99966181 same same 0.0 1.8e-16 2.6e-14
al 80,640 B −315436.25588951 same same 0.0 1.9e-16 2.6e-14

The file is set by the real-space grid (mask 40×40, W̃ 56×56), not the number of visibilities: 48 B/vis × 1e5 would be 4.8 MB. fits.info lists MASK, NUFFT_PRECISION_OPERATOR, DIRTY_IMAGE, DIRTY_BEAM and SPARSE_TERMS_SCALARS, all float64, and no transformer_class.json.

Also fixed here (review finding, pre-existing): agg_util.mask_header_from read PIXSCAY for both axes, so any saved mask with non-square pixel scales reloaded square; it now reads (PIXSCAY, PIXSCAX) with the old behaviour only for files lacking PIXSCAX. A second pre-existing finding — database-path searches never register dataset.fits through save_fits, so Fit.value("dataset") is None there — is out of scope and filed in Mind.

API Changes

Additive plus small behaviour changes on array-free datasets only. New public name autogalaxy.interferometer.model.analysis.interferometer_hdu_list_from(dataset) and SPARSE_TERMS_HEADER_KEYS. FitInterferometer.profile_visibilities / profile_subtracted_visibilities may now return None (array-free dataset, no ordinary light profile) or raise aa.exc.DatasetException (array-free dataset with one). save_attributes writes EXTNAME-tagged HDUs (same arrays as before for in-memory datasets) and skips transformer_class.json without a transformer. The aggregator loader reads by EXTNAME with positional fallback and re-attaches the sparse operator for array-free outputs; legacy dataset.fits files load as before.
See full details below.

Test Plan

  • pytest test_autogalaxy — 1287 passed (1284 before the review fixes)
  • New: array-free pixelization-only fit == in-memory sparse fit at rel 1e-8 (numpy + jax.jit); linear-light-only fit works; ordinary light profile on an array-free dataset raises; interferometer_hdu_list_from EXTNAMEs/header for both dataset kinds (in-memory arrays identical to before); save_attributes → _interferometer_from round trips for array-free and in-memory datasets through af.DirectoryPaths in tmp_path
  • Red check: with the old loader restored, the array-free round trip fails (noise_map must have the same shape as data; got (6, 6) and (7, 7) — the positional read took W̃ as data); in-memory round trip still passes
  • Witness (table above)
  • CI green on unittest 3.12 / 3.13 / nojax + docs

Heart RED override (development only)

Heart verdict at ship: RED, exact reason release validation FAILED (stage integrate) (verdict 2026-09-30T12:42Z, unrelated release-integrate leg). The live human authorized the development-only override for issue #638 in-session on 2026-09-30 ("Authorize override for #638"): commit, push and the pending-release PRs only. Branch gates before the ask: test_autogalaxy 1287 passed, test_autolens 770 passed + 1 xfailed, round-trip red-check, independent Codex (gpt-6-astra) review: FINDINGS (3) — #1 FITS header cards truncate float64 scalars (reproduced: 1.2345678901234567e+20 → 1.23456789012345e+20): fixed in-branch with a lossless SPARSE_TERMS_SCALARS float64 HDU (header cards kept as readable copies, loader falls back to them for older files), red-checked; #3 agg_util.mask_header_from read PIXSCAY for both axes (pre-existing): fixed in-branch with a non-square round-trip test, red-checked; #2 database-path searches never register dataset.fits via save_fits (pre-existing, in-memory path too): filed as draft/bug/autogalaxy/database_paths_dataset_fits_not_registered.md. This PR does not claim to repair Heart; merge needs its own explicit human command with every required check green.

Full API Changes (for automation & release notes)

Added

  • autogalaxy.interferometer.model.analysis.interferometer_hdu_list_from(dataset) -> HDUList — the dataset.fits contents for in-memory (four arrays) or array-free (SparseTerms extensions) datasets
  • autogalaxy.interferometer.model.analysis.SPARSE_TERMS_HEADER_KEYS — header key map (SUMW, DATATERM, NOISENRM, NVIS, EPS, TRNSFRMR); SPARSE_TERMS_SCALARS_ORDER — field order of the lossless SPARSE_TERMS_SCALARS HDU

Changed Behaviour

  • FitInterferometer.profile_visibilities → None on an array-free dataset with no ordinary light profile; raises aa.exc.DatasetException with one. profile_subtracted_visibilities → None when the fit has no data. inversion_with_data → the inversion itself when the fit has no data.

  • AnalysisInterferometer.save_attributes — writes via interferometer_hdu_list_from (EXTNAME-tagged; array-free datasets write SparseTerms extensions + header scalars); transformer_class.json only when a transformer exists

  • autogalaxy.aggregator.interferometer.interferometer._interferometer_from — EXTNAME lookup with positional fallback; rebuilds array-free datasets via aa.Interferometer.from_sparse_terms with the operator re-attached

  • autogalaxy.aggregator.agg_util.mask_header_from — reads PIXSCAX for the x pixel scale (was PIXSCAY for both axes); falls back to PIXSCAY only when PIXSCAX is absent

Migration

  • None for in-memory datasets. Code that fits array-free datasets should not use ordinary (non-linear) light profiles until phase 4.

Generated by the PyAutoLabs agent workflow.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JZZksyZ8LTA4LLxoZjQMNF

…array-free datasets (streaming phase 2, #638)

profile_visibilities / profile_subtracted_visibilities / inversion_with_data
guard the transformer-less dataset (pixelization and linear-light fits run;
ordinary light profiles raise until phase 4). save_attributes writes
dataset.fits via the new interferometer_hdu_list_from: array-free datasets
persist their SparseTerms as FITS extensions (W~, dirty image, dirty beam,
scalars + provenance in the header), in-memory datasets the same four arrays
as before, now EXTNAME-tagged. The aggregator loader reads by EXTNAME with a
positional fallback and rebuilds array-free datasets via from_sparse_terms
with the sparse operator re-attached, so a reloaded fit reproduces
log_evidence exactly. Scalars ride a lossless float64 HDU (FITS header
cards truncate exponent-form float64); agg_util.mask_header_from now reads
PIXSCAX for the x axis (pre-existing bug found in review).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JZZksyZ8LTA4LLxoZjQMNF
@Jammy2211

Copy link
Copy Markdown
Collaborator Author

Independent review (Codex gpt-6-astra) of this branch pair before PR-open

  1. P2 — FITS serialization loses float64 precision (in-diff). analysis.py:96 writes the scalar terms as ordinary numeric header cards. A byte-stream Astropy round trip converts 1.2345678901234567e+20 to 1.23456789012345e+20. Thus SUMW, DATATERM, and NOISENRM do not reliably survive exactly; changed data/noise terms change the reloaded evidence. The header test compares unserialized values, where this truncation has not happened. Store lossless representations or binary float64 scalars.

  2. P2 — Database searches cannot reload the saved dataset (pre-existing, still affects this feature). Galaxy analysis.py:350 and Lens analysis.py:368 write directly to disk without registering the HDU list through DatabasePaths.save_fits. For a search using database paths, database Fit.value("dataset") returns None; mask loading then fails on [0]. Unlike SearchOutput, database Fit.value does not scan the image directory. Both new round-trip tests bypass this failure with _FitStub and DirectoryPaths.

  3. P2 — Non-square pixel scales silently corrupt reloaded geometry (pre-existing helper defect, inherited by the new loader). agg_util.py:104 reads PIXSCAY for both axes. A saved mask with scales (0.1, 0.2) reloads as (0.1, 0.1). The new interferometer.py:89 also copies these incorrect scales into SparseTerms, so from_sparse_terms’ geometry check passes while the stored operator describes the original geometry. Subsequent fits can produce different reconstructions and evidence.

FINDINGS (3)

Disposition: FINDINGS (3) — #1 FITS header cards truncate float64 scalars (reproduced: 1.2345678901234567e+20 → 1.23456789012345e+20): fixed in-branch with a lossless SPARSE_TERMS_SCALARS float64 HDU (header cards kept as readable copies, loader falls back to them for older files), red-checked; #3 agg_util.mask_header_from read PIXSCAY for both axes (pre-existing): fixed in-branch with a non-square round-trip test, red-checked; #2 database-path searches never register dataset.fits via save_fits (pre-existing, in-memory path too): filed as draft/bug/autogalaxy/database_paths_dataset_fits_not_registered.md.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

pending-release PR queued for the next release build

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(interferometer): pixelization-only fits, save_attributes and aggregator reload on array-free datasets (streaming phase 2)

1 participant