Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
f81007c
ENH: deterministic product-table schemas + DDL single-home ratchet
RBhupi Aug 21, 2026
6a63cf1
FIX: strict tracked-cell columns in cell_tracks upsert
RBhupi Aug 21, 2026
0d090ae
FIX: cells_by_scan upsert uses union of row columns
RBhupi Aug 21, 2026
de05e56
FIX: TSE priority score raises on NaN components
RBhupi Aug 21, 2026
97d8be1
FIX: yaml_writer quotes keys, keeps empty dicts, rejects list-of-dicts
RBhupi Aug 21, 2026
089a61f
REFACTOR: delete dead _FIXED_CBS_COLS and unused execution-history DDL
RBhupi Aug 21, 2026
9eceaa5
ENH: log dashboard variable fallback instead of silent substitution
RBhupi Aug 21, 2026
0e55aae
ENH: reader.field_map and reader.fields config schema
RBhupi Aug 21, 2026
4240fe0
FIX: mypy type-var in renderer fallback log
RBhupi Aug 21, 2026
0050261
ENH: thread field_map/fields into IngestConfig
RBhupi Aug 21, 2026
f24f2cd
ENH: canonicalize source variable names at ingest via reader.field_map
RBhupi Aug 21, 2026
d093fbc
DOCS: ingest canonicalization via reader.field_map; track ARCHITECTUR…
RBhupi Aug 21, 2026
e8b27b4
ENH: stat_column helper + CELL_LABELS_VAR canonical constant
RBhupi Aug 21, 2026
4c0f0f9
REFACTOR: var_names -> global_.tracking_field role knob; cell_labels …
RBhupi Aug 21, 2026
00c73ee
ENH: resolve-time invariant for tracking_field availability
RBhupi Aug 21, 2026
567b3bf
FIX: REFLECTIVITY_VAR alias maps to reader.field_map; user tracking_f…
RBhupi Aug 21, 2026
0e237f3
ENH: structural grid contract; ZDR optional; detection boundary check
RBhupi Aug 21, 2026
11f22b0
ENH: field-generic tracker via stat_column; cell_uid v2 = hash(scan_i…
RBhupi Aug 21, 2026
4019fe4
TEST: quarantine ratchet for reflectivity stat-column literals
RBhupi Aug 21, 2026
eb41186
ENH: StoreClient.run_config provenance accessor
RBhupi Aug 21, 2026
8fe714c
REFACTOR: TSE role names (field_max/weights.field); provenance-driven…
RBhupi Aug 21, 2026
dec6b68
ENH: renderers and hover resolve tracking field from run provenance
RBhupi Aug 21, 2026
2c8e7f8
DOCS: variable names + tracking field guide; API provenance notes
RBhupi Aug 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,9 @@ results/
# Ignore user config overrides
scripts/user_config.py

# local planning folder — never committed
plan/

# ignore scratch notes but keep docs and repo-standard markdown/text
*.md
!README.md
Expand Down
136 changes: 136 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# Adapt Architecture Contract

This is the authoritative statement of Adapt's architecture. It is short on
purpose: every rule here is machine-enforced, and the enforcement point is
named next to the rule. If you change the architecture, change the enforcement
and this document in the same commit — `tests/test_architecture.py` fails if
this file, `.importlinter`, and the source tree disagree.

## Layer stack

A package may import only packages in layers **below** its own. Packages on
the same line (joined by `|`) are independent siblings and never import each
other. Enforced by `lint-imports` (`.importlinter`, contract "Adapt layer
stack", `exhaustive = true` — a new top-level package fails CI until it is
deliberately placed here **and** there).

```layers
cli
consumers | visualization
runtime
configuration
api | execution
modules | persistence
contracts | downloaders | utils
```

| Package | Role |
|---|---|
| `cli` | Outermost shell (`adapt <object> <action>`); may import anything, nothing imports it. |
| `consumers/` | Dashboard, target selection. Read **only** via `adapt.api.StoreClient`. |
| `visualization/` | Plotting. Core never imports it; matplotlib is confined here + consumers. |
| `runtime/` | Composition root: orchestrator, processor, postprocessor. The only place everything is wired together. |
| `configuration/` | Pydantic schemas, three-tier resolution, frozen `InternalConfig`, `defaults.yaml` module list. |
| `api/` | `StoreClient` — the read-only public facade over the store. |
| `execution/` | Graph builder/executor + `nodes/`: the one sanctioned wiring layer over modules. |
| `modules/` | Scientific modules (ingest, detection, projection, analysis, tracking, …). Deterministic, no store I/O. |
| `persistence/` | SQLite catalogs + NetCDF objects. No science, no orchestration. |
| `contracts/` | Frozen dataclasses, Protocols, and `check_*` validators. Zero adapt imports. |
| `downloaders/` | All boto3/S3 calls. Zero adapt imports. |
| `utils/` | Shared pure functions. Zero adapt imports. |

Optional plug-in modules are loaded via the `extensions:` list of dotted import
paths in the run configuration — there is no in-tree extensions package.

## The eight rules

1. **Scientific modules** (`modules/*`) import only `contracts`, `utils`
(plus `downloaders` for acquisition). They never touch the store, config,
runtime, or each other. Config arrives by injection as a frozen
`<name>_config` context entry.
*(lint-imports layers; `tests/test_architecture.py` module-independence tests)*
2. **Modules are deterministic**: no wall clock, no global RNG — identical
input + config ⇒ identical output.
*(AST bans + run-twice determinism tests)*
3. **Inter-module data flows through the context dict** via declared
`inputs`/`outputs`. Every output gets a `check_*` validator in `contracts/`;
the uncontracted allowlist only shrinks.
*(context-key coherence tests + `_UNCONTRACTED_OUTPUTS` ratchet)*
4. **Adding a module** = module package + `execution/nodes/` wrapper + one
line in `configuration/defaults.yaml` + contract validators. Zero edits
anywhere else. If you are editing `runtime/` to add a module, stop — the
design is wrong.
*(registration tests; new module auto-discovered by all fitness tests)*
5. **Consumers read only through `adapt.api.StoreClient`.** No
sqlite/duckdb, no NetCDF opens, no globs, no store paths.
*(lint-imports forbidden contract + consumer AST fitness test)*
6. **Every heavy third-party dependency has exactly one home** —
`_DEP_HOMES` in `tests/test_architecture.py` is the authoritative list.
A new dependency = `pyproject.toml` entry + `_DEP_HOMES` entry in the
same change.
*(dependency-home test + declared-dependencies test)*
7. **Public API = `adapt.api.__all__`** plus the documented consumer entry
points (`adapt.consumers.live.main`, target-selection exports). Everything
else is internal and may change without notice.
8. **Fail loudly.** No fallbacks, no bare `except`, no silent defaults.

## Product tables: core + extras

Science outputs vary by data source (NEXRAD and ARM radars ship different
variables), so product-table schemas are frozen at the first write of each
run — per run, per source, the column set is stable. Two rules keep that
flexible without being fragile:

- **Core identity columns** (`run_id`, `scan_id`, `scan_time`,
`scan_time_unix`, `valid_time`, `cell_uid`, `cell_label`) have contracted
types, fixed by name in `persistence/products.py` (`_CORE_COLUMN_TYPES`).
The writer stamps scan identity; modules never pass it. Readers may
hardcode these names.
- **Everything else is an extra**: source-dependent, frozen by a
deterministic rule (numeric → REAL, bool → INTEGER, else TEXT) so the
schema is a function of the column *names* — never of the values in
whichever scan happened to arrive first. Readers discover extras through
the `table_schemas` snapshots, never by hardcoding.

Static bookkeeping tables (scans, runs, artifacts, history, tracking core)
are defined in `configuration/schemas/*.sql` — the single DDL home. A
shrink-only ratchet in `tests/test_architecture.py` pins the remaining
legacy exception (`track_store.py`) and rejects any new `CREATE TABLE`
outside it.

Variable names are canonicalized once at ingest via `reader.field_map` /
`reader.fields`; downstream code never maps names. The applied mapping is
recorded in each grid's `source_fields_json` attribute as provenance.

## Decisions not to relitigate

See `docs/design/` for the full decision records.

- `scan_id` = sha256[:16] of the raw file bytes is the join key everywhere;
`scan_time` is single-format ordering/display metadata
(`adapt.utils.time.to_scan_iso` is the only serializer).
- `StoreClient` is read-only with typed filters — consumers never write.
- Module registration lives in `defaults.yaml`, not entry points.
- Contracts are validators and frozen dataclasses, not ABCs.
- The store is SQLite catalogs + NetCDF objects; conversion happens once at
the boundary.
- Two name bindings, each in one place: source→canonical names via
`reader.field_map` (applied inside ingest); role→field via
`global_.tracking_field` (which canonical field the core runs on).
Stat columns are minted only by `adapt.contracts.stat_column`.
- `cell_uid` v2 = hash(scan_id, cell_label): deterministic, collision-free,
field-free. Stored runs keep their minted uids; nothing recomputes them.

## Evidence required with any change

All green in CI before a change is accepted:

```
ruff check src tests
lint-imports
pytest # integration tests are opt-in: pytest -m integration
mypy src/adapt
```

New behavior ships with a test. A new invariant ships with a fitness function.
An allowlist never grows silently — extend the mechanism instead.
7 changes: 4 additions & 3 deletions docs/design/modules/ingest.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,11 +39,12 @@ A `Path` object pointing to a downloaded NEXRAD Level-II file (`.gz` or uncompre

The resolved `InternalConfig`. Ingest reads:
- `config.regridder.*` — grid shape, spatial extent, interpolation parameters
- `config.global_.var_names` — canonical field name mappings
- `config.reader.field_map` — source-to-canonical variable renames applied once here
- `config.reader.fields` — explicit canonical keep-list (empty = keep all)

### Output: `grid_ds`

A 3-dimensional `xarray.Dataset` on a regular Cartesian grid `(z, y, x)`. All coordinates are in metres from the radar origin. Reflectivity and velocity fields are renamed to canonical names defined in `config.global_.var_names`.
A 3-dimensional `xarray.Dataset` on a regular Cartesian grid `(z, y, x)`. All coordinates are in metres from the radar origin. Source variable names are renamed to canonical names via `config.reader.field_map`, and only `config.reader.fields` are kept (empty list = keep all) — inside `RadarDataLoader`, the single canonicalization point. The applied mapping is recorded in `attrs["source_fields_json"]`.

```
grid_ds
Expand Down Expand Up @@ -83,7 +84,7 @@ Polar Radar object
▼ Cressman regridder (via adapters/)
3D Cartesian grid
│
▼ Field renaming (via config.global_.var_names)
▼ Field renaming + selection (via config.reader.field_map / fields)
Canonical grid_ds
│
▼ z-slice at config.regridder.analysis_level
Expand Down
1 change: 1 addition & 0 deletions docs/readme.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@

installation
USAGE
variable_names
dashboard_reference
cli_reference
vision
Expand Down
124 changes: 124 additions & 0 deletions docs/variable_names.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# Variable Names and the Tracking Field

Adapt works with any radar source and, in principle, any 2-D scalar field.
Two independent bindings make that possible, and each lives in exactly one
place:

| Binding | Config knob | Applied where |
|---|---|---|
| Source name → canonical name | `reader.field_map`, `reader.fields` | Once, inside the ingest module |
| Role → canonical field | `global.tracking_field` | Read by detection / projection / tracking |

Everything downstream of ingest — module code, contracts, stored column
names, the dashboard, the Target Selection Engine — speaks **canonical
names only** and never maps names again.

## Canonical names

The canonical vocabulary is simply Adapt's established names:
`reflectivity`, `velocity`, `differential_reflectivity`,
`spectrum_width`, `cross_correlation_ratio`, `differential_phase`, …
The segmentation label grid is always `cell_labels` (pipeline-internal,
not configurable). Per-cell statistics columns are always minted as
`radar_<field>_<stat>` (e.g. `radar_reflectivity_max`) — the `radar_`
prefix is a namespace, not a physics claim.

## Renaming source variables (`reader.field_map`)

If your input files name a field differently — ARM CMAC calls reflectivity
`corrected_reflectivity` — declare the rename once:

```yaml
reader:
field_map:
corrected_reflectivity: reflectivity
corrected_differential_reflectivity: differential_reflectivity
fields: [reflectivity, differential_reflectivity]
```

- `field_map` maps **source name → canonical name** and is applied inside
ingest, immediately after regridding. Map entries whose source variable
is absent from a given file are ignored.
- `fields` is the explicit keep-list (canonical names, checked after the
renames). An empty list keeps every field the file provides. If a
listed field is missing, the run fails loudly naming what was available.
- The legacy `REFLECTIVITY_VAR: dbz` user setting still works and means
exactly this: `field_map: {dbz: reflectivity}`.

The applied mapping is recorded as provenance: in each grid's
`source_fields_json` attribute and in the run's stored configuration
(`registry.db → runs.config_json`), readable via
`StoreClient.run_config(run_id)`.

## Choosing the tracked field (`global.tracking_field`)

Detection thresholds on, projection flows on, and tracking links cells on
one canonical field — `reflectivity` by default:

```yaml
global:
tracking_field: reflectivity # or any canonical field, e.g. pressure
analyzer:
radar_variables: [reflectivity, velocity] # must include tracking_field
```

Two rules are enforced at configuration time (the run refuses to start
otherwise):

1. `tracking_field` must be in `analyzer.radar_variables` — otherwise the
per-cell statistics the tracker reads are never computed.
2. If `reader.fields` is non-empty, it must include `tracking_field`.

Dual-polarization fields are **optional**: a source without
`differential_reflectivity` runs the full pipeline; ZDR-derived columns
simply do not exist for that run (consumers discover available columns
through the store's `table_schemas`).

## Cell identity (uid v2)

Cell uids are `hash(scan_id, cell_label)` — a pure function of the scan's
raw bytes and the cell's label. Identical input and configuration always
mint identical uids, for any source and any tracked field. Runs recorded
before v2 keep their stored uids (uids are minted once and never
recomputed; every table scopes them by `run_id`). Cross-run comparisons
join on `scan_id` + `cell_label`.

## How consumers resolve the field

Consumers never guess. The run's provenance is the truth:

```python
client.run_config(run_id) # full resolved config (dict)
client.run_tracking_field(run_id) # the canonical tracked field
```

`run_tracking_field` also understands pre-v2 provenance (runs that
recorded `var_names.reflectivity`), so old stores stay readable. The
dashboard backdrop, hover statistics, and the Target Selection Engine's
`CellSnapshot.field_max` all resolve through it.

## Compatibility notes

- **Existing stores** remain readable unchanged; canonical names equal the
names Adapt always used.
- **Changed field sets need a new collection**: a collection's product
tables freeze their column set on first write, so a run whose
`fields` / `field_map` / analyzer whitelist produces different columns
is rejected by an existing collection (by design — create a new one).
- **Saved TSE YAML**: the priority weight key `reflectivity:` is now
`field:` (it weights the tracked field's max). Old files fail loudly at
load with a message naming the unknown key.
- **`cell_tracks.max_reflectivity`** (and `Track.max_reflectivity_dbz`,
`FilterSpec.max_refl_*`): the value is the maximum of the run's
*tracked field*; the name is kept for store compatibility and only
means dBZ for reflectivity-tracked runs.

## Onboarding a new source, in practice

1. Write (or reuse) an ingest path that reads the format.
2. Add the source's `reader.field_map` and an explicit `reader.fields`.
3. Pick `global.tracking_field` and match `analyzer.radar_variables`.
4. Run into a **new collection**.

No detection, projection, tracking, persistence, api, or consumer code
changes are required.
3 changes: 3 additions & 0 deletions src/adapt/api/domain.py
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,9 @@ class Track:
origin_type: str # INITIATION | SPLIT | MERGE | UNKNOWN
termination_type: str # TERMINATION | MERGED | ACTIVE_AT_END | UNKNOWN
max_area_km2: float
# Max of the run's TRACKED field over the cell's life (see
# StoreClient.run_tracking_field) — dBZ only when that field is
# reflectivity. Column name kept for store compatibility.
max_reflectivity_dbz: float


Expand Down
2 changes: 2 additions & 0 deletions src/adapt/api/selection.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,8 @@ class FilterSpec:
n_scans_min: int | None = None
max_area_min_km2: float | None = None
max_area_max_km2: float | None = None
# Filters the cell_tracks.max_reflectivity column: max of the run's
# TRACKED field (dBZ only for reflectivity-tracked runs).
max_refl_min_dbz: float | None = None
max_refl_max_dbz: float | None = None
origin_types: frozenset[str] | None = None
Expand Down
43 changes: 43 additions & 0 deletions src/adapt/api/store_client.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@

from __future__ import annotations

import json
import sqlite3
from collections.abc import Mapping
from dataclasses import dataclass
Expand Down Expand Up @@ -176,6 +177,48 @@ def run(self, run_id: str) -> Run:
raise StoreError(f"Run '{run_id}' does not exist")
return self._run_from_row(row)

def run_config(self, run_id: str) -> dict:
"""The resolved configuration this run was executed with.

Parsed from the run provenance record (registry runs.config_json) —
the single authoritative record of per-run settings such as
``global_.tracking_field`` and ``reader.field_map``. Keys use the
InternalConfig field spelling (``global_``, not ``global``).
"""
row = (
self._registry_conn()
.execute("SELECT config_json FROM runs WHERE run_id = ?", (run_id,))
.fetchone()
)
if row is None:
raise StoreError(f"Run '{run_id}' does not exist")
raw = row["config_json"]
if not raw:
raise StoreError(f"Run '{run_id}' carries no config provenance")
return json.loads(raw)

def run_tracking_field(self, run_id: str) -> str:
"""The canonical field this run detected/projected/tracked on.

Interprets both provenance formats — the single home for this
schema knowledge:
- current runs record ``global_.tracking_field``;
- legacy runs recorded ``global_.var_names.reflectivity`` (a name
knob whose value WAS the field the run tracked).
"""
cfg = self.run_config(run_id)
global_cfg = cfg.get("global_", {})
field = global_cfg.get("tracking_field")
if field:
return str(field)
legacy = global_cfg.get("var_names", {}).get("reflectivity")
if legacy:
return str(legacy)
raise StoreError(
f"Run '{run_id}' provenance records no tracking field "
"(neither global_.tracking_field nor legacy var_names)"
)

@staticmethod
def _run_from_row(row: sqlite3.Row) -> Run:
started = _parse_iso(row["started_at"])
Expand Down
Loading
Loading