REF: data store, Target V0.2; import-linter now 2-layers only - #75
Merged
Merged
Conversation
…rovisioning
{root}/registry.db + logs/ + collections/{id}/{catalog.db,products.db,objects/},
created only by adapt init. Store.open fails loudly on uninitialized or obsolete
(adapt_registry.db) roots. StoreRegistry records runs with full config_json/
config_hash/pipeline_version/environment, per-module executions (upsert), and
run events; collections carry source_kind + location. validate_collection
compares schema fingerprints (sha256).
Content-addressed objects: staged write (.staging/), sha256 + os.replace commit, catalog row + lineage edges registered atomically with the file. Catalog owns artifacts/artifact_lineage/scans/scan_products: scan_id = sha256(raw bytes)[:16], (run_id, scan_id) composite keys, scan completeness recomputed from REQUIRED_SCAN_PRODUCTS (gridded3d, segmentation2d, cell_stats) on every product link; auxiliary products never complete a scan. Cross-run raw lookup (find_artifact_by_scan) enables object reuse between runs.
…n products.db Module tables freeze on the first non-empty frame; identical schema snapshots in products.db and catalog.db. No ALTER TABLE ever: novel columns rejected after freeze. TableWriter stamps run_id/scan_id/scan_time from PersistenceMeta on scan-granular tables (modules never pass identity), derives granularity from the primary key (valid_time → time, scan_id → scan, else run), rejects duplicate keys and cross-module writes, and upserts on the declared key. TrackStore moves to products.db behind the same ledger (required ctor arg); its legacy schema-inference/ALTER path is gone.
StoreOutputRouter is the one dispatch from module outputs into the store: NetcdfArtifact → object store with identity attrs stamped and lineage to the raw volume (missing raw parent raises), ProductTableWrite → TableWriter, TrackTablesWrite → TrackStore; every write links scan_products. StoreExecutionHistory aggregates per-module timings into run_modules, routes warnings/errors to run_events, and finalizes the run row with scan counts.
StoreAcquirer is the only entry for raw data: idempotent per source_uri, reuses
already-stored objects across runs, registers the scan for this run, and mints
the queue contract {artifact_id, scan_id, scan_time, queued_at}. The local
replay source and the AWS downloader both go through it — the downloader stages
to a tempfile and never writes into the store itself; unparseable timestamps
fail loudly instead of being skipped.
…rocessor Orchestrator opens the store, registers the collection, begins/resumes the run (config lives in registry.runs.config_json — continuation reloads it from there), and finalizes through execution history. Processor keys on the queue's artifact_id, skips complete scans, routes all outputs through StoreOutputRouter, records failures via mark_scan_failed, and injects the in-memory grid to enrichment modules (no NetCDF re-read). PostProcessor runs post-hoc modules into an existing store attached to the original run_id; mask readers iterate the catalog. Modules no longer compute paths or carry identity columns; ingest gained CF-time re-encoding so pyart gridding works on live datasets. Rerun cleanup, config globbing, and output-dir setup are deleted.
adapt init creates the store layout (the only thing that ever does). --rerun is gone; postprocess resolves the latest run and its config from the registry. PNG plots are produced only when --plot-dir is given and land outside the data plane; PlotConsumer discovers artifacts through the catalog, keyed on scan_id, with no wall-clock fallback.
Everything scientists and the dashboard need from one read-only client:
collections/runs/latest_run/pipeline_progress; complete-scans-only listing and
scan_timeline (ScanRef spine); typed scans_since watermark for live follow;
open_scan_raster fully in-memory (zero retained file handles); cells/
cells_at_scan/track_history/tracks/track/track_graph (split-merge connected
component via recursive CTE); any product table with validated operator
filters {eq,gt,ge,lt,le,in}; arbitrary read-only SQL with the catalog ATTACHed;
annotations as the one write exception. Cross-run reads (run_id=None) union and
carry run_id per row. One lazy connection per db per collection; close()
releases everything. Synthetic-store test fixture builds through the real
writers — zero hand-copied DDL.
AppContext wraps one cached StoreClient per root (collection/active run/ timeline/open_raster); the Latest Scan tab runs on a ScanRef spine instead of file paths — every raster is fully in-memory, so frequent redraws retain zero file handles. Movie export and the TSE replay open and close one raster per frame. Volume stats and cells load through client.table/client.cells (raw sqlite3 gone from consumers); the TSE snapshot builds from client.cells/tracks. Store detection is registry.db; legacy roots surface the obsolete-layout error. The pipeline tab bootstraps fresh directories via adapt init (subprocess) and never clobbers an existing store.
AST fitness test over src/adapt/consumers/**: forbids sqlite3/duckdb imports, direct dataset/file opens (open_dataset, load_dataset, read_parquet, glob, rglob, iterdir, connect), and store-internal path literals (catalog.db, products.db, adapt_registry.db) — _utils.py alone owns store-root detection. Dependency ratchet updated: duckdb fully removed, ingest grid path no longer an uncontracted output.
DataRepository, RadarCatalog, RepositoryRegistry, ModuleOutputWriter, SqliteTable, the old OutputRouter/ExecutionHistory, RepositoryClient, FileProcessingTracker, output-directory setup, and both legacy schema files are gone, along with their ~28 test files. duckdb dependency dropped. Clean break: the only remaining reference to the old layout is the loud legacy-root detection error.
Integration gate (CI-excluded, ADAPT_REALDATA_DIR): adapt init → real 3-volume pipeline run → exact on-disk layout, every object cataloged with matching sha256 + lineage to its raw volume, complete-scan semantics (the first scan of a run stays pending — no pair, no analysis), fd-stable raster loops, watermark follow, tracks/filters/SQL, volume stats, movie with per-frame closes, TSE frame, after-the-fact xlma postprocess, validate_collection, loud legacy-root failure. The gate caught a bug synthetic data never hit: cell_stats froze cell_centroid_projection3_x as INTEGER on the first frame; a later scan's NaN promoted the column to float64 and the writer refused it mid-run. INTEGER and REAL now interchange under the frozen schema (SQLite affinity is lossless; NaN binds as NULL).
add_basemap cached tiles only on success, so an offline dashboard re-attempted the full network fetch on every 500 ms loop frame — on a black-holing network each frame stalls for the connect timeout on the UI thread. A failed fetch is now cached per extent exactly like a success (frame renders without the basemap); a zoom to a new extent retries.
Quick start now begins with adapt init (nothing else creates the store); outputs section documents the registry/collections/objects layout, content- derived scan identity, and completeness semantics; the API example moves from DataClient to StoreClient (typed filters + read-only SQL); troubleshooting covers the uninitialized and obsolete-layout errors. CLI reference gains adapt init and swaps --rerun/--no-plot for --plot-dir. Dashboard reference updates the store layout, StoreClient read path, and module map. Sphinx autodoc targets now point at the store modules (deleted-module targets would have broken the docs build).
docs/design/ was excluded by the *.md gitignore rule, so the architecture documents, per-module documents, decision records, and strategy notes were never visible to contributors. Track them explicitly (like USAGE.md), current as of the data store redesign: 00-08 layer documents updated to the store (02 runtime, 03 persistence rewritten as-built, 05 StoreClient, 08 phase checkoffs + deviations), the 2026-08-21 data store decision record (settled requirements, temporal-column rule, what real data taught us, verification, known gaps), module docs point at StoreClient, ARCH-001 marked resolved by the store refactor, strategy notes updated where they named deleted components. Contributors guide now opens with where to start reading.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR tags (required)
What changed
scan_idas the single identity key for the entire acquisition pipeline--plot-dirWhy
Centralizes data state, ensures idempotency at the source boundary, and eliminates duplication. Consumers now access data through a single queryable interface with full lineage and schema governance. We will test if this has any issues going forward. But we expect it to be better than current liberal schema.
Issues closed
Affects #70
Target version (required)
Aiming for:
v0.2.0Validation
AI usage & manual verification
Reviewer guide
adapt_core/store/→runtime/Impact note: Breaking change.
scan_idreplaces file-based identity. Existing runs need migration. Newadapt initestablishes store layout.