Skip to content

Test hybrid multi-head backend parity - #1537

Merged
aacostadiaz merged 1 commit into
ACEsuit:developfrom
darthjaja6:fix/mh1-cueq-parity-1298
Aug 17, 2026
Merged

aacostadiaz merged 1 commit into
ACEsuit:developfrom
darthjaja6:fix/mh1-cueq-parity-1298

Conversation

@darthjaja6

Copy link
Copy Markdown
Contributor

Summary

  • extend the CUDA multi-head parity regression from pure cuEquivariance to the hybrid cuEquivariance/OpenEquivariance conversion
  • exercise both DFT and MP2 heads
  • compare energy, forces, and stress against the e3nn reference

Context

Issue #1298 reports large and non-deterministic MACE-MH-1 disagreements with both CuEq and hybrid acceleration. The historical multi-head readout leakage is already fixed, and current tests cover pure CuEq, but the reported hybrid path had no equivalent regression coverage.

Validation

  • MACE_REQUIRE_CAPS=gpu,cueq,oeq pytest tests/backends -q — 52 passed
  • Black, isort, pylint, and git diff --check

@darthjaja6
darthjaja6 force-pushed the fix/mh1-cueq-parity-1298 branch from 8164244 to 56e58f1 Compare August 8, 2026 19:17
@darthjaja6
darthjaja6 marked this pull request as ready for review August 8, 2026 20:04
@aacostadiaz

Copy link
Copy Markdown
Collaborator

@cursor review

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown

Bugbot couldn't run

Bugbot is not enabled for your user on this team.

Ask your team administrator to increase your team's hard limit for Bugbot seats or add you to the allowlist in the Cursor dashboard.

@aacostadiaz
aacostadiaz merged commit 31de5b4 into ACEsuit:develop Aug 17, 2026
27 checks passed
rosecers added a commit to cersonsky-lab/mace-cg that referenced this pull request Sep 4, 2026
* Add rank-2 rigid node feature ablations

* Add rank-2 rigid node feature ablations

* tqdm in eval

* Fix Zenodo badge (#1630)

* Fine tuning with additional arbitrary embeddings.  (#1420)

* inital commit

* warmup/fixes

* deals with various embedding specs

* resolve

* nit

* head slicing

* replace ordering with slices

* pre-commits

* Align joint-embedding head transfer with the model's own column layout

The head columns of GenericJointEmbedding.project[0] sit at the cumulative
offset over every spec, so building the model-side offsets from the matching
specs alone shifted the copy whenever a non-matching spec came first. Build
them over all specs, mirroring the foundation side.

The spec-by-spec transfer was also undone by the shape-matched state-dict
copy at the end of load_foundations_elements, which re-aligned the head
columns positionally whenever both heads had the same shape. Skip
joint_embedding there, as readouts already are.

Also use the public named_children() accessor rather than _modules, which
tripped pylint's protected-access check, and add tests for a foundation and
target whose embedding specs only partly overlap.

---------

Co-authored-by: DrDavie1 <101512473+DrDavie1@users.noreply.github.com>
Co-authored-by: Alejandro Acosta <127198532+aacostadiaz@users.noreply.github.com>

* Fix cuEquivariance fusion device detection (#1641)

* Test hybrid multi-head backend parity (#1537)

* Fix bug in symmetric contraction (#1622)

* Fix bug in symmetric contraction

* Pin the symmetric contraction weight slot to the order forward() contracts

path_weight is indexed ascending by correlation order while self.weights is
built descending, so the zeroing flag needs remapping between the two. Getting
it wrong zeroes a valid correlation order and leaves the unreachable one
trainable, which changes trained numbers with no error to point at.

The case is reachable whenever hidden_irreps asks for an l that nu = 1 cannot
produce, for instance max_ell=1 with max_L=2. The test builds the smallest
Contraction of that shape and asserts each weight's zeroed state against the
nu that forward() actually pairs it with.

---------

Co-authored-by: Alejandro Acosta <127198532+aacostadiaz@users.noreply.github.com>

* Add learned full rank-2 rigid body tensor mode

* Add the golden harness and the first two anchors (#1643)

* Anchor the numerics on committed goldens instead of a retrained fixture

Every claim the rewrite will have to make -- that a converted checkpoint
still produces the same energy, that an accelerated kernel agrees with the
reference one, that the new training loop is not worse -- needs a number
that was recorded while the current stack was the live one. Today there is
none. The closest thing, the session-scoped trained_tiny_model_path fixture,
retrains on every pytest session, so it can tell you that two runs in one
process agree and nothing at all about two machines a year apart.

This adds tests/golden: a comparison harness, six structures, two tiny
anchor checkpoints and their outputs. The harness depends on the standard
library, numpy and ase and never imports the package under test, because the
suites that will consume it are forbidden from importing it and a single
convenience import would make the shared machinery unusable to exactly the
tests it exists for; two tests enforce that, one by grepping the source and
one by importing the module with the package blocked on sys.meta_path.

The tolerance table lives in the harness and nowhere else, and it is
reconciled against the two that already exist rather than becoming a third
that quietly disagrees with them. The fp64 row is the 1e-6 that
backend_parity already uses for implementation-against-implementation
agreement; the fp32 row adopts the 5e-5 absolute floor the polar regression
measured, because that file documents 5e-6 failing in CI and inventing a
different number here would throw that evidence away. A test scans every
other module in the directory for a literal atol, so the next golden ticket
cannot start a fourth table by accident.

There are two anchors because the short-range repulsion term enters the two
model classes differently and only anchoring both turns that into a number:
plain MACE appends the pair term next to e0 and never scales it, while
ScaleShiftMACE folds it into the sum that goes through scale_shift. On the
short-dimer fixture, deleting the term moves the total by exactly 1.000000x
the raw pair sum in one and 0.478465x -- the model's own scale -- in the
other, and both ratios are asserted. The plain anchor is built by direct
instantiation rather than by the CLI, because --model MACE returns a
ScaleShiftMACE with the scale taken from the dataset std, so a CLI recipe
would silently anchor the wrong class; the same whitelist is why the
accelerated-backend work has to run on the ScaleShiftMACE anchor, and that
refusal is pinned here as a contract instead of being discovered later.

The six fixtures are chosen so that each reaches a different branch of the
neighbour-list layer, which is what decides the cell a stress is divided by.
The two slabs differ only in whether the vacuum row is all zeros, and their
stresses consequently differ by exactly the volume ratio -- the branch that
would divide by a zero volume if the row were not patched.

The checkpoints and structures are committed, which the global *.xyz and
*.model patterns would have refused, so the ignore file gains a narrow
exception: a golden that is not in the clone is not a golden. Regeneration
goes through one script that refuses to run without an explicit
acknowledgement, and is byte-identical run to run, so a diff in these files
always means something moved. tests/golden joins the ci-core unit job rather
than a nightly, because a change that breaks the numerical contract should
fail the pull request that made it.

* Refuse to snapshot an output the golden schema does not know

The harness dropped unknown keys twice over -- once while scraping the
calculator's results dict, once while encoding the snapshot -- and said
nothing either time. That is the exact failure the inventory work exists to
prevent, reproduced inside the machinery meant to prevent it.

It was not hypothetical, because the registry was keyed on the names the
model's forward returns and the ASE calculator writes four of them
differently: LES_alphas / LES_kappas / bec / MACE_magmoms against
latent_alphas / latent_kappas / BEC / equilibrated_magmom. A LES or magnetic
golden taken through the calculator therefore recorded energy, forces and
stress, discarded the rest, and committed a reference that claimed to pin the
family. It would have passed forever while pinning nothing.

So an unrecognised key now raises, naming the key and the three ways to
resolve it. The alternative spellings live in tests/golden/calculator_keys.py
-- the harness itself must stay ignorant of the framework, and a test
re-derives the calculator's key set from the source so a fifth divergence
cannot appear unnoticed. Skipping is still possible but only one key at a
time, in writing: the committee spread statistics are on an allowlist with a
reason each, since their leading axis is the committee size and stress_var is
not even the same layout in the two calculators.

Three more holes closed while the schema was open:

* edge_forces and hessian could not be expressed at all -- their leading axes
  are the edge count and 3N, and every per-atom kind pins that axis to the
  atom count -- so the only thing the harness could have done with them was
  drop them. They get kinds. The remaining calculator outputs (energies,
  fukui_functions, stresses, virials, the polar family) get channels.
* the inputs block was written into every reference and never compared, so a
  snapshot with every magmom rewritten to [9,9,9] matched. Inputs are now
  compared in both directions and unconditionally: a different input is a
  different measurement, not a drift, and there is no flag to turn it off.
* magmom was read from ase's initial-moments attribute, which no forward pass
  looks at; the magnetic calculator reads atoms.arrays["REF_magmom"]. The
  harness recorded provenance from an array the model never sees. It now
  reads the same one, and refuses a structure carrying only the ase attribute
  rather than guessing.

The two anchor references are regenerated because "energies" is now visible
to the schema. Nothing else moved: every pre-existing channel is byte
identical, that one channel is the whole diff.

* Cover the model surface too, not just the calculator's

The alias map and its guard test were both derived from
mace/calculators/mace.py alone, so all 31 calculator result keys resolved
and 13 of the 43 model-forward keys resolved to nothing. That is not a
cosmetic gap: edge_forces and hessian are returned by every energy model
and by no calculator, so any golden that wants either has to go through
the golden_outputs hook -- and a hook hands the harness the forward's
whole dict, all 43 keys. The first such golden would have died on a key
the schema had never been shown.

Two of the thirteen were the same name-divergence defect one layer up,
and the obvious repairs are both wrong. The model emits atomic_stresses
as (n_atoms, 3, 3); the calculator calls the same physics `stresses` and
stores it Voigt-6. A plain alias is a shape failure. A channel each is
worse than the bug: both would hold the same quantity, and a golden taken
through the model route and one taken through the calculator route would
then never compare -- the silent split, arrived at deliberately instead
of by accident.

So a registration now names its surface, and may carry a conversion. The
per-atom stress is canonicalised on ingest: one channel holding the 3x3,
with the calculator's Voigt-6 expanded as it arrives. Voigt-6 cannot
carry an asymmetric tensor, so that direction was measured rather than
assumed. It is lossless because get_atomic_virials_stresses symmetrises
explicitly at mace/modules/utils.py:382; on the tiny_scaleshift anchor
over all six fixtures in float64 the asymmetry, the Voigt round trip and
the gap between the two routes are each exactly 0.0. A test re-measures
all three, so dropping that symmetrisation fails there rather than in a
golden quietly pinning a symmetrised copy of an asymmetric tensor.

Naming the surface also settles a collision a flat map cannot express:
`virials` is the graph-level virial in every forward and the per-atom
virial in the calculator's results, which has no key for the graph one at
all. The channel was declared per-atom, so a model-route snapshot of it
would have failed on the shape.

The guard test now derives its expectation from both surfaces -- the
calculator by regex, the forwards by AST, since the models build their
return dict three different ways and a regex finds one of them. A second
test guards the extractors themselves, because the failure that matters
is finding nothing and reporting coverage.

On scf_steps: the note that nothing emits it is wrong.
mace/modules/extensions.py:2102 sets it on MagneticSCFMACE's output, next
to scf_energy_history and equilibrated_magmom. It is one of a family of
three histories, all metadata for the same reason -- their extent is the
iteration count, so a run that converges one step sooner changes them
without changing any physical number. The fixed point itself stays
pinned.

Two input-side holes in the same module, fixed in the same pass because
they are the same kind of silence:

An input the model reads and the harness does not record is as bad as an
output nobody pins, and the fixed table of spellings could not know what
a given instance was built with. magmom_key and charges_key are
constructor arguments, so a structure whose moments live elsewhere
matched nothing, was recorded as no magmom at all, and then compared
clean -- there was nothing on either side to disagree about. The harness
now asks the object it was handed. A configured key replaces the defaults
rather than joining them: the entries in INPUT_ARRAY_KEYS are synonyms
whose disagreement is an error, while a stale REF_magmom beside a
configured key is an array the evaluation ignored, not a second opinion.

And inputs were compared at whichever row the outputs used, which lets a
change smaller than the output tolerance through. At the fp32 row a 2e-3
nudge to a 2.2 muB moment -- bcc iron, not a contrived magnitude -- fits
under the bound, and the snapshot then passes while having been taken at
moments the reference was not. There is nothing to trade off: an input is
read verbatim off the committed fixture rather than computed, so two
reads agree exactly or the fixture changed. Inputs now compare at a new
`exact` row, whatever the outputs use.

Measured while doing this: the model's forward reads the process-wide
default dtype, not only AtomicData. Building the graph in a float64 scope
and running the forward outside it costs ~2e-8 relative -- under the fp64
row, so it reads as agreement while making a bit-exact comparison
impossible. _batch now builds inside the scope instead of casting a
float32 graph up afterwards.

* Say so when the forced stress does not land in the results dict

The Voigt-to-3x3 conversion now happens once, in the registered
calculator-surface rule, so `_evaluate` calls `get_stress` only to make
the calculator compute it and then reads the value back out of
`results`. That is the right shape -- one representation change in one
place, rather than the accessor's return value and the results entry
being converted separately and free to drift -- but it introduces a way
to lose the stress silently: any calculator that computes it without
caching it would produce a snapshot with no stress at all for a periodic
structure, and nothing downstream distinguishes that from a model that
has none.

ase's own machinery always caches, so this is a guard rather than a fix.
It costs one lookup and turns a silently thinner reference into a
failure at the point of evaluation.

* Make MACECalculator's charges_key actually select the charges array

KeySpecification.arrays_keys maps property name -> atoms.arrays key:
config_from_atoms iterates `for name, atoms_key in arrays_keys.items()`
and stores atoms.arrays[atoms_key] under properties[name]. The calculator
wrote the pair the other way round, as {charges_key: "charges"}, so
"charges" was never a property name at all. config_from_atoms then looked
for an atoms.arrays field literally called "charges" (ASE calls that one
initial_charges, and nothing in MACE writes it), found nothing, and
AtomicData.from_config fell back to its zeros default.

The visible effect is that charges_key did nothing: whatever a user
passed, the batch carried zeros. DipoleMACE and EnergyDipoleMACE feed
data["charges"] into compute_fixed_charge_dipole, so their fixed-charge
dipole baseline was always zero and the predicted dipole was missing the
sum(q_i r_i) term the model was trained to be corrected by.

MagneticMACECalculator already had the direction right, which is why only
the plain calculator was affected.

Training is unaffected: --charges_key goes through
update_keyspec_from_kwargs, which builds arrays_keys["charges"] = <key>.

* Derive the input and eval surfaces too, and make the guard a real gate

The output surface got a derivation from mace/ last round; the input surface
and the eval CLI were still hand-written tables, and both were already
wrong in the way a table is always wrong -- silently, while reporting
coverage.

The input surface. INPUT_ARRAY_KEYS named two constructor arguments,
magmom_key and charges_key, and by naming them said nothing about
external_field, total_charge or total_spin. A calculator-route reference
recorded no field, and two references taken at different external fields
compared clean -- the field enters the energy and the BEC force correction
(mace/calculators/mace.py:809-859), so those were different numbers wearing
each other's name. Enumerating the missing three would have been the same
mistake with a longer list, so the harness now reads the mapping the reader
is handed: both calculators keep a KeySpecification's info_keys/arrays_keys
on the instance and config_from_atoms iterates them, so register_keyspec_
source reads the reader and the next constructor argument is picked up with
no change here. A property no channel covers now fails the snapshot instead
of being skipped; a training label is dismissed in writing; external_field,
which is written into the batch after the graph is built and appears in no
array, is recorded from the instance; and an input the structure carries
under a key nobody reads is refused, which generalises the older
ase-initial-moments refusal. The keyspec is read after the evaluation, not
before, because MACECalculator fills in its charges entry inside the call.

The eval surface. mace_eval_configs writes thirteen names onto its
structures and three resolved to nothing, so any ticket pinning it stopped
at authoring time. BO_contributions and node_energies are renames of
contributions and node_energy; BEC arrives flattened to (n_atoms, 9) and is
unflattened on ingest, the same treatment the calculator's Voigt-6 per-atom
stress gets. descriptors is decided rather than deferred: the channel is
node_feats and only the unaggregated per-atom form is pinnable, because the
other two --descriptor_aggregation_method settings are reductions of it and
per_element_mean is a dict keyed by chemical symbol that no kind describes.
The eval stress needs no alias, which is the surface scoping earning its
keep: it is already 3x3, and a flat map would have expanded it again.

The guard. The extractors are now in tests/golden/surface_scan.py and cover
all three surfaces. The calculator scan follows self.results[k], the aliased
local in both assignment directions, .update() with a literal or a named
dict, a whole-dict assignment, and the committee suffixes -- whose bases
come from the results_store_ensemble set literal the source guards them
with, not from four literals copied into the test. The model scan discovers
the files instead of naming two of them: four define a forward returning a
dict, and the two that were missed are the deployment wrappers a deployment
golden evaluates. total_energy_local, from the LAMMPS wrapper, was the one
key in the package that no channel described; it is a channel now, not an
alias for energy, because it excludes the ghost atoms' site energies.

Every scan also reports the writes it could not resolve, because an
extractor that finds nothing reports perfect coverage. One is allowed to
stand, with a reason: the torchsim wrapper forwards the wrapped model's dict
key by key, so its key set is the model surface's. Synthetic sources
exercise each write form directly, so a form stays tested after the real
file stops using it.

Includes the MACECalculator charges_key fix as a prerequisite: the input
derivation reads arrays_keys as property -> key, which is the direction
config_from_atoms iterates and the direction the calculator now writes.

One table survives and cannot be derived: the fallback spellings in
INPUT_ARRAY_KEYS / INPUT_INFO_KEYS, because deriving them means reading the
framework's key tables and harness.py may not import the framework. A guard
test derives the truth from DefaultKeys, update_keyspec_from_kwargs and each
calculator's own __init__ defaults, and asserts the table matches -- which
is how Qs joined REF_charges there.

* Let each family of goldens own its regeneration, so two of them can land independently

regenerate.py enumerated its targets in three places at once: a dict, the
argparse choices and the module docstring. README.md enumerated them in three
more. A family of goldens is added by one branch at a time, so every pair of
branches that added one edited the same six lines and conflicted on both
files.

The driver now discovers targets/*.py and reads ORDER, HELP, run() and an
optional IN_ALL off each module, and the README describes the rule instead of
listing the families. A family is a module in targets/ and a page in docs/,
and touches nothing shared.

--target all regenerates every committed artifact byte-identically, anchor
retraining included.

* Checkpoint interaction neighbor normalization (#1538)

Co-authored-by: Alejandro Acosta <127198532+aacostadiaz@users.noreply.github.com>

* Write ML-IAP results back through the coupling LAMMPS actually has (#1644)

* Write ML-IAP results back through the coupling LAMMPS actually has

`_update_lammps_data` wrote its results through the KOKKOS ML-IAP coupling's
API: it read `data.eatoms` and called `update_pair_forces_gpu`. The plain
coupling in `src/ML-IAP/mliap_unified_couple.pyx` has neither. There `eatoms`
is a `write_only_property`, and the only force entry point is
`update_pair_forces`, which takes a contiguous host `double[:, ::1]`.

conda-forge builds every CPU `lammps` with `PKG_KOKKOS=OFF`, so that is the
coupling any stock build gets, and the nightly real tier had therefore never
passed a single test. The symptom points nowhere useful: the model evaluates
fine, then LAMMPS reports `mliap_unified.cpp:71 compute_forces failure` over a
Python `property 'eatoms' of 'MLIAPDataPy' object has no getter`. Scoping the
tier to a single interaction layer did not help, because this is a second
KOKKOS-only call, downstream of the `forward_exchange` one already guarded.

The writeback now branches on the coupling. Both branches reach the same C
`update_pair_forces`; only where the pointer comes from differs, so results are
unchanged where KOKKOS is present. The float32 special case collapses into an
unconditional `.double()`, since LAMMPS reads both buffers as double whatever
the model's dtype is.

A mismatch between `nlistatoms` and `nlocal` now raises here, naming both
counts. The setter fills exactly `nlistatoms` doubles and Cython rejects a
length mismatch from inside the `.pyx`, naming neither; the KOKKOS branch has
always assumed the same equality silently.

`test_mliap_writeback.py` pins both paths with stubs that reproduce each
`.pyx`'s property shape, which is what the bug turned on: `eatoms` as
`property(fget=None, fset=...)`, and a force setter that accepts only a
C-contiguous float64 array. Without the fix it reproduces the nightly failure
exactly. `StubMACE` moves into `_harness.py` so both LAMMPS contract files
share one definition.

* Let the LAMMPS real tier fail the nightly

`continue-on-error: true` is why nobody noticed this job failing every single
night from the day it first ran a test: the nightly reported success by
absence. A job that cannot turn the run red is indistinguishable from a job
that does not exist, which is the opposite of what a CI test is for.

The guidance in tests/integrations/README.md that told a new real tier to start
soft and get promoted later goes too. LAMMPS is the argument against it.

* Pin the radial, neighbour list, key parsing and batching behaviour by value (#1646)

* Add a tolerance row for closed-form checks instead of scattering numbers

The pure-math units -- radial bases, cutoff envelopes, distance transforms
-- are checked against the analytic expressions they implement, which is a
different measurement from every row the table had: nothing crosses a
machine, a device or a kernel, only the order of the fp64 arithmetic
differs. Without a row for it each of those tests has to invent its own
atol, which is exactly the drift the single table exists to prevent, and the
numbers that get invented are the loose ones borrowed from a golden (1e-6),
three orders too slack for a formula check.

The bound is measured, not guessed: over every case in
tests/unit/test_radial.py the worst deviation is 1.8e-15 absolute and
7.1e-15 relative, so 1e-12 clears it by three orders while still failing any
real change to a formula. The row is deliberately not usable as a golden
row, and the guard test names it so a fifth row cannot appear silently.

* Pin the radial, parsing, neighbour-list and batching behaviour by value

These four layers are pure functions that port into the rewrite nearly
verbatim, so what protects them is not that the current code runs but that
the numbers are written down. Until now they were not: the radial module had
three tests at rtol=1e-2 and no coverage of the Chebyshev, Gaussian or
polynomial-cutoff paths, the brute-force neighbour reference was a private
closure in one test file, and the returned cell -- the thing a stress
divides by -- was asserted in one regime out of three.

The neighbour reference moves to tests/neighbour_oracle.py so the rewrite's
pluggable backends can be measured against the same oracle rather than each
growing its own, and it compares integer unit shifts rather than
displacement vectors, which removes a tolerance nobody should have to pick.
All three returned-cell regimes now have a test each plus one that they are
actually three, because collapsing any two of them is the simplification a
port reaches for first and it silently rescales every slab stress. The
committed golden fixtures are read through their manifest, so a structure
retagged there fails here instead of quietly pinning something else.

Three findings came out of writing it, all pinned as characterization rather
than fixed: true_self_interaction does nothing, because the matscipy call is
made without self_interaction and the edge it claims to keep is never
produced; the default dipole key cannot be read back from extxyz at all,
since ase reserves that name as a calculator property, which is why every
dipole workflow in the tree overrides it; and the non-linear first
interaction block cannot be TorchScripted, so it can never reach the LAMMPS
export path.

The three superseded radial tests are gone rather than kept beside the new
ones -- each is covered at a real tolerance now, in float64 and against the
published formula instead of against a remembered output.

* Fix LAMMPS export for models saved before avg_num_neighbors became a buffer (#1645)

* Adopt a pickled avg_num_neighbors as the buffer it now is

Making `InteractionBlock.avg_num_neighbors` a registered buffer with a declared
`torch.Tensor` type gave TorchScript something to check, and gave
`_load_from_state_dict` a legacy path for checkpoints loaded as state dicts. It
left the other way MACE ships a model uncovered: `torch.load` of a whole pickled
module, which is what the ASE calculator, every `mace/cli` tool and the
pretrained artifacts under `mace/calculators/foundations_models/` all use.
Unpickling restores `__dict__` directly, so neither `__init__` nor
`_load_from_state_dict` runs and the block keeps the plain Python float it was
pickled with.

Eager forward passes never notice -- a float promotes against whatever it
divides -- so the whole golden tier passes on such a model. Scripting it does
not:

    RuntimeError: Could not cast attribute 'avg_num_neighbors' to type Tensor:
    Unable to cast 4.59375 to Tensor

which is `mace_create_lammps_model` refusing every checkpoint written before the
buffer landed: everything users already have on disk, every published foundation
model, and the three `.model` files tracked inside this repo. Reproducible on a
clean tree with

    python mace/cli/create_lammps_model.py tests/golden/models/tiny_scaleshift.model

CI could not see it. The only export test trains a model first, so its blocks
are constructed through `__init__` and get the buffer; it passes either way. The
uncovered case is exactly "export a checkpoint you did not just build", which is
what deployment does.

So `__setstate__` promotes the float, and the dtype comes from the module's own
weights rather than from `torch.get_default_dtype()`. That distinction is the
substance of the fix, not a detail: as a float the value participated at the
precision of whatever it was divided into, so taking the process default would
silently round the normalization of a float64 model loaded under a float32
default -- a changed number, in a factor that scales every message. The new
test measures that rather than asserting it, comparing the promoted buffer's
messages against the float's bit for bit.

Five of the eight new cases fail without the change, including the scripting
one. The three that do not are guards on what must stay still: a model pickled
after the buffer landed is left alone, and the promotion is not allowed to move
the numbers at either precision.

* Put the promoted buffer on the device the weights are on

A Python float has no device, so the attribute this replaces was indifferent to
where the model lived. A buffer is not, and the promotion built it from a dtype
alone: `torch.load(..., map_location="cuda")` puts the weights on the
accelerator, and the new buffer stayed on the CPU. The module was then split
across two devices for any caller that does not go on to `.to()`, which is the
raw-model foundation loaders and anything handing a freshly loaded model
straight to DDP.

The division itself would have survived. `avg_num_neighbors` is zero-dim, so
torch's cpu-scalar rule applies and `cuda_tensor / cpu_scalar` is legal --
verified on the analogous MPS path. What does not survive is everything that
assumes a module's tensors are colocated.

Both the dtype and the device now come from the same weight tensor, which is
the rule `set_avg_num_neighbors` two methods up already follows.

Two cases cover it. The plain one asserts colocation on any host, where it
cannot fail; the `gpu` one loads through a `map_location` and checks the buffer
followed, which is where the assertion has teeth.

* Scope the device assertion to the buffer this change owns

The gpu case asserted that every buffer in the model sits on the parameter
device. Two do not, and neither is this change's doing: e3nn keeps its Wigner 3j
tables inside `TensorProduct._compiled_main_left_right`, a `torch.fx.GraphModule`
that unpickles by regenerating itself, so `_w3j_1_1_2` and `_w3j_1_2_1` come back
on the CPU whatever `map_location` said. It reproduces on a model pickled long
after the promotion existed, with none of this code on the path.

So the assertion now names the buffers this change is answerable for, across
every interaction layer rather than only the first, and the docstring records why
the broader claim is not made -- it is the one a reader reaches for next, and it
fails for somebody else's reason.

The two assertions that measure the fix passed on the MI210 run that caught this.

* Add learned full rank-2 rigid body tensor mode

* Remove the L-BFGS reload that cannot run (#1653)

`run_train` carried a second checkpoint load for L-BFGS resumes, placed after the
optimiser swap and guarded by a flag set only when *both* earlier `load_latest`
calls raise. They cannot both raise, so the block never executed.

The intended trigger is knowable: a checkpoint whose L-BFGS optimiser state the
freshly built Adam cannot accept. That raises a `ValueError` which the checkpoint
loader catches and downgrades to a warning, so the load returns normally and the
`except` never fires. Nor can absence of a checkpoint trigger it -- `load_latest`
returns `None` in that case rather than raising, so the second call is not even
attempted. Instrumenting both points and running three resume scenarios,
including `--restart_latest --lbfgs` after a stage-two run that left both a
stage-one and a stage-two checkpoint, the flag was never set once.

Deleting it changes nothing a user sees: the same scenario before and after
produces byte-identical logs, resuming from the stage-two checkpoint at the
recorded epoch with a fresh L-BFGS optimiser.

The flag was doing one thing incidentally, and it was harmful. Without `--lbfgs`,
both loads failing left a run that had asked to resume starting from scratch in
silence, because the flag it set was only ever read inside the `--lbfgs` branch.
That case now says so.

* Fail where the batch size is wrong, not four hundred lines later (#1648)

A training set smaller than one batch cannot train: `drop_last` throws away the
incomplete final batch, that is the only batch, and the loader comes out empty.
The run then died in `compute_avg_num_neighbors` on a `torch.cat()` of an empty
list, naming neither the batch size nor the set. `--batch_size` defaults to 10,
so any set of fewer than ten configurations reached this on default flags, which
is an ordinary size for a first experiment or a small fine-tuning set.

There was already a diagnosis for it, and it could not serve. It sits behind a
check for ase-readable inputs, so the hdf5 and lmdb paths never see it, and it
only logs -- the run carried on regardless and failed elsewhere. Testing
`len(train_loader) == 0` instead tests the condition rather than a proxy for it
and covers every data path.

Two routes reach it and they need different advice. Without a distributed
sampler the loader's own `drop_last` is responsible, so the batch size is what to
change. With one, the loader keeps its partial batch and the *sampler* holds the
`drop_last`: it splits the set over the ranks and discards the remainder, so a
set smaller than the world size leaves a rank with nothing. Blaming the batch
size there is not merely unhelpful but can be false -- two configurations over
four ranks with `--batch_size 1` reported a set "fewer than --batch_size (1)".

The pre-existing log had the same blind spot in a milder form, asserting that the
full training set survives as one batch whenever the incomplete batch is kept.
That holds for `--lbfgs` and not per rank under distribution, so the three cases
are now separate, and only the fatal one is an ERROR.

Four cases: the set that cannot train fails naming `--batch_size`, the same set
trains under `--lbfgs`, a batch that fits is untouched, and a rank with no data
names the sampler instead. The first and last fail without this change.

* Make pseudolabelled replay data usable, and refuse it when it is not (#1649)

* Give a generated label a weight that lets it train

Pseudolabelling wrote the label and left the loss weight alone, for every
property except stress. That is fine when the replay file already carried
labels, and it makes the whole operation pointless in the case it exists to
serve: an unlabelled replay set arrives with every property weight at zero, so
the generated labels land on configurations that contribute nothing to the
gradient.

The reported consequence is worse than "it does not learn". Measured on a
one-epoch multihead run over an unlabelled replay set, the pt_head reports a
loss of exactly 0.0 with rmse_e_per_atom and rmse_f both absent -- the best
number on the page, from a head doing none of the work. With the weights set it
reports 1.8e-6 and real metrics, matching what the same run produces from a
labelled replay file to within a few percent.

The weight is a floor rather than an assignment, so a value the user chose
deliberately survives. `stress` already did exactly this; the rule is now shared
rather than stated once and forgotten five times.

This does not touch the reader that lets an unlabelled file through in the first
place. That behaviour is deliberate -- the labels are about to be generated --
and it is only harmful because the generated labels were unusable, which is what
this fixes.

* Refuse a partly relabelled replay set instead of mixing two label sources

A batch that failed during pseudolabelling had the file's own labels substituted
for it, and the loop carried on. One call could therefore return a replay set
carrying two levels of theory, with nothing in the return value to say which
configurations held which -- and the summary line went on reporting every
configuration as relabelled, because it counted the list rather than the work.
Measured with a transient failure injected into batch 2 of 3: twelve
configurations back, four of them still on file labels, and
"Generated pseudolabels for 12 configurations" in the log.

Controlling which labels the replay head sees is the entire point of replay, so a
partial result is not a weaker version of the right answer, it is a set nobody
can use. It raises now, naming the batch and the configuration range. A caller
that wants to tolerate transient failures can retry the call; one that cannot
tell the two label sources apart could do neither.

The one `logging.error` that used to mark this was the only trace, and it is easy
to lose in a training log that runs to thousands of lines.

* Keep the two invariants the swallowed exception used to preserve

Refusing a partly relabelled set means `generate_pseudolabels_for_configs` can
now leave early, and two things were relying on it never doing so.

**The model's gradients.** The function clears `requires_grad` on every parameter
and restored it after the batch loop, which the old swallow-and-continue always
reached. The raise is a non-local exit past that point, so a failed batch handed
the caller back a model with every parameter's gradient disabled -- and the model
is the caller's, used for more than this. The restore is in a `finally` now.

**The caller's two splits.** `apply_pseudolabels_to_pt_head_configs` replaced
`collections.train` as soon as train relabelling succeeded and only then started
on valid, so a failure there left train on foundation labels and valid on the
file's, returned False, and had `run_train` report that it was "continuing with
original configurations". That is the same mixing of label sources this refuses
inside a single split, one level up, and with a message asserting the opposite.
Both splits are relabelled before either is committed, so the caller's claim is
true again: either both move or neither does.

Two tests, one per invariant, both failing without this. The stress condition
lost a nesting level to keep the function inside the block-depth limit; the
`finally` added one.

* Fix eval_configs on ragged structure sets and on a dtype the checkpoint disagrees with (#1654)

* Write node energies per structure, as the descriptors beside them already do

`--return_node_energies` failed on any set whose structures differ in size:

    ValueError: setting an array element with a sequence. The requested array
    has an inhomogeneous shape after 1 dimensions.

The per-structure splits were accumulated with `append`, so the list held one
entry per *batch* rather than per structure, and the `np.concatenate` that
followed had to build a rectangular array out of it. That succeeds only while
every structure has the same number of atoms, which is why a uniform set worked
and a mixed one did not.

`--return_descriptors` sits four lines above and had this right already, with the
reason written down -- it accumulates with `extend` and asserts on the length,
"no concatentation - elements of descriptors_list have non-uniform shapes". Node
energies now do the same, which makes the fix a one-word change plus deleting the
concatenation.

Asserted by value rather than by exit code: on a 3, 6, 3 atom set each
structure's array has that structure's length and sums to the total energy
reported for it. The uniform case is bit-identical to before, so the path that
worked is untouched.

* Reconcile the evaluation CLI's dtype with the checkpoint, as the calculator does

`--default_dtype float32` against a float64 checkpoint died several frames inside
a scripted tensor product:

    RuntimeError: both inputs should have same dtype

Nothing in the output named the flag, the checkpoint, or a mismatch, so there was
nothing to tell the user what to change. Meanwhile the ase calculator, handed the
same model and the same request, warns and converts and returns an energy. Two
shipped inference routes, one request, opposite outcomes -- and one of them
already had the answer.

The CLI now does what the calculator does, with the same warning, so the choice
of route stops being a choice about whether the run works. The test asserts the
two agree to the bit rather than checking a number of its own, because the
agreement is the contract.

* Assert the dtype conversion, not a float32 bit pattern

The route-agreement assertion demanded that the evaluation CLI and the ase
calculator return the same float32 energy to the bit. They do here and they did
not in CI, by 1.1e-7 relative -- less than one float32 eps.

The reason is structural rather than a flake to be re-run: the CLI runs in a
subprocess and the calculator runs in the test process, float32 kernel selection
on CPU depends on the shapes a process has already evaluated, and under xdist the
worker has evaluated a great many. Loosening the comparison alone would not have
been worth much either, since a run that failed to convert and stayed at float64
would agree to about that same order.

So what is asserted exactly is the thing the change actually does: the run
succeeds, and its output carries the "does not match model dtype" warning, which
only the conversion path emits. The energies are still compared, at rel=1e-5, as
a sanity bound rather than as the contract.

* Pin the training, evaluation, export and fine-tuning contracts from outside (#1655)

* Give the training contracts a set whose loss can actually go down

The end-to-end training comparison needs a dataset, and the shared water
fixture is not one: its labels are noise, so a full-epoch quasi-Newton step
fits them exactly and the held-out loss rises monotonically. Measured on the
twenty rattled waters, an L-BFGS run goes 23.8 -> 47.0 -> 112.6 -> 130.1 over
four epochs. No tolerance makes that read as a working optimiser, so any
contract asserting that training improves anything would have had to assert
something weaker instead.

So the committed set is labelled by one closed-form pair potential, with the
forces, the stress and the coordination-dependent partial charges all
differentiated from the same expression. It covers isolated atoms, molecules,
slabs and bulk cells because each reaches a different neighbour-list regime,
and it carries both electrostatic labels so a long-range model has something
to be wrong about. The analytic gradients are checked numerically rather than
trusted, both in the recipe and in a test, because a committed dataset whose
forces do not match its energies teaches a rewrite the wrong physics and
nothing else in the tree would notice.

The dipole goes under REF_dipole and not the package's default key: ase
recognises "dipole" as a calculator property, so a value written into info
comes back from an extxyz round trip inside calc.results, where the parser
never looks. The default name would have produced a label no training run
could read.

* Pin the training, evaluation and calculator command lines from outside

The CLI-observable behaviour of the legacy stack is the user contract the
rewrite has to honour, and almost none of it was asserted anywhere. These
tests drive the installed entry points and read only what a user can see, so
the live parity run can re-execute them against a different engine without
editing a line.

One surface has no coverage at all: mace_select_head is an installed entry
point with no test, and it is the tool that decides which head ships. The
other two here are already exercised, and not where it counts.
test_run_train_lbfgs drives the CLI with --lbfgs and compares energies, but
nothing touches the drop_last consequence, which is not cosmetic -- with a
batch larger than the training set the default regime drops the only batch and
refuses before training starts, while the identical command with --lbfgs
trains normally -- nor the one-step-per-epoch regime that --lbfgs selects.
test_calculator_padding already compares a padded run against an unpadded one;
what is added here is that comparison through the golden harness, channel by
channel with units and shapes checked, rather than field by field by hand.

Three recorded defects are pinned as they stand rather than fixed, because
this characterises the frozen stack: --return_node_energies dies on any set
whose structures differ in size, --default_dtype float32 against a float64
checkpoint fails several frames deep in a scripted tensor product with
nothing naming the flag, and the post-swap reload the L-BFGS resume path
carries cannot be reached -- the exception it waits for is caught and
downgraded inside the checkpoint loader, and the two load attempts partition
the checkpoints so one of them always succeeds. Each of those tests says in
its message that a failure means the defect was fixed and the test should go.

* Pin the fine-tuning mechanics where they will not be exiled to nightly

test_finetuning_pseudolabels.py carries a module-level network marker, so
anything added to it inherits the marker and runs only in the nightly
download job. The mechanics worth pinning need no download: a committed tiny
anchor works as a foundation model, so replay assembly, the frozen pt_head,
adaptation to a shifted level of theory and model-generated replay labels all
become required pull-request coverage instead of a skip that reads as a pass.

The two legacy behaviours the rewrite has already decided to drop are pinned
here so the removal is declared rather than discovered. The per-batch
pseudolabel fallback keeps the original file labels for whichever batch
failed and the stage still reports success, so a replay set can silently mix
model labels with foreign ones; the test exhibits it and records that the
only configuration reachable from the command line which makes a batch fail
also stops the run later, in the ordinary loader -- which is the measurement,
not a weakness of the test, because it shows the pseudolabel stage's failure
is invisible in its own output. The unlabelled-replay continuation is worse
than "it is allowed": the run completes with the replay head reporting a loss
of exactly zero and every error metric absent, because the generated labels
arrive carrying the file's zero property weights, so the head contributes
nothing while looking perfect.

Replay selection is driven through mace_finetuning_select, since the
--*_pt flags themselves only run against the four downloaded corpora; the
flags' own surface is pinned separately so a rename fails. Two findings fell
out: --disallow_random_padding refuses rather than returning the short set
its help implies, and the selection CLI writes its descriptor cache next to
wherever it was started, which put a stray array in the checkout until every
call got a working directory.

* Freeze what the LAMMPS export computes, and commit the input it computes it on

The plain TorchScript export format is on its way out, but what replaces it
has to produce the same physics, and nothing pinned the numbers -- the
existing contract test loads the artefact and checks it is finite.

The obstacle was a rule collision worth stating rather than quietly
resolving. Contract tests may not reach into the package for their
assertions, and building the domain-decomposed input LAMMPS hands to a pair
style needs a neighbour list, of which the package's own is the only one on
hand. Accepting the import makes the golden depend on the code it measures;
reimplementing the list commits a second one that can drift from the first
unnoticed. Neither is taken: the input is committed next to the outputs, the
import lives in the regeneration path, and the replay path imports nothing
but torch. A rewrite reproduces this golden by reading the file. A test
asserts that the import stays confined, because that property would otherwise
be lost in a commit that only meant to tidy the imports.

Four replicas per axis is a measurement, not a round number. The local
block's environment is only complete in the limit, and the residual against
the committed periodic reference is 1.2e-6 eV/Ang on the forces at three
replicas, 1.8e-7 at four and 2e-15 at five; three does not fit inside the
tolerance row and five costs 128 kB of committed edge list.

The ML-IAP format gets its declared interface snapshotted -- what a LAMMPS
input script has to agree with -- and not its numerics: past the first
interaction layer that path needs LAMMPS to exchange ghost node features, and
a stand-in would have to guess a protocol this repository has no reference
for. The refusal that produces is pinned instead, since without it the
failure was a bare AttributeError inside the second layer.

The paths filter learns about the goldens: they live with the other goldens
rather than with the test that reads them, so without this the one pull
request that moves them is the one that does not check them.

* Make the black-box rule a test instead of a paragraph

The contract suites are worth writing once only if they can be re-run against
a different engine unedited, and that survives exactly as long as nobody
reaches into the package for an assertion. It is the kind of property that
decays one convenient import at a time, each of which looks harmless and none
of which fails anything, so the failure mode is that the suite quietly stops
being portable and nobody finds out until the parity run cannot start.

The check reads imports out of the source rather than resolving them, and
refuses module names spelled as strings too, since that is the obvious way
around the first check. The ase calculator is allowlisted with its reason:
it is itself one of the contracts, there is no console script for it, and its
padding arguments exist nowhere else.

* Update the recorded defects that have since been fixed

Five of the contracts here pinned defects rather than behaviour, each saying that
a failure means the defect was fixed and the test should be updated. Four of them
have been, on develop, so this does that.

Two are deleted because the condition they recorded is gone and the repair they
named has happened:

* `--return_node_energies` on a ragged set. The test said to delete it and widen
  `test_eval_node_energies_sum_to_the_total_energy` to the ordinary fixture set,
  which is what happens here -- so the ragged case is now covered by the
  positive test rather than by a pin on the failure, and the equal-size fixture
  built to work around it goes with it.
* `--default_dtype` against a disagreeing checkpoint. The CLI converts now, as
  the calculator always did, and the contract for that lives in
  `tests/workflows/test_eval_configs_shapes_and_dtype.py` beside the change.

Three are rewritten, because the behaviour they exercise is still worth pinning
and only the outcome moved:

* The L-BFGS `drop_last` contrast stays -- it is real behaviour, not a defect --
  and now asserts that the refusal names `--batch_size` instead of matching the
  empty-tensor error from the average-neighbour statistic.
* A failing pseudolabel batch stops the stage. The assertions invert: the stage
  must report that it did not succeed, where before it had to report success.
  "continuing with original configurations" is now true, which is the point.
* An unlabelled replay set is relabelled into something that trains, so the
  `pt_head` reports real metrics where it used to report a loss of exactly zero
  and no metrics at all.

Checked against the pre-fix tree as well as this one: all four of the rewritten
or widened tests fail on `cf2f2f26`, so none of them has been loosened into
passing.

* Give the LAMMPS real tier's reference batch the model's dtype (#1658)

The real tier compares LAMMPS's energy against a python evaluation of the same
model, and that comparison had never once run. The test died earlier, inside
LAMMPS, on the KOKKOS-only writeback. With that fixed the assertion is finally
reached, and it fails: model_batch builds its AtomicData at the process-wide
default dtype, which is float32 under pytest, while the model is float64.

Nothing in the failure says so. It surfaces from inside the first scripted e3nn
graph, the node embedding's o3.Linear, as "RuntimeError: both inputs should
have same dtype", naming neither the batch nor which side is wrong.

test_export.py and test_ghost_parity.py each dodge this by setting the default
dtype to float64 before they load a model. Casting the batch to the model's own
dtype instead, read off r_max exactly as LAMMPS_MLIAP_MACE.__init__ does, keeps
the next caller of the harness from inheriting the trap.

* Scale a committee's spread by the square of the unit conversion (#1656)

`energy_units_to_eV` and `length_units_to_A` multiply every energy, force and
stress the calculator returns, and nothing exercised them. Closing that gap
turned up a bug in the surface next to it.

A variance carries the square of whatever scales the values:

    var(c * X) == c**2 * var(X)

`MACECalculator` applied the factor once, and to data it had already converted:
the ensemble array was scaled in place, `.numpy()` shares storage with a cpu
tensor, so the variance was then taken over converted values and multiplied
again. Measured on a two-member committee, `energy_var` came out at
`unit_conv**3` -- 8x rather than 4x for a factor of two, and 2x the variance of
the `energy_comm` array it claims to describe. `MagneticMACECalculator` has the
exponent half of the same mistake, without the aliasing.

None of it shows at the default factor of 1.0, where a number, its square and
its cube agree, which is why the existing committee test could assert
`energy_var == var(energy_comm)` and pass throughout. It is `energy_var` that an
active-learning loop reads to choose what to label next, so a spread wrong by a
constant selects badly and reports nothing.

The tests state it as an internal consistency -- the reported variance is the
variance of the reported members, whatever the conversion -- and once from
outside, that converting multiplies the members linearly and the variance
quadratically. The conversions themselves turn out to have been right all along:
those five cases pass on the unfixed tree, and only the two committee ones fail.

* Restore the process default dtype after changing it (#1659)

* Put the process default dtype back after changing it

`default_dtype` is a context manager over a process-wide setting, and it only
restored on the way out of a clean block. An exception inside left the whole
process on the scope's dtype, silently, because nothing reads it back: the next
tensor built without an explicit dtype simply came out at the wrong precision.

`MACECalculator` runs inside this scope in nine places, `_atoms_to_batch` among
them, so an unknown element in a structure was enough to leave a float32 caller
at float64. Under pytest the same thing contaminates every later test on the
worker, which is how a test passes alone and fails in a suite.

Adds `restores_default_dtype` for the callers that need the same guarantee but
set the default themselves rather than wrapping a block.

* Stop the backend converters changing the caller's default dtype

Each converter reads the dtype off the source model and sets it process-wide so
the submodules it builds inherit it, then never puts it back, on success or on
the way out of an exception.

These are library functions, not only CLI entry points. `MACECalculator(
enable_cueq=True)` calls `run_e3nn_to_cueq` at construction, `run_train`
converts mid-run at a dtype it has already chosen from `--default_dtype`, and
`eval_configs` and `create_lammps_model` call it too. In each case the caller
had picked a dtype and got the model's instead.

The decorator rather than a `with` around each body: the body still needs the
default set while it constructs, so the only thing that changes is that the
scope ends.

* Take the ML-IAP wrapper's buffer dtype from the model

`total_charge` and `total_spin` were built at `torch.get_default_dtype()`, and
the export reached the right dtype only as a side effect: `--format mliap`
converts to the cueq layout first, and that converter used to set the default
globally on its way through. With the converter restoring it, a float64 export
would get float32 buffers.

Nothing would have raised. They are handed to the forward beside the model's own
tensors, so promotion turns the mismatch into a quietly less accurate number,
and it is LAMMPS that consumes it. `r_max` is where `LAMMPS_MLIAP_MACE` reads
its own dtype, so the wrapper and the coupling cannot disagree.

* Add backend parity goldens for the cueq and oeq kernels (#1657)

* Pin the accelerated kernels against the committed CPU numbers, and prove one ran

The backend suite only checks cueq and oeq against an e3nn model built beside
them in the same process, which moves whenever the e3nn side moves. Nothing
held the accelerated path to a number somebody had to rewrite a file to
change, so this converts the ScaleShiftMACE anchor (and MACE-MP-0 small where
a job may download) on the accelerator and asserts the committed CPU e3nn
references at the shared accelerated tolerance row. No reference is added and
none is regenerated.

The hard part is that matching those references proves nothing about kernels:
they came from the unaccelerated path, so the unaccelerated path reproduces
them exactly. Two ways to get there silently, and both are real. Converting
with the default device leaves conv fusion off, because the converter ties it
to device == "cuda" -- the resulting model is genuinely cuEquivariance, runs
the unfused channelwise product, and reproduces the anchor reference to
1.7e-16, eleven orders of magnitude inside the row, so no tolerance can catch
it. And cuEquivariance downgrades to its naive implementation with a warning
and no error when the ops wheel will not import, which is the state of any
plain [cueq] install. A value-only golden is green in both cases while no
vendor kernel has run, which is the exact oracle failure this is meant to
prevent.

So every case audits the module tree for the accelerated implementation and
counts the calls into it with forward hooks, and the audit does not lean on
the refusal MACE already has in its fusion wrapper, since that is code under
test. Both halves are exercised on a host with no GPU: the unfused conversion
is reproduced end to end, and the naive downgrade is the real thing rather
than a stand-in wherever cuequivariance_ops_torch is absent, which is what
the CPU backend job has. The plain-MACE anchor stays out of parity because
the conversion whitelist refuses it; that both converters stop on the refusal
instead of returning a model is pinned too.

Markers keep the per-vendor selection intact: the network case is cueq-only,
because an oeq one would be collected by the AMD job, whose required
capabilities promise no network test it can reach. The extension workflow's
backend filter learns tests/golden, or no job with cuequivariance installed
would review a change to any of this.

* Drop the dtype scope the converters no longer need

`_convert` wrapped the conversion in `default_dtype` because `run` set the
process default from the source model and never put it back. That is fixed in
the converters themselves now (#1659), so the scope guarded nothing.

It could not have worked as described either: `run` sets the default from the
source model's parameters inside its own body, which overrides an outer scope,
so the accelerated modules were never reading the value this set. Removing it
rather than keeping it as insurance, because the docstring explaining it is now
wrong, and a wrong explanation of a no-op is worse than neither.

* Add golden references for the dipole, dielectric and polar models (#1660)

* Add golden references for the dipole, dielectric and polar models

Nothing in the tree fixed a non-energy observable numerically: the golden
anchors covered energy models only, and the electrostatics conversion work
was about to be measured against nothing. Three references close that. A tiny
seeded AtomicDipolesMACE pins the dipole assembly per pull request; the
published MACE-Polar model pins the energy surface together with the
electrostatics keys only that family has; MACE-MDP pins the polarizability,
which AtomicDielectricMACE is the only class in the tree to emit -- PolarMACE
never does, so a polarizability test aimed at the polar model would assert
nothing.

The calculator door turns out not to open for these models at all. A
dipole-only MACECalculator leaves no energy in its results, so the harness's
get_potential_energy entry point raises before any snapshot happens, and
dmu_dr, dalpha_dr and atomic_dipoles reach no results dict in any case. Two
adapters in tests/golden/routes.py drive calculate() directly and present a
forward as an evaluation; they forward attribute lookups to the wrapped
calculator because the harness recovers what an evaluation reads by
inspecting it, and an adapter that only passed outputs through would record
the inputs of a reader that never ran. The MDP calculator path is then pinned
against the same file the forward is, so one number answers for both doors.

Two divergences surfaced while measuring and are pinned rather than tidied.
The two fixed-charge baselines are not in the same unit -- AtomicDipolesMACE
divides by 1e-11/c/e and AtomicDielectricMACE has that division commented out
-- so the same channel carries Debye from one family and e*Ang from the
other, and a rewrite that unifies the helpers now fails on the factor instead
of silently rescaling a committed reference. And MACECalculator declares no
implemented_properties of its own, so every instance extends ase's
class-level list in place: the first version of the polar test asserted
membership there, passed alone, and failed as soon as the two golden files
ran in one process. The assertion counts what a construction adds instead,
which stays exact whether or not the leak is fixed.

The ci-extensions polar job gains tests/golden and network in its guaranteed
capabilities, because it is the only pull-request job with graph_longrange
and a skip there would read as a pass; the nightly foundations job gains the
same selection for the download-only MDP golden.

* Compare bit for bit within a process, not against a committed file

Two assertions in these goldens demanded bit-equality against numbers recorded
on another machine, which is not a property either of them was testing. Both
failed in CI on torch 2.13 while passing locally on 2.11, at 1e-16.

The calculator route was held bit-identical to the committed reference. The
claim it was making is that the calculator is the same forward with a dict
rename in front of it, so it now evaluates the forward route here too and
compares the two in this process. Drift is still caught against the file, at
the tolerance row, so both halves of the original intent survive: two routes
that drifted together would still fail the file comparison.

The anchor's Clebsch-Gordan buffers were compared with torch.equal. e3nn builds
that basis through torch linalg, so the two sides differ in the order of the
fp64 operations, and the extent is what distinguishes the reduced basis this
test exists to detect. The shape stays exact and the entries move to the
closed_form_fp64 row, which is the row for a formula evaluated twice. The
failure message prints the actual deviation, since a bound is only worth having
if a breach says by how much.

* Point the dipole anchor's sidecar at the target that rebuilds it

The dipole anchor and its reference are built by the dipoles target, not by
anchors, which runs the two energy anchors only. Following the recorded
command therefore rebuilds tiny_scaleshift and tiny_mace and leaves
tiny_dipoles.model and its reference exactly as they were, which is the worst
shape a stale golden can take: the regeneration reports success and the file
the reader meant to refresh never moves.

The sidecar is edited in place rather than regenerated, so the checkpoint
bytes and the reference stay the ones the tests were measured against.

The assertion that was meant to catch this only required the string to
mention regenerate.py, which every target satisfies. It names the family now.

* Enforce per-file coverage floors and add CPU benchmark cases to the nightly (#1661)

* Add CPU cases to the nightly benchmark job

The job has been publishing an empty baseline for its whole existence. The
only test in tests/benchmarks is marked gpu and network, so on the CPU runner
all sixteen parametrizations skipped, the job went green, and the uploaded
benchmark.json was a 0-byte file. "The rewrite is not slower" is only
falsifiable against an old number, and there was no old number.

The CPU cases fill it: the committed tiny anchor at three sizes, MP-small at
two, both dtypes, with dtype, device, torch version and backend recorded
beside every timing so the comparison still means something after the stack
that produced it is gone. The size spread is the point rather than the count.
Measured on the anchor in fp32, throughput is 6897, 8588 and 10658 atom-steps
per second at 216, 512 and 1728 atoms, so the smallest case runs at 65% of the
largest one's rate and about a third of its per-atom cost is work that does
not scale with the system. That is the overhead a faster-per-kernel rewrite
can still lose to, which is why the small sizes are there. The curve has not
flattened by 1728 atoms and the docstring says so rather than calling it
saturated. The usual framing of the small regime as a few-millisecond step is
a GPU statement and is not the wall-clock on a CPU runner.

Dropping continue-on-error is not by itself enough, and neither is
if-no-files-found on the upload: pytest-benchmark creates the --benchmark-json
path whatever happens, so the empt…

This branch was previously deployed

1 inactive deployment
gpu-external — 56e58f12 Deployed Aug 17, 2026 by darthjaja6 via gpu-amd #127
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants