Skip to content

feat(randomization): add mean_matched_uniform distribution kind - #203

Open
tkevinbest wants to merge 2 commits into
mainfrom
feat/mean-matched-uniform-distribution
Open

tkevinbest wants to merge 2 commits into
mainfrom
feat/mean-matched-uniform-distribution

Conversation

@tkevinbest

@tkevinbest tkevinbest commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Strictly opt-in: no shipped preset uses the new kind, so every existing config draws bit-identical values.

Two mass randomization presets use plain uniform ranges. link_mass_range: [0.9, 1.2] is a multiplicative scale on each randomized link, unperturbed at 1.0. added_mass_range: [-1.0, 3.0] is an additive kg offset on the torso, unperturbed at 0.0. A uniform's mean is the midpoint of its band, and neither band is centered on its unperturbed value: [0.9, 1.2] averages 1.05, [-1.0, 3.0] averages +1.0 kg. Averaged over environments, the policy trains against a robot whose links are 5% heavier than the URDF describes and whose torso carries an extra kilogram it does not have. A config diff shows none of this; the ranges look reasonable. The asymmetry is deliberate — the wide upper side covers carrying a payload — so narrowing or re-centering the bounds is not the fix.

mean_matched_uniform keeps the same bounds and puts the mean where it belongs. For low < mean < high it draws U[low, mean] with probability (high - mean) / (high - low) and U[mean, high] otherwise, so E[X] = mean exactly while the support stays [low, high].

"added_mass_range": {"kind": "mean_matched_uniform", "low": -1.0, "high": 3.0, "mean": 0.0},

No existing kind does this on an asymmetric band: on [0.9, 1.2], uniform gives 1.05, log_uniform 1.0428, and gaussian with mean=1.0, std=0.05 gives 1.0028, because that parameter is the pre-truncation mean. DistributionSpec.expectation() returns the closed-form mean of any spec, so a range's mean can be checked rather than assumed.

What holds the default behavior fixed:

  • Nothing outside distribution.py, sampler.py and their tests references the new kind.
  • A bare [lo, hi] pair still parses as uniform.
  • The uniform, log_uniform and gaussian branches of _inverse_cdf are untouched. The only edit to shared code moves _INV_SQRT2 and _P_EPS from sampler.py into distribution.py at identical values, so expectation() can share them.
  • The sampler is keyed, so switching a preset to the new kind would change drawn values at a fixed seed. That is a separate decision for consumers.

mean is required rather than defaulted to the midpoint, and validation requires strict low < mean < high. A mean at the midpoint routes to the plain-uniform path so the two agree bit-for-bit. The kind string is serialized into checkpoints and DistributionSpec.parse raises on unknown kinds, so renaming it later breaks eval of checkpoints trained with it.

98 unit tests (28 new), no_sim CI selection 551 passed / 9 skipped, pre-commit clean, IsaacSim DR matrix passing. IsaacGym and MuJoCo were not run and the kind has no cross-backend _dr_matrix.py cases, both because _inverse_cdf is the repo's only spec.kind dispatch.

A bare [lo, hi] range is uniform, so its mean is the band's midpoint. On a
band that is asymmetric about the nominal value that biases every env: a
link mass scale of [0.9, 1.2] averages 1.05, making every randomized link
5% heavier than its URDF, and an added-mass range of [-1.0, 3.0] puts
+1.0 kg on every torso. The asymmetry is deliberate (the wide side covers
payload), so rather than change the bounds this adds a distribution whose
shape spans the same band with its expectation on the nominal.

mean_matched_uniform is a two-piece uniform: with lo < m < hi, it draws
U[lo, m] with probability p = (hi - m) / (hi - lo) and U[m, hi] otherwise,
which is the lever rule and makes E[X] = m exactly. Its inverse CDF is
piecewise linear with a kink at (p, m), so it drops into the existing
inverse-CDF sampler; _inverse_cdf is the single spec.kind dispatch point
in the repo, so no backend code changes.

Also adds DistributionSpec.expectation(), the exact analytic mean for every
kind, so a range's bias is computable rather than eyeballed. No existing
kind can span an asymmetric band unbiased: on [0.9, 1.2] uniform gives
1.05, log_uniform 1.0428, and a gaussian with an explicit mean=1.0 still
lands at 1.0028, since that mean is the pre-truncation one.

Mechanism only -- no shipped preset changes, since flipping one alters
training dynamics at a fixed seed and is a separate decision.
The docstring edits in randomization/terms/locomotion.py and objects.py were
rewrites of existing prose in files this branch leaves functionally alone, so
they are reverted. What remains is limited to the two modules that implement
the new kind, plus its tests and a README note.

Also cuts the commentary that argued for the change rather than describing the
code: coverage percentages, restatements of adjacent error messages, and
docstring prose duplicated across surfaces.

Net effect on the diff: 6 files -> 4, and 18 pre-existing lines touched -> 9
(the Distribution Literal, the Fields docstring block, and the two constants
that had to move).
@tkevinbest
tkevinbest marked this pull request as ready for review September 14, 2026 16:20
@tkevinbest
tkevinbest enabled auto-merge (squash) September 14, 2026 16:21

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant