Skip to content

Make CI finish in about a third of the time - #7

Merged
fidetolabs merged 2 commits into
mainfrom
faster-ci
Sep 21, 2026
Merged

fidetolabs merged 2 commits into
mainfrom
faster-ci

Conversation

@fidetolabs

@fidetolabs fidetolabs commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Four of the six minutes a release takes were one pytest step, and nearly all of
that was test_backtest.py running 36 independent replays one after another.

Two changes, one commit each.

Run the suite across cores. pytest-xdist, with -n auto --dist loadgroup in
both workflows. The catch is tests/test_postgres.py: it shares one database and
resets by emptying public, so two of its tests on different workers drop each
other's tables. Plain --dist load gives four errors on one run and four failures
on the next, moving around between runs. The file carries an xdist_group mark
now, which keeps it on a single worker while everything else still spreads a test
at a time.

Replay over less history. A replay re-runs the whole pipeline once per
rebalance date, so what that file costs is stops, not rows -- and _window takes
half of the scaffold's 420 days, about 39 stops, thirty-six times over. 150 days
leaves around a dozen, which every assertion still has room for. Two tests write
their windows as dates rather than deriving them from the feed, and a short feed
ends after they start; they ask for days=420.

Measured

On a runner, test (3.12), the pytest step:

time
main, before 264s
this branch 72s

3.7x. The two changes compound -- the shorter feed is worth about 2x whatever the
core count, and xdist takes what is left. the Postgres tests must have actually run still passes, so the group did not turn them into skips.

On release.yml

That file runs only on a tag, so nothing in it can be proven by this PR. Its only
change here is the same pytest flag test.yml just ran green -- deliberately.
Folding announce into publish would have saved another ~45s of runner boot,
but it would have put an untested step after the PyPI upload, and PyPI will not
take the same version twice: publish succeeds, release page fails, and the tag
cannot be retried. Not worth 45 seconds. Left as two jobs.

https://claude.ai/code/session_01GN79GJbvpEiwuRB23ZDkXt

Four minutes of a six-minute release were one pytest step, and nearly all of that
was test_backtest.py running 36 independent replays one after another.

`-n auto` spreads them. The catch is tests/test_postgres.py: it shares one
database and resets by emptying `public`, so two of its tests on different
workers drop each other's tables -- `--dist load` gives four errors on one run
and four failures on the next, moving around. The file carries an `xdist_group`
mark now and `--dist loadgroup` keeps it on one worker while everything else
still spreads a test at a time.

Locally, with a server up: 81s serial, 27s across ten cores, 198 passed either
way and three times running.
A replay re-runs the whole pipeline once per rebalance date, so what this file
costs is stops, not rows -- and `_window` takes half of the scaffold's 420 days,
which at `rebalance: 5d` is about 39 stops, thirty-six times over.

150 days leaves around a dozen stops, which every assertion here still has room
for. The file goes from 68s to 29s on its own.

Two tests write their windows as dates rather than deriving them from the feed,
and a short feed ends after they start. They ask for `days=420`, which is what
the parameter is for.
@fidetolabs
fidetolabs merged commit b71797c into main Sep 21, 2026
3 checks passed
@fidetolabs
fidetolabs deleted the faster-ci branch September 21, 2026 12:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant