Repository navigation
Make ONNX Runtime optional and run CUDA natively with cuda-oxide kernels - #36
Conversation
PLDA setup needs two inverses and one generalized symmetric eigensolve on 128x128 matrices, once per pipeline. That alone pulled in ndarray-linalg with statically linked MKL or OpenBLAS and four backend features. A small Cholesky, pivoted LU, and tred2/tql2 solver now does this with no dependencies. Results match LAPACK to about 1e-14 on phi and 2e-12 on the transform, and diarization output is unchanged. Breaking: the default-linalg, intel-mkl, openblas-static, and openblas-system features are removed.
macOS builds only need CoreML, yet every build compiled, linked, and downloaded ONNX Runtime, and CoreML modes still built ORT sessions at load time. ONNX Runtime is now an optional dependency enabled only by the cpu, cuda, migraphx, and load-dynamic features. Each build must pick a backend explicitly; a build without one fails with a compile error that lists the choices. Each model holds a single backend chosen at load, so CoreML modes never create ORT sessions or download ONNX files. A crate-owned InferenceError replaces ort::Error in the API. CoreML and CPU diarization output is byte-identical to before on VoxConverse-dev samples. CoreML startup roughly halves and peak memory drops by about 90 MB. Breaking: no default backend, ExecutionMode::Cpu needs the cpu feature, and ort::Error is replaced in public signatures.
The native CUDA backend loads safetensors weights instead of ONNX models, and its kernels need layer-by-layer references to check parity against ONNX Runtime. export_weights.py converts the segmentation and multi-mask embedding ONNX models into the safetensors files the runtime loads, along with a manifest of each graph. make_reference.py records ONNX Runtime CPU inputs, outputs, and every intermediate tensor for the fixture audio. Both are deterministic and download only public models.
ONNX Runtime's CUDA provider ran the filterbank DFT on the CPU, copied data back and forth every batch, needed over 15 GB of GPU memory in one process, and tied users to a matching ORT, CUDA, and cuDNN install. CUDA modes now run natively through cudarc: cuBLAS and cuDNN for the heavy layers and cuda-oxide kernels, shipped as committed PTX, for the filterbank, pooling, SincNet front end, and epilogues. The cuda feature no longer pulls in ONNX Runtime and builds without a CUDA toolkit. Models load from safetensors weights. On all 216 VoxConverse-dev files, DER matches the ONNX Runtime path and diarization runs 1.81x faster in cuda mode and 1.62x faster in cuda-fast mode. Defaults: FP32 segmentation (TF32 worsened one file by 4.5 DER points), TF32 embedding, the persistent SmallH LSTM, and CUDA graphs. Kernels target sm_75; cuda-sm80, cuda-sm90, and cuda-sm120 opt in to newer GPU tiers. Breaking: CUDA modes need the new safetensors assets, a Turing or newer GPU, and cuDNN 9 at run time.
The CUDA benchmark binary no longer uses ONNX Runtime's CUDA provider, so the images stop downloading and copying the ORT GPU shared libraries. The native backend loads cuBLAS, cuDNN 9, and NVRTC from the CUDA runtime base image. These Dockerfile changes have not been built yet.
The gpuq canary staged models and datasets from a Tigris bucket and uploaded results there, so it could not run without that bucket and its AWS credentials. speakrs-bm now downloads the model assets for the selected CUDA modes anonymously from Hugging Face at the revision pinned in src/models.rs. Datasets come from their public sources through the existing xtask acquisition. Results stay on the worker and a compact summary is printed between SPEAKRS_RESULTS markers in the captured log. The workload needs no secrets. The canary's --impls speakrs now selects the native CUDA implementation; it previously matched no implementation.
The ONNX Runtime CUDA EP enabled TF32 and NHWC, not tf32=false. Note that CUDA modes now run on the native backend with per-model math.
Disable default CPU features in both CUDA build stages, since the default xtask feature enables ONNX Runtime. Check the GPU binary dependency tree in CI so ORT cannot return unnoticed.
Use exact power-of-two scaling outside the safe f64 range and restore eigenvalues and B-normalized vectors. Return typed errors when a scale would lose nonzero values. Keep normal-scale arithmetic unchanged. Cover extreme and generalized problems with relative residuals, and pin the shipped scale factors. Refresh only the linalg source hash in the qualification lock.
Show the CUDA PTX environment name as inline code and keep ORT-only items out of the CUDA feature labels. Resolve the VBx variant links so the strict docs build passes. Refresh only the four changed doc-file hashes in the qualification lock. Keep all kernel and qualification harness bytes unchanged.
|
Important Review skippedToo many files! This PR contains 190 files, which is 90 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. ⚙️ Run configuration
⛔ Files ignored due to path filters (6)
📒 Files selected for processing (190)
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
The host-only PLDA tests read the PLDA model files, which this job did not fetch.
Keep device limits distinct from real planning errors. Allow Library fallback in production, but fail explicit qualification selections.
Bind production PTX, source files and coverage to accepted records. Encode the Library-noise rule and reject incomplete evidence.
Require the source-evidence note on each manifest entry, correct the Scratch and SideStream signatures in the harness docs, and keep CI push runs on master.
Summary
This PR makes ONNX Runtime optional, so the CUDA modes no longer use it. They run on a native NVIDIA backend instead, sped up by three custom cuda-oxide kernels. Overall, the CUDA modes run about 2.96× (
cuda) and 2.69× (cuda-fast) faster thanmaster. CPU and CoreML output stays byte-identical tomaster, and DER does not change.What changed
cpu,migraphxandload-dynamicfeatures.coremlandcudadon't use it. Building with no backend feature is a compile error, so a backend has to be chosen explicitly.ExecutionMode::CudaandCudaFastrun on cudarc with cuBLAS and cuDNN 9, loaded at run time. Models load from safetensors weights, and CUDA graphs are used.RuntimeConfig. Segmentation runs in FP32 by default, and embedding in TF32.cuda-sm80,cuda-sm90,cuda-sm120) are opt-in features and off by default. cuDNN still handles every other layer, batch and precision combination.PersistStaticSmallHcargo xtask cuda-qualify,scripts/cuda/qualify/). Production selects a kernel only for the layer, batch and precision combinations the harness qualified, because it reads the same coverage declarations.Why
Dropping ONNX Runtime from CUDA and CoreML builds removes a large native dependency and its download step. It also lets the native backend run faster than ORT did, while keeping DER unchanged.
Results
VoxConverse-dev, all 216 files, run A/B/A/B:
cudacuda-fastmaster's ORT CUDA, matched per filePersistStaticSmallH, CUDA graphs), full-set RTFxmastermastermasterexceeded the GPU box's memory limit in the one-process full-set run, so its full-set RTFx could not be measured.master, the default precision choices were each gated per file, and the kernels change DER by +0.0003 points incudaand 0 incuda-fast.master(280 of 280 compared) and so is CoreML (120 of 120). CoreML also starts in about half the time and uses about 90 MB less memory.Validation
just fmt,just clippy,cargo test --workspace, and a warnings-as-errors docs build.cargo checkfor every supported backend combination with--no-default-features. The intended compile errors fire for invalid combinations.Known limits