Repository navigation
Add AVX2-baseline llama.cpp gfx1201 ROCm image + platform-aware build runner - #12
Merged
Merged
Conversation
…aware build runner Adds the first AMD ROCm image family member: an AVX2-baseline llama.cpp llama-server built from a pinned upstream commit (b10433) against a digest-pinned ROCm 7.2.3 base, targeting the gfx1201 (RDNA4) GPU backend. Stock upstream llama.cpp ROCm toolbox images ship a libggml-cpu compiled with an AVX-512 CPU baseline; on AVX2-only inference hosts ggml_cpu_init() executes an AVX-512 instruction at startup and llama-server dies with SIGILL before loading a model. This image compiles the CPU backend at an explicit x86-64-v3 (AVX2) baseline with AVX-512 disabled, and fails the build if any AVX-512 opcode remains in the CPU backend. Also makes the candidate build runner platform-aware so linux/amd64 images build natively on x64 runners instead of under emulation on the ARM64 builder.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the first AMD ROCm image family member and the CI plumbing to build it natively.
Image
llama-cpp-rocm-gfx1201-avx2(images/amd/rocm/llama-cpp-gfx1201-avx2/): an AVX2-baseline llama.cppllama-server, built from pinned upstream source (commit9b05354…, buildb10433) against a digest-pinnedrocm/dev-ubuntu-24.04:7.2.3base, targeting thegfx1201(RDNA4) GPU backend.Why
Stock upstream llama.cpp ROCm toolbox images ship a
libggml-cpucompiled with an AVX-512 CPU baseline. On an inference host whose CPU has AVX2 but not AVX-512,ggml_cpu_init()executes an AVX-512 instruction (vmovdqa32 %zmm0) at process startup andllama-serverdies with an illegal instruction (SIGILL) before it loads a model or touches the GPU. This image compiles the CPU backend at an explicitx86-64-v3(AVX / AVX2 / FMA / F16C) baseline with all AVX-512 paths disabled, and a build-time guard fails the image if any AVX-512 (zmm) opcode remains in the CPU backend. The GPU backend still targetsgfx1201, and the ROCm 7.2.3 base intentionally matches the proven-good production runtime.Pipeline change
build-candidate.yml's build job was hardcoded toruns-on: ubuntu-24.04-arm. The validation layer already supportsamd/rocm/linux/amd64, but an amd64 image on an ARM64 runner would build under QEMU — a non-starter for a ROCm/llama.cpp compile. The build job now selects its runner from the manifest'splatform(linux/amd64→ x64,linux/arm64→ ARM64), so this image builds natively.docs/ci-cd.mdupdated to match.Supply chain
sha256digest; llama.cpp pinned to an immutable commit with its GitHub archive verified by SHA-256.attribution/anddependency.lock.jsonrecord the full boundary.Status / test plan
Source-build candidate; hardware qualification pending.
build_enabled: trueso a candidate can be dispatched../scripts/validate-repository.shpasses (2 manifests)images/amd/rocm/llama-cpp-gfx1201-avx2/image.jsononmaster, then qualify the resulting digest ongfx1201hardware before promoting to a release tag.The first candidate compile may need one or two iterations on ROCm/CMake specifics; that's the intended build-in-CI loop, not a regression.