Skip to content

Add AVX2-baseline llama.cpp gfx1201 ROCm image + platform-aware build runner - #12

Merged
Aaronontheweb merged 1 commit into
masterfrom
feature/amd-rocm-llama-cpp-avx2
Aug 14, 2026
Merged

Aaronontheweb merged 1 commit into
masterfrom
feature/amd-rocm-llama-cpp-avx2

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Summary

Adds the first AMD ROCm image family member and the CI plumbing to build it natively.

Image llama-cpp-rocm-gfx1201-avx2 (images/amd/rocm/llama-cpp-gfx1201-avx2/): an AVX2-baseline llama.cpp llama-server, built from pinned upstream source (commit 9b05354…, build b10433) against a digest-pinned rocm/dev-ubuntu-24.04:7.2.3 base, targeting the gfx1201 (RDNA4) GPU backend.

Why

Stock upstream llama.cpp ROCm toolbox images ship a libggml-cpu compiled with an AVX-512 CPU baseline. On an inference host whose CPU has AVX2 but not AVX-512, ggml_cpu_init() executes an AVX-512 instruction (vmovdqa32 %zmm0) at process startup and llama-server dies with an illegal instruction (SIGILL) before it loads a model or touches the GPU. This image compiles the CPU backend at an explicit x86-64-v3 (AVX / AVX2 / FMA / F16C) baseline with all AVX-512 paths disabled, and a build-time guard fails the image if any AVX-512 (zmm) opcode remains in the CPU backend. The GPU backend still targets gfx1201, and the ROCm 7.2.3 base intentionally matches the proven-good production runtime.

Pipeline change

build-candidate.yml's build job was hardcoded to runs-on: ubuntu-24.04-arm. The validation layer already supports amd/rocm/linux/amd64, but an amd64 image on an ARM64 runner would build under QEMU — a non-starter for a ROCm/llama.cpp compile. The build job now selects its runner from the manifest's platform (linux/amd64 → x64, linux/arm64 → ARM64), so this image builds natively. docs/ci-cd.md updated to match.

Supply chain

  • Base image pinned by sha256 digest; llama.cpp pinned to an immutable commit with its GitHub archive verified by SHA-256.
  • Everything in the runtime is compiled from that pinned source — no inherited prebuilt inference binaries.
  • llama.cpp's MIT license is retained in the image; attribution/ and dependency.lock.json record the full boundary.

Status / test plan

Source-build candidate; hardware qualification pending. build_enabled: true so a candidate can be dispatched.

  • ./scripts/validate-repository.sh passes (2 manifests)
  • Release-note change policy passes
  • Next: dispatch Build inference image candidate for images/amd/rocm/llama-cpp-gfx1201-avx2/image.json on master, then qualify the resulting digest on gfx1201 hardware before promoting to a release tag.

The first candidate compile may need one or two iterations on ROCm/CMake specifics; that's the intended build-in-CI loop, not a regression.

…aware build runner

Adds the first AMD ROCm image family member: an AVX2-baseline llama.cpp
llama-server built from a pinned upstream commit (b10433) against a
digest-pinned ROCm 7.2.3 base, targeting the gfx1201 (RDNA4) GPU backend.

Stock upstream llama.cpp ROCm toolbox images ship a libggml-cpu compiled
with an AVX-512 CPU baseline; on AVX2-only inference hosts ggml_cpu_init()
executes an AVX-512 instruction at startup and llama-server dies with SIGILL
before loading a model. This image compiles the CPU backend at an explicit
x86-64-v3 (AVX2) baseline with AVX-512 disabled, and fails the build if any
AVX-512 opcode remains in the CPU backend.

Also makes the candidate build runner platform-aware so linux/amd64 images
build natively on x64 runners instead of under emulation on the ARM64 builder.
@Aaronontheweb
Aaronontheweb merged commit 111c1ae into master Aug 14, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant