Skip to content

ray-haproxy: Add version 2.8.25 - #2428

Draft
luhenry wants to merge 11 commits into
mainfrom
ray-haproxy
Draft

luhenry wants to merge 11 commits into
mainfrom
ray-haproxy

Conversation

@luhenry

@luhenry luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Bundles a prebuilt HAProxy 2.8.25 binary for Ray Serve, vendoring OpenSSL and Lua built from source plus PCRE/libxcrypt found via ldd. Upstream ships no riscv64 wheel.

Mirrors upstream's release.yml, run against manylinux_2_39_riscv64 in place of manylinux2014, which has no riscv64 image.

Differs from upstream

  • Builds and packages in one docker step - no separate host-side Python stage needed
  • Vendors PCRE2, not PCRE1 - Rocky 10 dropped the legacy pcre-devel package
  • Drops the CVE scan and 6-distro verify job for upstream's verify-vendoring.sh, run natively on riscv64

Matrix: 3.12/3.13/3.14 - no C ABI, so no free-threaded leg needed

Testing

  • same as upstream's smoke test, plus its verify-vendoring.sh run natively on riscv64

License: Wheel bundles HAProxy (GPLv2) plus vendored OpenSSL (Apache-2.0), Lua (MIT), PCRE2 (BSD) and libxcrypt (LGPL-2.1); license texts ship unchanged from upstream.

Patches

  • 0001-build-haproxy-dist-add-riscv64-and-build-against-PC.patch - Inappropriate [no legacy PCRE1 on manylinux_2_39_riscv64]. Without it the build exits at the arch check; riscv64-only.
  • 0002-verify-vendoring-recognise-riscv64-binaries.patch - Inappropriate [only riscv64 hits this check]. Without it vendoring verification always fails.
  • 0003-THIRD_PARTY_LICENSES-note-PCRE2-on-riscv64.patch - Inappropriate [only riscv64 vendors PCRE2]. Not a build break.

CI run in progress.

Ports HAProxy 2.8.25 for riscv64: ray-haproxy bundles a prebuilt HAProxy
binary (statically vendoring OpenSSL and Lua it builds from source, plus
PCRE and libxcrypt discovered via ldd) for Ray Serve. The build mirrors
ray-project/ray-haproxy's own release.yml, run against
manylinux_2_39_riscv64 instead of manylinux2014 (no riscv64 image exists
for that profile). Patches add a riscv64 arch branch and switch to PCRE2,
since manylinux_2_39_riscv64 (Rocky 10) dropped the legacy PCRE1 package.
luhenry added a commit that referenced this pull request Sep 28, 2026
…nary, public source repo found, builds from source for riscv64)
CI failed with OpenSSL's Configure script unable to find FindBin.pm:

    Can't locate FindBin.pm in @inc (you may need to install the FindBin module) ... at .../openssl-3.0.15/Configure line 15.

Rocky 10 (manylinux_2_39_riscv64) splits FindBin.pm out of core Perl into
perl-FindBin; manylinux2014's older Perl carries it without a separate
package, so upstream never needed it. Install it alongside perl-IPC-Cmd
in the riscv64 branch.
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://riseproject-dev.github.io/python-wheels/pr-preview/pr-2428/

Built to branch gh-pages at 2026-09-29 04:41 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

CI got past the FindBin fix (44611db) and now fails on the next
`use` line in OpenSSL's Configure script:

    Can't locate lib.pm in @inc (you may need to install the lib module) ... at .../openssl-3.0.15/Configure line 16.

Rocky 10 (manylinux_2_39_riscv64) splits the `lib` pragma out of core
Perl into perl-lib the same way it split FindBin.pm into perl-FindBin;
manylinux2014's older Perl carries it without a separate package.
Install it alongside perl-IPC-Cmd/perl-FindBin in the riscv64 branch.
build-haproxy-dist.sh's own post-patchelf sanity echo re-runs ldd on the
haproxy binary and fails with "not a dynamic executable" on
manylinux_2_39_riscv64, even though patchelf's own --print-rpath reads the
same file back cleanly one line above and the plain execve further down
(haproxy -v) runs it directly. Replace ldd's self-exec check with a static
readelf NEEDED-vs-vendored check on riscv64, leaving x86_64/aarch64
untouched.
The staged haproxy binary segfaulted with no output right when the build
script exec'd it, immediately after vendoring reported success. The
vendored libssl.so.3/libcrypto.so.3 turned out to be Rocky 10's system
OpenSSL (rpm reports 3.5.5), not the OpenSSL 3.0.15 this same script
builds from source and links HAProxy against a few lines earlier: ldd's
default resolution (no LD_LIBRARY_PATH/RPATH set yet at that point) finds
Rocky 10's own libssl.so.3 on manylinux_2_39_riscv64, something
manylinux2014 (x86_64/aarch64) never ships, so this only bites riscv64.
Point ldd at our own build before vendoring runs so it collects the
matching OpenSSL instead of Rocky 10's unrelated one.
… crash

Two inference-only fixes (ldd self-exec skip, vendor our own OpenSSL)
have not resolved the segfault at the end of build-haproxy-dist.sh on
riscv64. Add a temporary, riscv64-only diagnostic that runs the crashing
invocation under strace -f, and attempts a gdb backtrace against any core
dump produced, to get a real crash signal from the actual CI runner
instead of guessing further.

This commit is diagnostic-only and is expected to be superseded by the
real fix once the backtrace is in hand.
…limit -c

The diagnostic added in 288a119 called `ulimit -c unlimited`
unconditionally before running the crashing `haproxy -v` under strace.
On the real riscv64 runner the container's core-dump hard limit is
locked to 0, so that call itself failed:

    ./ci/build/build-haproxy-dist.sh: line 327: ulimit: core file size: cannot modify limit: Operation not permitted

and under this script's `set -euxo pipefail`, that nonzero exit aborted
the whole script before the strace-wrapped `haproxy -v`, the core-file
search, or the gdb backtrace ever ran. No diagnostic signal was
captured.

Make `ulimit -c unlimited` non-fatal by running it as an `if` condition
(which `set -e` does not treat as abort-worthy), logging the hard limit
and the outcome instead of relying on it to succeed. Since the hard
limit being locked to 0 may be a permanent constraint of this
runner/container, not just this run, don't depend on core dumps at all:
`strace -f`'s own output (last syscalls, mapped files, and the signal
delivery itself) is now the primary signal, and its full log is dumped
to the job log whenever the invocation exits nonzero. The core-file
search and gdb backtrace stay as a bonus path if a core does happen to
appear, but nothing depends on it anymore.

Still a TEMPORARY riscv64-only diagnostic, not a fix.

Upstream-Status: Inappropriate [temporary riscv64 CI diagnostic, not meant to be carried or upstreamed]
…as haproxy's exit code

The diagnostic added in 288a119/058f6396 wrapped the crashing haproxy -v
invocation in `strace -f -o ... env ...` unconditionally and took STATUS
from that wrapped command's exit code. The fixed diagnostic just ran for
the first time and reported real signal:

    haproxy -v exit status: 127
    ./ci/build/build-haproxy-dist.sh: line 338: strace: command not found

Exit 127 is bash's own "command not found" status, not haproxy's. `dnf
install -y gdb strace 2>/dev/null || true` silently left no working
`strace` on PATH, so the shell failed to exec `strace` itself before
haproxy -v ever ran (wrapped or otherwise). This diagnostic run produced
zero information about the original segfault.

Restructure so a missing diagnostic tool can't be mistaken for the thing
being diagnosed: run the plain, unwrapped invocation first and on its
own so STATUS is always its real exit code, then gather diagnostics
afterwards and independently of STATUS -- LD_DEBUG=all (glibc's own
dynamic-linker trace, built into every glibc so it can't hit this same
"tool missing" failure mode), a guarded `ldd`, and `strace` only when
`command -v strace` actually finds it on PATH.

Still a TEMPORARY riscv64-only diagnostic, not a fix.
…crash

haproxy -v really does segfault on riscv64 (status 139), plain and under
LD_DEBUG=all alike -- confirmed by the last CI run. But the LD_DEBUG=all
re-run was a bare command under `set -e`, unlike the plain invocation
above it which is guarded with `set +e` / capture-status / `set -e`. Its
own segfault therefore aborted the script immediately, before the
following echo/cat of /tmp/haproxy-v-lddebug.log could run -- the trace
glibc's dynamic linker had already written to that file (unbuffered, as
it goes) never reached the job log.

Guard the LD_DEBUG=all re-run the same way the plain invocation is
guarded, and print the log unconditionally afterwards regardless of the
re-run's own exit status, tailed to the last 500 lines since the trace
can be large and the crash-adjacent lines at the end are what matter.
The previous diagnostic round finally captured a clean, correctly-guarded
LD_DEBUG=all re-run, and it came back completely empty. glibc writes that
trace via direct, unbuffered write(2) calls rather than through stdio, so
the existing whole-process stdout/stderr redirect isn't expected to be
swallowing it. Add a second, independent re-run using LD_DEBUG_OUTPUT,
which glibc opens and writes to directly by path, to tell a buffering
artifact in this harness apart from a crash too early for ld.so's own
debug-print code to run at all.

Also confirmed from the build flags: no CPU= or ARCH= is passed to
HAProxy's make, HAProxy's own Makefile defaults for TARGET=linux-glibc add
no -march/-mcpu on an unset CPU/ARCH, and a leaked -march=native would
fail to compile outright on this toolchain (see build-pyscipopt.yml)
rather than merely segfault at runtime -- no evidence of an ISA mismatch
from the build flags themselves.

Still a TEMPORARY riscv64-only diagnostic, not a fix.
The LD_DEBUG_OUTPUT re-run finally ran cleanly last round but the log
still cut off right after its "exit status: 139 (informational only)"
line, with the job then failing at exit code 2 -- two bugs in the
diagnostic harness itself, not new signal about the crash:

- glibc appends the PID to LD_DEBUG_OUTPUT's path (e.g.
  "/tmp/haproxy-v-lddebug-output.<pid>"), so locating the file needs a
  glob.
- The glob was done via `ls ... | head -n1`, and this script inherits
  `set -euxo pipefail` from its `bash -c` wrapper (bash propagates
  SHELLOPTS to child scripts). With no file matching, `ls` exits 2 and,
  under pipefail, that becomes the pipeline's own status even though
  `head` succeeds -- which errexit then treats as a failed assignment
  and kills the script before the following `if` ever runs.

Replace it with a bare glob `for` loop (no pipeline to fail under
pipefail) that prints every matching file's content, and says so
explicitly when no file matches at all, distinct from a matched file
that's merely empty.

Still a TEMPORARY riscv64-only diagnostic, not a fix.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant