Repository navigation
RISC-V runner segfault with cache #1139
Description
Activity
Yeah, this is very odd.
I do not think this is one real bug for caching compiler-rt on the runner, as we have many jobs already, including in my repo forked, this is the first time to see it. But I do not have evidence either.
Or was another memory-intensive job running at the same time? I suspect this. BTW, I will do my best to avoid running other jobs on the host.
I'll just close this since I don't think there is something actionable. Just something to keep an eye on if it happens again.
I don't think there is any problem running other jobs. Shouldn't segfault either way of course :)
Something similar with a different cache point, this time SIGILL https://github.com/rust-lang/compiler-builtins/actions/runs/23731213183/job/69125229566

hmm, I will add one new riscv64 hardware as the runner to see if the situation improves.
Another one, this time in the post-run cache https://github.com/rust-lang/compiler-builtins/actions/runs/23532708278/job/69228667733
Wonder if some cache-related tool is miscompiled for the arch. I'll reopen since this seems to be recurring.
It's happening quite often https://github.com/rust-lang/compiler-builtins/actions/runs/23770665652/job/69261142904?pr=1142, https://github.com/rust-lang/compiler-builtins/actions/runs/23767883361/job/69251906199.
Maybe you could enable core dumps for the runners? To do this I think you can edit
/proc/sys/kernel/core_patternto contain/var/crash/core.%e.%p.%ton the host (can't be done from docker) then addulimit -c unlimitedas one of the first Docker run commands. Then a job to unconditionally upload like the following can work:- name: Upload core dumps uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6 if: always() with: name: core-dump path: "/var/crash/core.*"
at least that should tell us what command is actually failing and give us somewhere to look.
Reacted by vimer
The run at https://github.com/rust-lang/compiler-builtins/actions/runs/23674147811/job/68973834693?pr=1138 failed with the following:
I don't think it's happened consistently but fyi in case you have any ideas @yuzibo.