Skip to content

Judge the race step by what the detector found - #50

Merged
ShawnChen-Sirius merged 2 commits into
mainfrom
fix/race-step-verdict
Aug 13, 2026
Merged

ShawnChen-Sirius merged 2 commits into
mainfrom
fix/race-step-verdict

Conversation

@ShawnChen-Sirius

@ShawnChen-Sirius ShawnChen-Sirius commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

main has been red since #45 on one step, and not for anything in this repository.
go test -race reports every test as PASS and then the process segfaults on the way
out, so the step fails on an exit status that says nothing about the tests.

What the core dump says

#0  runtime.raise
#1  runtime.raisebadsignal (sig=11)
#2  runtime.badsignal (sig=11)
#3  runtime.sigtrampgo
#4  runtime.sigtramp

runtime.badsignal is Go's path for a signal on a thread it does not own — a C
thread. __tsan_fini and __run_exit_handlers are on the stack, and 33 threads
are parked in __pthread_cond_wait_common inside libchdb
.

So the engine's process-global ClickHouse pools are still alive at exit, because
nothing can stop them: v26.7.0's C ABI has 50 functions and none shuts the engine
down. Under -race, Go's exit path runs racefini → __tsan_fini, tearing the
sanitizer runtime down while those threads can still wake into instrumented code.
One faults, Go re-raises because it is not a Go thread, and the binary dies after
every test has passed.

This is why it only happens under -race: without it, the same leaked threads are
killed by exit_group and nobody notices.

Waiting does not fix it

Worth measuring rather than assuming. Crashes over 48 runs, six in parallel:

wait before exit crashes
none 5
500ms 3
2s 7
5s 3

No trend. BackgroundSchedulePool re-arms its tasks on a timer, so the threads
never go quiet — a longer wait just picks a different moment to gamble on. (Those
counts are non-zero exits, which at six-way parallelism include some genuine
contention failures, so they overstate the segfault rate. The absence of a trend is
the point.)

What changes

The step fails on what the race detector is for: a reported data race, a failing
test, a panic. A segfault at exit with nothing else wrong passes, with a warning.

Coverage is unchanged — this is only how the step reads its own result — and it is
not a blanket exemption. A run that segfaults and fails a test still fails.

Verified

In a Linux container and on macOS, against a build that reproduces the crash:

  • ordinary pass → 0
  • hits the exit segfault → 0, and says so
  • hits it while also failing a test → 1
  • injected data race → 1
  • injected failing test → 1
  • injected panic → 1

When to remove this

Once the engine can be shut down. chdb-core has the machinery already —
GlobalThreadPool::shutdown() in src/Common/ThreadPool.h — and no C entry point
calls or exposes it. Then the plain exit status is trustworthy and this wrapper
should go.

🤖 Generated with Claude Code

Note

Judge the race detector test step by inspecting output rather than exit status

  • Adds race-test.sh, a wrapper around go test -race that parses output to determine pass/fail instead of relying on the raw exit code.
  • Fails the step on data races, test failures, panics, or fatal in-test signals; passes (with a warning) on a post-PASS exit-time segfault matching PASS immediately followed by signal: segmentation fault.
  • Updates chdb.yml to call the new script in both Linux and macOS race test steps.
  • Behavioral Change: a specific exit-time segfault that previously failed the CI step now exits with status 0 and emits a ::warning:: instead of a ::error::.

Macroscope summarized 12cb828.

main has been red since #45 on one step, and not for anything in this repository.
`go test -race` reports every test as PASS and then the process segfaults on the
way out, so the step fails on an exit status that says nothing about the tests.

From a core dump: the crashing thread's stack is runtime.raise <-
raisebadsignal <- badsignal <- sigtrampgo, which is Go's path for a signal on a
thread it does not own — a C thread. __tsan_fini and __run_exit_handlers are on
the stack, and 33 threads are parked in __pthread_cond_wait_common inside
libchdb. So: the engine's process-global ClickHouse pools are still alive at exit
because nothing can stop them — v26.7.0's C ABI has 50 functions and none shuts
the engine down — and under -race Go's exit path tears the sanitizer runtime down
while those threads can still wake into instrumented code. One faults, Go
re-raises because it is not a Go thread, and the binary dies after passing.

Waiting before exit does not fix it, which was worth measuring rather than
assuming: crashes at waits of none, 500ms, 2s and 5s came out 5, 3, 7 and 3 out
of 48, no trend. BackgroundSchedulePool re-arms its tasks on a timer, so the
threads never go quiet — a longer wait just picks a different moment to gamble on.

So the step now fails on what the race detector is for. A reported data race, a
failing test, a panic: red. A segfault at exit with nothing else wrong: green,
with a warning. Coverage is unchanged — this is only how the step reads its own
result — and it is not a blanket exemption: a run that segfaults *and* has a real
failure still fails, which is one of the cases exercised below.

Verified in a Linux container and on macOS, against a build that reproduces the
crash: an ordinary pass exits 0; a run that hits the exit segfault exits 0 and
says so; a run that hits it while also failing a test exits 1; an injected data
race exits 1; an injected failing test and an injected panic exit 1.

Remove this once the engine can be shut down. chdb-core has the machinery already
— GlobalThreadPool::shutdown() in src/Common/ThreadPool.h — and no C entry point
calls or exposes it. Then the plain exit status is trustworthy again and this
wrapper should go.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread .github/scripts/race-test.sh Outdated
"The log mentions a segfault" covered more than intended. A fault during a test —
#46, which was a real one — also says signal: segmentation fault somewhere, and
would have been waved through with no failing test or panic to catch it.

The two are distinguishable, and by more than wording. A fault on a thread Go owns
gets the runtime's own handler, which prints SIGSEGV: segmentation violation and a
goroutine dump before dying. The exit-time one lands on a libchdb thread, so Go's
badsignal re-raises and there is no dump at all — only the harness line. Checked
against both: #46's CI log has the runtime header and no harness line; the
exit-time crash has the harness line and no header.

So a runtime signal header now fails the step outright, and the exemption wants
positive evidence — a PASS line with the segfault on the very next one, which is
what "every test in this package passed, then the binary died" looks like.

Re-ran the injected data race, failing test and panic against the tightened
version, and replayed both real crash logs through the new rules offline: the
in-test crash is rejected, the exit-time crash is exempted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ShawnChen-Sirius

Copy link
Copy Markdown
Contributor Author

@wudidapaopao please review this PR

@ShawnChen-Sirius
ShawnChen-Sirius merged commit bfccbd0 into main Aug 13, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant