This repository vendors third-party resources and adapts them into test cases
for RocJITsu regression coverage. The main pytest entrypoint is tests/test_corpus.py.
corpus/
iree/ IREE run-module cases and target configs.
kernels/ HIP kernel reproducers, CMake runners, and vendored sources.
cts/ HIP semantic tests organized by target family.
dbt/ Offline DBT translation profiles.
semantics/ Standalone target-specific HIP semantic programs.
llama/ llama.cpp test-backend-ops cases and vendored GGML sources.
tensile/ gfx1250 TensileLite configs and generated artifacts.
tests/
test_corpus.py Unified pytest entrypoint.
test_suites/ Suite adapters used by test_corpus.py.
support/ Shared discovery, target, build, and run helpers.
scripts/
run_gfx1250_regression.sh
extract_gfx1250_hsacos.py
... additional corpus helper scripts
requirements.txt Python packages for pytest and corpus helpers.
corpus/iree/: IREE HIP e2e and matmul cases described by JSON files. The suite compiles withiree-compileand runs withiree-run-module.corpus/kernels/: standalone HIP kernel reproducers. Current backends includehip-stream-k,hip-matmul,hipkittens, androcblas.corpus/cts/: HIP semantic tests organized by target family, including FPSan-derived floating-point cases and standalone integer ISA cases. The gfx1250 integer suite covers the portable RDNA4 integer families, both gfx1250 integer matrix forms, and selected gfx1250-specific instructions. Its build-time coverage gate extracts the gfx1250 image from each linked test executable and verifies that every expected opcode is present in the code that the runtime test executes.corpus/semantics/: standalone HIP programs with deterministic inputs, source-ISA coverage, and typed results that can be captured under any externally selected launch configuration.corpus/llama/: selectedllama.cpptest-backend-opscases with the pinned GGML sources they were measured against. Seecorpus/llama/README.md.corpus/tensile/: gfx1250 TensileLite YAML configs, manifests, numeric smoke lists, and generated HSACO/code-object artifacts. This corpus is run by the Tensile scripts, not bytests/test_corpus.py.
tests/test_corpus.py discovers and runs the iree, kernels, cts, dbt,
semantics, and llama suites. By default it uses target gfx1201 and selects
the first three; dbt, semantics, and llama are opt-in.
The gfx1250-only memory CTS contains selected deterministic representatives for scalar, S_BUFFER, global/FLAT, buffer, scratch, LDS, atomic, block, asynchronous, transpose, and tensor-memory behavior. Its buffer coverage executes stride-scale, swizzle-enable, and OOB-select descriptor modes with host-side address and bounds oracles. Run an individual case through the normal CTS entrypoint and an appropriate gfx1250 runtime or simulator wrapper:
rocjitsu --config /path/to/gfx1250.json -- \
pytest tests/test_corpus.py \
--target gfx1250 \
--suite cts \
--case memory_isa_gfx1250_tensor_test \
--timeout 30The default CMake build for corpus/cts runs
memory_isa_gfx1250_source_coverage_test. That gate extracts the linked
gfx1250 offload image with the selected SDK llvm-objdump and verifies the
manifest's required opcodes and material modifiers. It establishes instruction
presence, not functional correctness; the runtime GTests provide the semantic
oracles. After configuring a CTS build directly, build the linked-image gates
and run the complete memory collection with:
cmake --build <cts-build-directory>
rocjitsu --config /path/to/gfx1250.json -- \
ctest --test-dir <cts-build-directory> \
--output-on-failure \
-R '^memory_isa_gfx1250_'Coverage is representative rather than exhaustive. The current suite excludes
the exhaustive cross-product of S_BUFFER and VBUFFER widths, address modifiers,
and descriptor modes; the exhaustive FLAT width, atomic, and modifier matrix
across address spaces; the full atomic type/address-space matrix; cluster-mask
cross-products for clusters larger than two workgroups; physical multicast request
coalescing and timeout fidelity; cache, transaction-count, and prefetch behavior;
MWAIT; fault delivery; and timing fidelity.
Representative FLAT B32 routing across global, LDS, and scratch
addresses is covered. Cluster coverage pins clustered dispatch, workgroup ranks,
B32 load and async-to-LDS execution, and emitted M0, wait, and barrier lowering. All four
M0[1:0] workgroup-mask values are emitted and executed; M0=0, 1, and 2 have
destination-selection oracles, while M0=3 checks visible results for both requesting
workgroups. The suite does not prove that matching requests were physically coalesced. The
manifest records the same boundary for machine-readable audits.
The scratch misalignment case targets
SH_MEM_CONFIG.alignment_mode=UNALIGNED: its base-plus-immediate address is byte
offset 157, and its byte-exact oracle is not the expected result for modes that
automatically align DWORD accesses.
# Install Python packages in your virtual environment.
python -m pip install -r requirements.txt
# Check that the IREE tools are available.
command -v iree-compile iree-run-module
# Check ROCm SDK discovery, or set ROCM_PATH explicitly.
rocm-sdk path --rootPreview the selected pytest cases without running them:
pytest --collect-only tests/test_corpus.pyRun through RocJITsu with all valid test suites for that target:
rocjitsu --config /path/to/gfx1201.json -- \
pytest tests/test_corpus.py \
--target gfx1201 \
--timeout 15Run the RocJITsu corpus matrix for all configured gfx targets:
ROCM_VENV=path/to/.venv \
ROCJITSU_WORKSPACE=path/to/rocjitsu-workspace \ # contains configs and rocjitsu binary
ROCJITSU_EXE=path/to/rocjitsu-binary \
./scripts/run_rocjitsu_corpus_matrix.shRun selected suites:
rocjitsu --config /path/to/gfx1201.json -- \
pytest tests/test_corpus.py \
--target gfx1201 \
--suite kernels,cts \
--timeout 15Run selected cases:
rocjitsu --config /path/to/gfx1201.json -- \
pytest tests/test_corpus.py \
--target gfx1201 \
--suite cts \
--case fpsan_wmma \
--timeout 15Run suite runtime commands through a wrapper:
pytest tests/test_corpus.py \
--target gfx1201 \
--run-wrapper "rocjitsu --config ${CONFIG} --"Run with a list of tests to skip:
rocjitsu --config /path/to/gfx1201.json -- \
pytest tests/test_corpus.py \
--target gfx1201 \
--skip-tests-config tests/gfx1201_skip_tests.example.jsonUseful selectors:
--target <gfx target>: target to run, for examplegfx942,gfx950,gfx1201, orgfx1250.--suite <iree|kernels|cts|dbt|semantics|llama>: include a suite. Repeat or pass comma-separated values.--exclude-suite <suite>: exclude a suite.--backend <backend>: include a kernel backend such ashipkittens.--exclude-backend <backend>: exclude a kernel backend.--case <selector>: include a case by case id or selector name.--exclude-case <selector>: exclude a case.--artifact-directory <path>: write build artifacts, logs, and generated outputs somewhere other than.pytest-artifacts.--run-wrapper <command>: prepend a shell-style command prefix to supported suite runtime commands.--comparison-run-wrapper <command>: run each selected semantic program through a second wrapper and compare its typed results exactly with the self-checking--run-wrapperresults. This option requires selecting only thesemanticssuite.--comparison-required-stderr <text>: require text in each comparison run's stderr; repeat for multiple external activation checks.--timeout <seconds>: fail an individual pytest case if it exceeds this runtime. This is provided bypytest-timeoutand is not a timeout for the entire script; use--session-timeout <seconds>for a whole-session limit.--skip-all-runs: build or compile only where supported.--dbt-corpus <path>: packaged-HSACO extraction root for the opt-indbtsuite; defaults toROCJITSU_HSACO_CORPUS.--dbt-translator <path>:rj_dbt_translateexecutable; it can also be resolved fromRJ_DBT_TRANSLATE,ROCJITSU_BUILD_DIR,ROCJITSU_BUILD, orPATH.--dbt-llvm-objdump <path>: gfx1250-capable TheRockllvm-objdumpused to disassemble each successful translated object; it can also be resolved fromROCJITSU_DBT_LLVM_OBJDUMP, beside the translator, orPATH.--dbt-package-lock <path>: consumer-owned producer-package versions, file-only manifest digest, and extraction-size rules.--dbt-expected-failures <path>: consumer-owned strict expected-failure manifest.--dbt-expected-rewrites <path>: consumer-owned manifest recording how many source instructions require rewriting in selected inputs.--dbt-timeout <seconds>: override the profile's per-object translation timeout.--dbt-memory-limit-mib <MiB>: override the profile's per-object translator resident-memory limit.--dbt-allow-incomplete-corpus: explicitly consume an extraction that records"complete": falsesolely because CCOB materialization was skipped.
The opt-in semantics suite builds standalone gfx1250 HIP programs, verifies
their declared source instructions, and runs them directly or through the
repository-wide wrapper:
ROCM_PATH=/path/to/rocm-sdk \
pytest tests/test_corpus.py \
--target gfx1250 \
--suite semantics \
--run-wrapper "/path/to/launcher --config /path/to/config.json --"The package also provides generic capture and comparison utilities. External
workflows choose simulator or physical hardware, original or prepared
binaries, environment variables, and provenance checks. See
corpus/semantics/gfx1250/README.md for
the build, capture, comparison, and hardware-golden flow.
The opt-in llama suite runs 535 selected test-backend-ops cases that
compare the GGML HIP backend against its CPU reference. corpus/llama/README.md
covers how the inventory was selected and what the vendored sources contain.
Pytest builds test-backend-ops once for the selected target and then runs one
process per case. Under xdist, one worker holds a cross-worker build lock while
the other workers wait and reuse the completed executable. The CMake job count
matches the pytest worker count by default. Set LLAMA_CORPUS_BUILD_WORKERS on
the pytest command to override only the CMake -j value.
The build needs a ROCm SDK root providing lib/cmake/hip/hip-config.cmake,
hipBLAS, and rocBLAS:
export ROCM_PATH="$(rocm-sdk path --root)"
LLAMA_CORPUS_BUILD_WORKERS=16 \
pytest tests/test_corpus.py \
--target gfx1201 \
--suite llama \
--timeout 15 \
-n 8Run the same cases through RocJITsu. The run wrapper applies to each harness process, while pytest remains responsible for test timeouts:
LLAMA_CORPUS_BUILD_WORKERS=16 \
pytest tests/test_corpus.py \
--target gfx1201 \
--suite llama \
--run-wrapper "rocjitsu --config /path/to/gfx1201.json --" \
--timeout 15 \
-n 8The llama adapter has no internal timeout. Use the pytest-timeout option
--timeout <seconds> when a limit is needed. Every completed case writes
<artifact-directory>/llama/<target>/cases/<op>.<digest>.log with the
command, outcome, and captured output, plus a matching *.outcome.json record.
The slowest gfx1201 case under RocJITsu needs about 14 seconds
serially, so eight workers contending for one GPU can push it over the limit and
report a timeout that a serial run does not reproduce. Re-run a lone unexpected
timeout without -n to confirm it, or raise pytest's --timeout.
Select cases by operator name, by case digest, or by the exact case string:
# Every MUL_MAT case.
pytest tests/test_corpus.py --suite llama --case MUL_MAT
# One case, by the digest that appears in its test id.
pytest tests/test_corpus.py --suite llama --case 67c2b413a2ab--case splits on commas, so exact OP(params) strings only work through the
JSON lists of --run-tests-config and --skip-tests-config.
The packaged-HSACO workflow inventories the gfx1250 code objects shipped in an
active TheRock/PyTorch venv. It deliberately stops at extraction so the same
content-addressed corpus can feed separate offline translation and analysis
workflows. The large generated corpus stays in ignored results-* directories
rather than being committed.
Install the explicit gfx1250 PyTorch device package into the same venv as
TheRock. Installing plain torch can select a build without the gfx1250
payload. The consumer supplies the requirements file so package updates and
their corresponding DBT expectations stay in one change:
export ROCM_VENV=/path/to/venv
export DBT_REQUIREMENTS=/path/to/requirements-gfx1250-dbt.txt
uv pip install \
--python "$ROCM_VENV/bin/python" \
-r "$DBT_REQUIREMENTS"
"$ROCM_VENV/bin/rocm-sdk" init
"$ROCM_VENV/bin/python" -c \
'import torch; print(torch.__version__, torch.version.hip, torch.cuda.get_arch_list())'Extract the packaged code objects:
"$ROCM_VENV/bin/python" scripts/extract_gfx1250_hsacos.py \
--environment "$ROCM_VENV" \
--destination results-gfx1250-packaged \
--materialize-ccobAdditional unpacked ROCm distributions can be scanned into the same corpus. The extractor combines these files with the selected venv package roots and deduplicates identical code objects by their complete-HSACO SHA-256. For example, the gfx1250 DBT corpus currently uses the July 26 multi-architecture test distribution:
export TEST_DIST_VERSION=7.15.0a20260726
export TEST_DIST_ARCHIVE="$ROCM_VENV/therock-dist-linux-multiarch-tests-${TEST_DIST_VERSION}.tar.gz"
export TEST_DIST_ROOT="$ROCM_VENV/test-distribution-${TEST_DIST_VERSION}"
curl --fail --location \
--output "$TEST_DIST_ARCHIVE" \
"https://rocm.nightlies.amd.com/tarball-multi-arch/therock-dist-linux-multiarch-tests-${TEST_DIST_VERSION}.tar.gz"
printf '%s %s\n' \
0842af33a89376555df161b6a8d122dc32083fc2f4867ee7236b2258b479c567 \
"$TEST_DIST_ARCHIVE" | sha256sum --check --strict
mkdir "$TEST_DIST_ROOT"
tar -xzf "$TEST_DIST_ARCHIVE" -C "$TEST_DIST_ROOT"
"$ROCM_VENV/bin/python" scripts/extract_gfx1250_hsacos.py \
--environment "$ROCM_VENV" \
--additional-root "$TEST_DIST_ROOT" \
--destination results-gfx1250-packaged \
--materialize-ccob--additional-root is repeatable and each root must be beneath the selected
venv. This makes paths in manifests/provenance.jsonl independent of the
checkout or CI workspace location. Additional roots must not overlap.
The extractor covers:
- TheRock and PyTorch KPACK archives;
- loose loadable AMDGPU ELF code objects;
- gfx1250 entries in HIP offload bundles;
- directly embedded gfx1250 AMDGPU ELFs;
- AOTriton ZIP/AKS2 image stores;
- CCOB containers materialized through HIP.
All methods except CCOB are file-only. --materialize-ccob needs a visible
gfx1250 GPU and a working /dev/kfd; omit it for a fully offline extraction
subset. The resulting summary.json then has "complete": false, making the
omission explicit.
Deduplication is by the SHA-256 of the complete HSACO bytes. Each unique object
is written once as objects/<sha256>.hsaco; duplicate KPACK members, loose
files, embedded images, and AOTriton records retain separate entries in
manifests/provenance.jsonl. summary.json reports both source_records and
unique_code_objects, while manifests/SHA256SUMS,
manifests/NON_CCOB_SHA256SUMS, and manifests/packages.json preserve
integrity, the pinned file-only baseline, and environment details.
The extractor is deterministic for an unchanged environment: filesystem,
archive, record, and checksum orderings are canonical; JSON keys are sorted;
and generated manifests contain no timestamps or destination-dependent paths.
Two runs against the same venv should have identical objects/ contents and
byte-identical files under manifests/ plus an identical summary.json.
Verify every extracted object against its content-addressed filename:
(
cd results-gfx1250-packaged
sha256sum --quiet -c manifests/SHA256SUMS
)Consumers should use manifests/SHA256SUMS or the flat objects/*.hsaco
directory as their input list. Translation results and policy belong in their
own result directory and should not be mixed back into the extraction corpus.
Translation is a separate, read-only consumer of a completed extraction. The
dbt suite sends each unique object to rj_dbt_translate with input revision
b0 and output revision a0. It verifies each input hash and checks
code-object output has bounded ELF tables, executable loadable content, and a
non-empty AMDGPU metadata note. Hash-pinned objects selected by the consumer additionally
run in diff mode and must reproduce the expected number of source instructions
requiring rewrite. The output instruction sequence remains free to change. The
pinned profile gives each translator process a 30-second timeout and 4 GiB
resident-memory limit, while stdout and stderr are spooled to temporary files
so large outputs do not accumulate in pytest worker memory. Every successful
output must also disassemble with the selected TheRock llvm-objdump,
independently exercising LLVM's ELF and gfx1250 ISA readers.
The consumer-owned package lock ties the extraction to its producer packages. It records the relevant TheRock, PyTorch, and Triton versions, the file-only manifest digest, and object-count rules. Collection fails before any xfail or rewrite expectation is applied if the extraction does not match that lock. Expected failures and rewrites are consumer-owned as well, so package and translator updates can adjust all coupled inputs in one change.
Keep extraction and translation as separate commands, but sequence them so pytest only runs after extraction succeeds:
export ROCM_VENV=/path/to/venv
export ROCJITSU_HSACO_CORPUS="$PWD/results-gfx1250-packaged"
"$ROCM_VENV/bin/python" scripts/extract_gfx1250_hsacos.py \
--environment "$ROCM_VENV" \
--destination "$ROCJITSU_HSACO_CORPUS" \
--materialize-ccob &&
"$ROCM_VENV/bin/python" -m pytest tests/test_corpus.py \
--target gfx1250 \
--suite dbt \
--dbt-translator /path/to/build/tools/rj_dbt_translate \
--dbt-llvm-objdump /path/to/therock/lib/llvm/bin/llvm-objdump \
--dbt-package-lock /path/to/package_lock.json \
--dbt-expected-failures /path/to/expected_failures.json \
--dbt-expected-rewrites /path/to/expected_rewrites.json \
-n 8An offline extraction that intentionally omitted GPU-only CCOB materialization
has "complete": false. Opt in explicitly when testing that subset:
pytest tests/test_corpus.py \
--target gfx1250 \
--suite dbt \
--dbt-translator /path/to/build/tools/rj_dbt_translate \
--dbt-llvm-objdump /path/to/therock/lib/llvm/bin/llvm-objdump \
--dbt-package-lock /path/to/package_lock.json \
--dbt-expected-failures /path/to/expected_failures.json \
--dbt-expected-rewrites /path/to/expected_rewrites.json \
--dbt-allow-incomplete-corpus \
-n 8Expected failures are keyed by the full input SHA-256 and additionally constrain the failure class, process return code, and diagnostic. An unlisted failure, changed diagnostic, timeout of the wrong object, or successful translation of an xfailed object fails the suite. Only expected and unexpected failures write diagnostic logs. Successful output stays file-backed during validation and is deleted afterward.