ttl-insert-cb-sync (with #1108) rejects a block that is still held when a nested region (for example a loop) acquires the same DFB again, because the nested acquisition returns the held block's slot. The check is skipped when any release of the DFB precedes the nested region (hasOwnedReleaseBeforeBoundary in TTLInsertCBSync.cpp), even if that release covers only part of the held block. An external call can declare such a partial release (ttl.DFBEffect.pop(dfb, tiles=1) on a two-tile block lowers to ttl.opaque_call ... dfb_effects [#ttl.dfb_protocol_effect<pop, 0, 1>]), so the remaining tile stays held while the loop re-acquires the DFB, and a later use of the block reads a slot the loop has consumed. Main has no nested-acquisition check and accepts both variants below.
To Reproduce
The IR is what the frontend emits for a data-movement kernel that waits for a two-tile block, calls an external function declaring a one-tile pop, waits on the same DFB in a loop, and reads the held block afterwards.
// ttlang-opt partial_release_nested_partial.mlir --pass-pipeline='builtin.module(func.func(ttl-insert-cb-sync))'
func.func @held_block_across_loop(%arg0: tensor<1x6x!ttcore.tile<32x32, bf16>, #ttl.layout<shape = [32, 192], element_type = !ttcore.tile<32x32, bf16>, buffer = system_memory, grid = [1, 1], memory = interleaved>>)
attributes {ttl.kernel_thread = #ttkernel.thread<noc>} {
%c0 = arith.constant 0 : index
%c1 = arith.constant 1 : index
%c2 = arith.constant 2 : index
%cb = ttl.bind_cb{cb_index = 0, block_count = 3} : !ttl.cb<[1, 2], !ttcore.tile<32x32, bf16>, 3>
%held = ttl.cb_wait %cb : <[1, 2], !ttcore.tile<32x32, bf16>, 3> -> tensor<1x2x!ttcore.tile<32x32, bf16>>
// An external call pops one of the held block's two tiles.
ttl.opaque_call "partial_pop" dfb_dependencies(%cb : !ttl.cb<[1, 2], !ttcore.tile<32x32, bf16>, 3>) dfb_effects [#ttl.dfb_protocol_effect<pop, 0, 1>] () {header = "shim.hpp"} : () -> ()
scf.for %i = %c0 to %c2 step %c1 {
%next = ttl.cb_wait %cb : <[1, 2], !ttcore.tile<32x32, bf16>, 3> -> tensor<1x2x!ttcore.tile<32x32, bf16>>
%slice = ttl.tensor_slice %arg0[%c0, %c0] : tensor<1x6x!ttcore.tile<32x32, bf16>, #ttl.layout<shape = [32, 192], element_type = !ttcore.tile<32x32, bf16>, buffer = system_memory, grid = [1, 1], memory = interleaved>> -> tensor<1x2x!ttcore.tile<32x32, bf16>, #ttl.layout<shape = [32, 192], element_type = !ttcore.tile<32x32, bf16>, buffer = system_memory, grid = [1, 1], memory = interleaved>>
%x = ttl.copy %cb, %slice : (!ttl.cb<[1, 2], !ttcore.tile<32x32, bf16>, 3>, tensor<1x2x!ttcore.tile<32x32, bf16>, #ttl.layout<shape = [32, 192], element_type = !ttcore.tile<32x32, bf16>, buffer = system_memory, grid = [1, 1], memory = interleaved>>) -> !ttl.transfer_handle<write>
ttl.wait %x : !ttl.transfer_handle<write>
ttl.cb_pop %cb : <[1, 2], !ttcore.tile<32x32, bf16>, 3>
}
// The held block is read after the loop re-acquired the DFB.
%element = ttl.raw_element_read %held[%c0, %c0] : tensor<1x2x!ttcore.tile<32x32, bf16>> -> bf16
func.return
}
ttlang-opt partial_release_nested_partial.mlir --pass-pipeline='builtin.module(func.func(ttl-insert-cb-sync))'
Observed with #1108 (9ac464a3c) and on main 546f6b20a: accepted; the output keeps ttl.cb_wait, the external call, the loop's ttl.cb_wait/ttl.cb_pop, and the read after the loop, with no release of the held block's remaining tile.
The same program without the external call (the ttl.opaque_call line removed) is rejected by #1108:
error: dataflow buffer block is still acquired when a nested region acquires the same buffer again; release it before that region (pop it) or acquire the block inside the region
Expected: the partial-release variant is rejected the same way. Count the tiles released before the nested region and bypass the nested-acquisition analysis only when they cover the whole block (on every path), as #1108 already does for the data-movement check before the next same-level acquisition (hasReleaseBefore).
Found by Copilot on #1108; #1108 lists it as a known limitation.
ttl-insert-cb-sync(with #1108) rejects a block that is still held when a nested region (for example a loop) acquires the same DFB again, because the nested acquisition returns the held block's slot. The check is skipped when any release of the DFB precedes the nested region (hasOwnedReleaseBeforeBoundaryinTTLInsertCBSync.cpp), even if that release covers only part of the held block. An external call can declare such a partial release (ttl.DFBEffect.pop(dfb, tiles=1)on a two-tile block lowers tottl.opaque_call ... dfb_effects [#ttl.dfb_protocol_effect<pop, 0, 1>]), so the remaining tile stays held while the loop re-acquires the DFB, and a later use of the block reads a slot the loop has consumed. Main has no nested-acquisition check and accepts both variants below.To Reproduce
The IR is what the frontend emits for a data-movement kernel that waits for a two-tile block, calls an external function declaring a one-tile pop, waits on the same DFB in a loop, and reads the held block afterwards.
ttlang-opt partial_release_nested_partial.mlir --pass-pipeline='builtin.module(func.func(ttl-insert-cb-sync))'Observed with #1108 (
9ac464a3c) and on main546f6b20a: accepted; the output keepsttl.cb_wait, the external call, the loop'sttl.cb_wait/ttl.cb_pop, and the read after the loop, with no release of the held block's remaining tile.The same program without the external call (the
ttl.opaque_callline removed) is rejected by #1108:Expected: the partial-release variant is rejected the same way. Count the tiles released before the nested region and bypass the nested-acquisition analysis only when they cover the whole block (on every path), as #1108 already does for the data-movement check before the next same-level acquisition (
hasReleaseBefore).Found by Copilot on #1108; #1108 lists it as a known limitation.