Skip to content

[ttl][dfb] Sync inserts an extra release for a guarded block that a nested if pops #1156

Description

@brnorris03

ttl-insert-cb-sync inserts a release for a guarded acquisition (an acquisition in an scf.if then-region) when it finds no owned release at the guard's level, but it does not see releases nested in a further scf.if/else inside that region (findOwnedDFBReleases in DFBAcquireReleaseAnalysis.cpp skips a release whose projection into the ordering block is the guard itself). A block that is popped on both branches of a nested if therefore gets an extra inserted pop: two pops for one wait. The next wait of the DFB then reads a block that was never published to it, and the consumer runs one block ahead of the producer.

To Reproduce

import os

os.environ["TTLANG_COMPILE_ONLY"] = "1"

import torch
import ttl
import ttnn
from ttl import ttl_api

ttl_api._device_target_arch = lambda _runtime_args: "blackhole"


@ttl.operation(grid=(1, 1))
def guarded_nested_pop(inp, out):
    x = ttl.make_dataflow_buffer_like(inp, shape=(1, 1), block_count=4)
    o = ttl.make_dataflow_buffer_like(inp, shape=(1, 1), block_count=2)

    @ttl.compute()
    def compute():
        node_x, node_y = ttl.node(dims=2)
        r = o.reserve()
        if node_x == 0:
            first = x.wait()
            r.store(first)
            if node_y == 0:  # both branches pop the guarded block
                first.pop()
            else:
                first.pop()
            for i in range(2):
                second = x.wait()
                r += second
                second.pop()
        r.push()

    @ttl.datamovement()
    def reader():
        for j in range(3):
            with x.reserve() as blk:
                ttl.copy(inp[0, j], blk).wait()

    @ttl.datamovement()
    def writer():
        with o.wait() as blk:
            ttl.copy(blk, out[0, 0]).wait()


tensors = [
    ttnn.from_torch(torch.zeros((32, 96), dtype=torch.bfloat16), dtype=ttnn.bfloat16, layout=ttnn.TILE_LAYOUT)
    for _ in range(2)
]
guarded_nested_pop(*tensors)
print("COMPILED")
python repro_guarded_nested_pop.py

Observed on main 546f6b20a: the program compiles to

if (node_x == 0) {
  cb_ctarg_1.wait_front(1);
  cb_ctarg_1.pop_front(1);        // inserted
  if (node_y == 0) {
    cb_ctarg_1.pop_front(1);      // the program's pop
  } else {
    cb_ctarg_1.pop_front(1);      // the program's pop
  }
  for (...) { cb_ctarg_1.wait_front(1); cb_ctarg_1.pop_front(1); }
}

With #1085 the DFB lifecycle verifier rejects it instead ("logical DFB 0 has an incomplete or misordered consumer lifecycle ... the consumer kernel performs a pop before a matching wait"). The same program without the outer if node_x == 0: compiles correctly on both.

Expected: no inserted pop, since every path through the nested if pops the block; the program compiles to one wait and one pop per path. Related forms from the #1108 review (probes G5, G6, G13, G14, D12, D13): one branch pops and the other does not, or the nested if is followed by direct waits.

Found by the review of #1108 (finding S5-3).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    DFBbugSomething isn't workingcompilerMLIR analysis, verification, transformation, and lowering after the Python frontend

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions