Skip to content

Reduce Triton TBE unweighted backward reduction width (#4488) - #4488

Closed
stashuk-olek wants to merge 5 commits into
meta-pytorch:mainfrom
stashuk-olek:export-D114403703
Closed

Reduce Triton TBE unweighted backward reduction width (#4488)#4488
stashuk-olek wants to merge 5 commits into
meta-pytorch:mainfrom
stashuk-olek:export-D114403703

Conversation

@stashuk-olek

@stashuk-olek stashuk-olek commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 5, 2026
@meta-codesync

meta-codesync Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

@stashuk-olek has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114403703.

stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
@meta-codesync meta-codesync Bot changed the title Reduce Triton TBE unweighted backward reduction width Reduce Triton TBE unweighted backward reduction width (#4488) Aug 5, 2026
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 5, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
@stashuk-olek
stashuk-olek force-pushed the export-D114403703 branch 2 times, most recently from 9db3ae0 to bdb975b Compare August 6, 2026 23:03
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 6, 2026
)

Summary:

We could drop that, small tuning on buffer_size.

Reviewed By: axeisghost

Differential Revision: D114403703
@meta-codesync meta-codesync Bot changed the title Reduce Triton TBE unweighted backward reduction width (#4488) Reduce Triton TBE unweighted backward reduction width Aug 7, 2026
Summary:

Use BoundsCheckMode.V2 for Triton TBE standalone index validation. This preserves the existing warning and offset-repair behavior while reducing validation overhead before Triton forward.

Reviewed By: axeisghost, TroyGarden

Differential Revision: D114279989
Summary:

Validate and repair indices inside supported unweighted, non-VBE Triton forward kernels. Unsupported configurations retain the CUDA v2 fallback, while fused checked loads preserve warning accumulation without a separate full-index pass.

Reviewed By: TroyGarden, axeisghost

Differential Revision: D114284802
)

Summary:

Convert long bags over small tables into per-bag row histograms followed by a tensor-core dot. Generic row, dimension, and bag-size predicates select the route, and checked loads preserve fused bounds behavior.

Reviewed By: axeisghost

Differential Revision: D114273027
Summary:

Restructure short- and long-run gradient accumulation to reduce serialized reduction pressure. The new reduction schedule exposes more independent work while preserving the existing exact rowwise optimizer update.

Reviewed By: axeisghost

Differential Revision: D114273026
)

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703
@meta-codesync meta-codesync Bot changed the title Reduce Triton TBE unweighted backward reduction width Reduce Triton TBE unweighted backward reduction width (#4488) Aug 8, 2026
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 8, 2026
)

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 8, 2026
)

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 8, 2026
)

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703
stashuk-olek added a commit to stashuk-olek/torchrec that referenced this pull request Aug 8, 2026
)

Summary:

Reduce the dout reduction tile from 16 to 8 in the unweighted short-run, separate long-run accumulation, and fused long-run kernels. The smaller tile reduces live vector state and register pressure on B200 while retaining FP32 accumulation and the existing rowwise optimizer update. Weighted kernels remain unchanged.

Reviewed By: axeisghost

Differential Revision: D114403703
@meta-codesync meta-codesync Bot closed this in 50809e4 Aug 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant