接入 API · 个人 AI 解读连接自己的模型解读资讯,浏览新闻无需配置。
返回资讯列表
算力与芯片官方来源国际

Triton 编译器发布: gfx950-tutorial-v2.0

Triton 编译器发布 · 来源更新
今日摘要使用自己的 API,仅供个人查看

来源摘要

Adds the two compiler changes the Gluon Flash Attention kernels need, on top of gfx950-tutorial-v1.1. The LLVM pin is unchanged (850a2b1), so the out-of-tree LLIR-scheduler plugin does NOT need rebuilding. [Gluon] gl.warp_predicate — a per-wave masked-skip region lowering to s_and_saveexec + s_cbranch_execz, with no cross-wave reduction and no barrier. This is what lets a wavefront skip the flash-attention accumulator rescale when none of its own rows advanced their running max. [AMD] warp-pipeline barriers — always emit LOCAL cluster barriers rather than the minimal set, place the loop-carried wrap-around barrier at the top of the loop body, and force a hard LDS drain there with amdgpu.memory_counter_wait(ds = 0). Together these keep s_waitcnt lgkmcnt out of the MFMA stages. Verified against the branch these were developed on: both attention kernels and all four inter_wave GEMM kernels (a16w16 fp16/bf16, a8w8, a4w4) compile to byte-identical assembly and pass their correctness checks.

阅读原始来源
来源
Triton 编译器发布 · 官方来源
来源更新
2026/07/29 16:47
首次采集
2026/09/19 12:56

本文为公开信息索引与摘要,详情及后续变化请以原始来源为准。

把 AI 雷达放到桌面

在支持安装的浏览器中,可以将本站作为应用打开。

安装入口取决于浏览器;应用和网站使用同一份最新内容。

查看完整安装指南