大模型技术
FlashAttention:从 IO-Aware Attention 到 GPU Work Partitioning
剖析 FA 与 FA-2 两篇论文中的 IO 优化、Softmax 与 GPU 并行设计。
标签
共 1 篇文章
剖析 FA 与 FA-2 两篇论文中的 IO 优化、Softmax 与 GPU 并行设计。
Triton 实现 Softmax Attention 算子的基本思路。重点思考怎样不用存完整 attention matrix,仍然精确算出 softmax attention.
全站检索
输入关键词开始搜索