09-092026大模型技术FlashAttention:从 IO-Aware Attention 到 GPU Work Partitioning剖析 FA 与 FA-2 两篇论文中的 IO 优化、Softmax 与 GPU 并行设计。