框架接入¶
AscendNPU IR支持框架(PyTorch/TensorFlow/MindSpore)接入,有两种方式:
DSL接入方式:通过Triton、TileLang等领域特定语言接入,将算子编译为AscendNPU IR。
IR接入方式:通过IR表示接入,支持Torch IR、Linalg/HFusion IR、HIVM IR多层级接入,支持自动算子融合和切分,生成昇腾亲和的高性能算子。
DSL接入方式¶
AscendNPU IR向上支持与Triton、TileLang等语言或框架的对接,使能三方DSL支持昇腾硬件,在NPU上运行自定义算子。
接入方式 |
说明 |
|---|---|
使用Triton编写高性能内核,通过Triton Ascend在昇腾NPU上运行。含安装、环境、算子映射及昇腾扩展说明。 |
|
使用TileLang Ascend(基于tile-lang/TVM的DSL)开发面向昇腾NPU的内核(如GEMM、向量运算、attention)。含环境、构建与快速开始。 |
IR接入方式¶
AscendNPU IR支持多层级IR接入,不同层级在抽象程度和控制粒度上有所差异(详见IR接入简介-多级IR抽象架构):
Torch IR:框架层ATen算子,经Pass转换为Linalg/HFusion。
Linalg/HFusion IR:通用张量代数层与硬件感知融合层,标准MLIR dialect表达算子语义,HFusion自动完成融合、切分和调度。
HIVM IR:NPU指令层,直接映射硬件指令,显式控制存储层级(GM/UB/L1/L0)和计算流水线(Vector/Cube/MTE),支持精细粒度调优。
Torch IR接入¶
直接使用Torch dialect的ATen算子,通过convert-torch-to-hfusion等Pass自动转换为Linalg/HFusion Named Op,再进入自动融合和调度流程。
Torch至AscendNPU IR转换流程¶
Torch IR通过torch-backend-to-named-op-backend-pipeline转换流水线接入AscendNPU IR。BiShengIR自定义的convert-torch-to-hfusion Pass优先将Torch ATen算子转换为Linalg/HFusion Named Op,未覆盖的算子回退到上游torch-mlir的标准lowering通路。主要转换阶段如下:
convert-torch-to-hfusion:BiShengIR自定义转换,覆盖55+个ATen算子到Linalg/HFusion Named Op。convert-torch-to-linalg:上游torch-mlir转换,处理剩余算子。convert-torch-to-scf/arith/tensor:上游torch-mlir完成控制流、算术、tensor等转换。func-backend-type-conversion:将Torch类型(!torch.vtensor)转换为标准builtin类型(tensor)。
用例torch.mlir:
func.func @torch_mul(%arg0: !torch.vtensor<[4096],f16>, %arg1: !torch.vtensor<[1,56,4096],f16>) -> !torch.vtensor<[1,56,4096],f16>
attributes {hacc.entry, hacc.function_kind = #hacc.function_kind<DEVICE>} {
%0 = torch.aten.mul.Tensor %arg0, %arg1 : !torch.vtensor<[4096],f16>, !torch.vtensor<[1,56,4096],f16> -> !torch.vtensor<[1,56,4096],f16>
return %0 : !torch.vtensor<[1,56,4096],f16>
}
调用方式:有两种方式,二者共享同一套编译pipeline。
分步转换:先将Torch IR转换为Linalg/HFusion IR,适用于需要缓存或查看中间IR的场景。转换完成后,可将
torch_to_hfusion.mlir作为输入,按Linalg/HFusion IR接入流程继续编译生成二进制。命令:
bishengir-opt -torch-backend-to-named-op-backend-pipeline torch.mlir -o torch_to_hfusion.mlir预期产物:MLIR文本文件(
.mlir格式),内容为转换后的Linalg/HFusion IR。例如:
func.func @torch.aten.mul_tensor(%arg0: tensor<4096xf16>, %arg1: tensor<1x56x4096xf16>) -> tensor<1x56x4096xf16> attributes {hacc.entry, hacc.function_kind = #hacc.function_kind<DEVICE>} {
%0 = tensor.empty() : tensor<1x56x4096xf16>
%broadcasted = linalg.broadcast ins(%arg0 : tensor<4096xf16>) outs(%0 : tensor<1x56x4096xf16>) dimensions = [0, 1]
%1 = linalg.elemwise_binary {fun = #linalg.binary_fn<mul>} ins(%broadcasted, %arg1 : tensor<1x56x4096xf16>, tensor<1x56x4096xf16>) outs(%0 : tensor<1x56x4096xf16>) -> tensor<1x56x4096xf16>
return %1 : tensor<1x56x4096xf16>
}
端到端编译:使用
bishengir-compile直接将Torch IR编译为可执行二进制,完整经过Torch → HFusion → HIVM IR编译pipeline。命令:
bishengir-compile -enable-torch-compile=true -enable-hfusion-compile=true -enable-hivm-compile=true -target=Ascend910B1 torch.mlir -o torch_kernel.o预期产物:Ascend NPU算子二进制文件(
.o格式),可与CANN runtime配合在设备端运行。
支持的Torch算子¶
Elementwise Binary¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Elementwise Unary¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
分解为 |
|
分解为 |
Compare¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
Reduction¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Data Movement¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
其他¶
Torch Op |
转换目标 |
|---|---|
|
|
|
|
|
|
Linalg/HFusion IR接入¶
使用Linalg/Tensor、HFusion等标准MLIR dialect表达算子语义,直接进入Linalg/HFusion IR层级的自动融合和调度流程。
用例hfusion.mlir:
func.func @hfusion_reduce_mul(%arg0: tensor<40960xf32>, %arg1: tensor<40960x1024xf32>, %arg2: tensor<40960x1024xf32>, %arg3: tensor<40960x1024xf32>) -> tensor<40960xf32>
attributes {hacc.entry, hacc.function_kind = #hacc.function_kind<DEVICE>} {
%1 = tensor.empty() : tensor<40960x1024xf32>
%3 = linalg.elemwise_binary {fun = #linalg.binary_fn<mul>} ins(%arg1, %arg2 : tensor<40960x1024xf32>, tensor<40960x1024xf32>) outs(%arg3: tensor<40960x1024xf32>) -> tensor<40960x1024xf32>
%4 = tensor.empty() : tensor<40960xf32>
%sum = linalg.reduce {arith.addf} ins(%3 : tensor<40960x1024xf32>)
outs(%4 : tensor<40960xf32>) dimensions = [1]
%5 = tensor.empty() : tensor<40960xf32>
%6 = linalg.elemwise_binary {fun = #linalg.binary_fn<mul>} ins(%arg0, %sum : tensor<40960xf32>, tensor<40960xf32>)
outs(%5: tensor<40960xf32>) -> tensor<40960xf32>
return %6 : tensor<40960xf32>
}
调用方式:
命令:
bishengir-compile -enable-hfusion-compile=true -enable-hivm-compile=true -target=Ascend910B1 hfusion.mlir -o hfusion_kernel.o预期产物:Ascend NPU算子二进制文件(
.o格式),可与CANN runtime配合在设备端运行。
自动融合:
Linalg/HFusion IR接入后,HFusion编译流程会对符合融合条件的算子执行自动融合与调度:将多个算子合并进同一kernel执行,使中间结果在片上内存中复用,减少全局内存读写;并根据融合模式与算子特征,自动选择Tiling方案和调度策略,生成面向Ascend NPU的高效执行schedule。融合后的IR经Tiling、循环生成、Transform Dialect应用等步骤,最终下推至HIVM并生成可执行二进制。
支持的Op类型包括:
ElemwiseBroadcastReduceTransposeConcat
关于自动融合的算法原理、约束能力、架构设计等详细说明,请参阅HFusion AutoSchedule自动融合与调度。
HIVM IR接入¶
对于需要精细控制硬件行为的场景,可以直接使用HIVM dialect编写kernel,显式管理存储层级和计算流水线。
用例hivm.mlir:
func.func @hivm_vadd(%valueA: memref<16xf16, #hivm.address_space<gm>>,
%valueB: memref<16xf16, #hivm.address_space<gm>>,
%valueC: memref<16xf16, #hivm.address_space<gm>>)
attributes {hacc.entry, hacc.function_kind = #hacc.function_kind<DEVICE>} {
%ubA = memref.alloc() : memref<16xf16, #hivm.address_space<ub>>
hivm.hir.load ins(%valueA : memref<16xf16, #hivm.address_space<gm>>)
outs(%ubA : memref<16xf16, #hivm.address_space<ub>>)
%ubB = memref.alloc() : memref<16xf16, #hivm.address_space<ub>>
hivm.hir.load ins(%valueB : memref<16xf16, #hivm.address_space<gm>>)
outs(%ubB : memref<16xf16, #hivm.address_space<ub>>)
%ubC = memref.alloc() : memref<16xf16, #hivm.address_space<ub>>
hivm.hir.vadd ins(%ubA, %ubB : memref<16xf16, #hivm.address_space<ub>>,
memref<16xf16, #hivm.address_space<ub>>)
outs(%ubC : memref<16xf16, #hivm.address_space<ub>>)
hivm.hir.store ins(%ubC : memref<16xf16, #hivm.address_space<ub>>)
outs(%valueC : memref<16xf16, #hivm.address_space<gm>>)
return
}
HIVM层使用#hivm.address_space标注存储层级:gm(Global Memory)、ub(Unified Buffer)、l1(L1 Buffer)、l0a/l0b/l0c(L0 Buffer)。通过hivm.hir.load/hivm.hir.store进行显式DMA搬运,通过hivm.hir.vadd等指令在片上完成计算。
调用方式:HIVM层无需使能HFusion编译流程,默认的HIVM编译流程会完成同步插入、内存规划等优化。
命令:
bishengir-compile -enable-hfusion-compile=false -enable-hivm-compile=true -target=Ascend910B1 hivm.mlir -o hivm_kernel.o预期产物:Ascend NPU算子二进制文件(
.o格式),可与CANN runtime配合在设备端运行。
关于IR层概念、公共编译选项及其他接入路径(如Triton、TileLang),请参阅IR接入简介。