Reduction ops for TMA shared-to-global bulk tensor copies.
Declared in <llvm/IR/NVVMIntrinsicUtils.h>
enum class TMAReductionOp : uint8_t;
These map to the cp.reduce.async.bulk.tensor.* family of PTX instructions.
| Name | Description |
|---|---|
ADD | Element-wise add reduction. |
MIN | Element-wise minimum reduction. |
MAX | Element-wise maximum reduction. |
INC | Saturating increment reduction. |
DEC | Saturating decrement reduction. |
AND | Bitwise AND reduction. |
OR | Bitwise OR reduction. |
XOR | Bitwise XOR reduction. |
| Name | Description |
|---|---|
getTMATensorReductionOpName | Return the PTX name for a TMA tensor reduction operation. |