NVIDIA Blackwell
Cluster Launch Control (CLC) for Blackwell (SM100+) dynamic persistent kernels. |
|
Add two tensors using a native two-lane packed instruction. |
|
Allocate tensor memory. |
|
Order generic-proxy and asynchronous-proxy shared memory accesses. |
|
Perform a native two-lane packed fused multiply-add. |
|
Select the maximum with a native packed half-precision instruction. |
|
Select the minimum with a native packed half-precision instruction. |
|
Multiply two tensors using a native two-lane packed instruction. |
|
Subtract two tensors using a native two-lane packed instruction. |
|
Represents a tensor memory descriptor handle for Tensor Core Gen5 operations. |
|
Describes the layout for tensor memory in Blackwell architecture. |
|
Describes the layout for tensor memory scales in Blackwell architecture. |