NVIDIA Rubin
Cluster Launch Control (CLC) for Blackwell (SM100+) dynamic persistent kernels. |
|
Allocate tensor memory. |
|
Store a tensor to shared memory asynchronously and signal an mbarrier on completion. |
|
Issue a fence to complete asynchronous shared memory operations. |
|
Represents a tensor memory descriptor handle for Tensor Core Gen5 operations. |
|
This instruction causes the provided mbarrier to be arrived-on with a count of 1 when all async tcgen05 MMA and copy instructions previously issued by the thread are complete. |
|
Start an asynchronous copy from shared memory to tensor memory. |
|
Emit a 5th generation TensorCore MMA instruction. |
|
Calculate the number of CTAs that will commit the tcgen05 MMA instruction. |
|
Emit a 5th generation TensorCore MMA scaled instruction. |
|
Describes the layout for tensor memory in Blackwell architecture. |
|
Describes the layout for tensor memory scales in Rubin architecture. |