NVIDIA Blackwell

async_copy

clc

Cluster Launch Control (CLC) for Blackwell (SM100+) dynamic persistent kernels.

mbarrier

tma

add2

Add two tensors using a native two-lane packed instruction.

allocate_tensor_memory

Allocate tensor memory.

fence_async_shared

Order generic-proxy and asynchronous-proxy shared memory accesses.

fma2

Perform a native two-lane packed fused multiply-add.

max2

Select the maximum with a native packed half-precision instruction.

min2

Select the minimum with a native packed half-precision instruction.

mma_v2

mul2

Multiply two tensors using a native two-lane packed instruction.

sub2

Subtract two tensors using a native two-lane packed instruction.

tensor_memory_descriptor

Represents a tensor memory descriptor handle for Tensor Core Gen5 operations.

tensor_memory_descriptor_type

TensorMemoryLayout

Describes the layout for tensor memory in Blackwell architecture.

TensorMemoryScalesLayout

Describes the layout for tensor memory scales in Blackwell architecture.