AMD CDNA 4
AMD buffer load from global memory via a scalar base pointer and a tensor of offsets instead of a tensor of pointers. |
|
AMD buffer store a tensor directly to global memory via a scalar base pointer and a tensor of offsets instead of a tensor of pointers. |
|
Compute an efficient padded shared layout that avoids bank conflicts. |
|
Get the scale layout for MFMA scaled operands. |
|
Load M/N-packed fp4 bytes from shared memory into a K-packed MFMA dot operand layout. |
|
Computes matrix multiplication |
|
AMD Scaled MFMA operation. |
|
Scale and convert FP16, BF16, or FP32 values to a low-precision MX format, dividing by the raw E8M0 |
|
Upcast an fp4 or fp8 tensor and fold raw E8M0 scale payload into the CDNA4 scaled-upcast op. |