triton.experimental.gluon.language.amd.cdna4.async_copy

Functions

global_load_to_shared

AMD global load to shared operation.

buffer_load_to_shared

AMD buffer load to shared operation.

commit_group

Commit oustanding async operations.

wait_group

Wait for outstanding commit groups.

load_shared_relaxed

Load a tensor from shared memory with extra hints for the underlying compiler to avoid emitting unnecessary waits before loading from the target shared memory.