Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
Posted by matt_d 6 days ago
Comments
Comment by touisteur 2 days ago
Comment by saagarjha 1 day ago
Comment by broadsidepicnic 2 days ago
Comment by touisteur 2 days ago
Comment by lostmsu 2 days ago
Comment by tpurves 2 days ago
Comment by lostmsu 1 day ago
Unless MI350P will be much cheaper than 6000 Pro, which is unlikely given its VRAM, it will lose to it, because you'd need 2 of either.
Comment by touisteur 1 day ago
I was remarking on the MI350P because I've had a hard time procuring "small" CDNAx systems (for e.g. development, experiments and lower-profile servers) and OAM seemed very niche (not if you're aiming for density and training/inference...).
I hope the MI350P fills a lower part of the spectrum and I can start massively porting CUDA stuff or at least work on HIP and ROCm and what I need to make most or some of our CUDA stuff run on AMD HW, then how make it run fast.
Comment by boroboro4 2 days ago
Comment by smallerize 2 days ago
Comment by touisteur 1 day ago
Same for gpudirect (more useful for scale-out or training).
Everyone is focused on AI but these are two interesting techs trickling down from the NVIDIA tree, relatively "easy" to use there, which maybe exist in AMD world but I somehow missed the docs and APIs and demos on how to use them...