Pinned Loading
-
io-aware-attention-trainium
io-aware-attention-trainium PublicA memory and IO aware implementation of FlashAttention on AWS Trainium, investigating how streaming softmax, tiling, and multi-island execution adapt to non GPU accelerator constraints.
Python
-
KITE-Kernel-Intelligence
KITE-Kernel-Intelligence PublicKITE: Kernel Intelligence-per-watt Tree Explorer
Python 1
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

