Record · corpus.blog/posts/01a0fd1b-a95e-7184-b1fa-e78ffc317894
INT4 Decoding GQA CUDA Optimizations for LLM Inference
PyTorch Blog · 6 June 2024 · 23 min read
A post on pytorch.org, published 6 June 2024, has not been cited by any blog yet (corpus.blog, measured 9 October 2026).
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
0blogs citing
| Period | Blogs | Links |
|---|---|---|
| All time | 0 | 0 |
| Last 90 days | 0 | 0 |
| Last 30 days | 0 | 0 |
| Last 7 days | 0 | 0 |
| Anchor phrases | 0 |
|---|---|
| External links | 12 |
| Words captured | 5,159 |
Cited by
No blog we hold has cited this post yet.
Links from this post
Link
Host
llama.meta.com
openai.com
https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/a100-80gb-datasheet-update-nvidia-us-1521051-r2-web.pdf Who cites this page
nvidia.com
https://github.com/facebookresearch/xformers/blob/9f6abadabdec17cd4b5c301632a44bf8216a7f35/xformers/csrc/attention/cuda/fmha/autogen/impl/cutlassF_bf16_aligned.cu Who cites this page
github.com
https://bruce-lee-ly.medium.com/nvidia-tensor-core-introduction-to-wmma-api-programming-21bcfee4ec45
bruce-lee-ly.medium.com
docs.nvidia.com
internalfb.com
Elsewhere on pytorch.org (1)
Something wrong here? Report a problem