Record · corpus.blog/posts/01a0fdec-2f34-70ca-8cf3-fc4523aa4866
The Path to Achieve Ultra-Low Inference Latency With LLaMA 65B on PyTorch/XLA
PyTorch Blog · published 28 June 2023
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
0blogs citing
- First cited
- —
- Most recent
- —
- Rank this month
- not ranked
- External links
- 22
- Words captured
- 2,347
Cited by
No blog we hold has cited this post yet.
Links from this post
Link
Host
arxiv.org
ai.facebook.com
openai.com
arxiv.org
arxiv.org
https://github.com/huggingface/transformers/blob/v4.27.2/src/transformers/models/opt/modeling_opt.py
github.com
storage.googleapis.com
storage.googleapis.com
arxiv.org
arxiv.org
arxiv.org
huggingface.co
github.com