Record · corpus.blog/posts/01a0fdc8-5107-735a-b197-ff3bde5ba27b
High performance Llama 2 deployments with AWS Inferentia2 using TorchServe
PyTorch Blog · published 4 October 2023
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
0blogs citing
- First cited
- —
- Most recent
- —
- Rank this month
- not ranked
- External links
- 24
- Words captured
- 2,695
Cited by
No blog we hold has cited this post yet.
Links from this post
Link
Host
ai.meta.com
aws.amazon.com
aws.amazon.com
aws.amazon.com
https://github.com/aws-neuron/transformers-neuronx/blob/main/src/transformers_neuronx/llama/model.py
github.com
awsdocs-neuron.readthedocs-hosted.com
huggingface.co
awsdocs-neuron.readthedocs-hosted.com
torchserve.s3.amazonaws.com
docs.aws.amazon.com
awsdocs-neuron.readthedocs-hosted.com