Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
· By Antonio Sedino, CTRO · Published by Reinventy Solutions Corp.
Amazon SageMaker HyperPod now supports model caching for inference, pre-loading model weights and container images onto cluster nodes for local NVMe storage access.
Amazon SageMaker HyperPod now supports model caching for inference, pre-loading model weights and container images onto cluster nodes to enable pods to read from local NVMe storage instead of downloading over the network.
