Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM
· By Antonio Sedino, CTRO · Published by Reinventy Solutions Corp.
Walkthrough covers cluster provisioning, NVFP4 quantization, OpenAI-compatible endpoint with reasoning, tool calling and speculative decoding.
The AWS Machine Learning Blog explains how to deploy the 2.4-trillion-parameter open-weight model Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod using vLLM. The guide covers cluster provisioning, NVFP4 quantization, and setting up an OpenAI-compatible endpoint with built-in reasoning, tool calling and native MTP speculative decoding.
