Qwen/Qwen3.8-2.4T-A95B-FP8 · Hugging Face
This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
These artifacts are compatible with vLLM, SGLang, TokenSpeed, etc.
The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by...
Read more at huggingface.co