本来叫 nano 的,后来发现装不下 Qwen3.5,就改名叫 big 了
-
Updated
May 6, 2026 - Python
本来叫 nano 的,后来发现装不下 Qwen3.5,就改名叫 big 了
French course summary of “Fast & Efficient LLM Inference with vLLM,” covering inference, quantization, LLM Compressor, PagedAttention, Continuous Batching, benchmarking, and evaluation.
Using vLLM compressor to quantize Qwen3-0.6B using GPTQ to W4A16
NVFP4 and W4A16 quantization on consumer Blackwell (sm_120): recipes, an eval harness that detects the damage aggregate benchmarks hide, reproducible numbers.
Add a description, image, and links to the llm-compressor topic page so that developers can more easily learn about it.
To associate your repository with the llm-compressor topic, visit your repo's landing page and select "manage topics."