v1.3.0rc29
Highlights Model Support Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519 Optimize Gemma4 vision rotary embeddings without a GEMM operation #19571 Enable the CuTeDSL MoE backend for Kimi K3 NVFP4 SiTU #19003 Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151 Extend MiniMax-M3 piecewise CUDA graphs through model and executor…