Qwen 3.8 Flash Next
Qwen3.8-Flash-Next is a 176.94B multimodal MoE model (512 routed experts, 10 active, plus a shared expert) built on a hybrid attention stack of 36 Gated DeltaNet layers and 12 Qwen Sparse Attention layers, with a 51.2B-parameter n-gram PLE embedding. It loads as model_type: qwen4_exp, a different architecture from Qwen3.8-27B, which is qwen3_5.
This guide shows how to fine-tune it with Axolotl with multi-turn conversations and proper masking.
Getting started
Install Axolotl following the installation guide.
Install Cut Cross Entropy to reduce training VRAM usage.
Install
torchvision. The model loads a video processor even when your dataset is text-only.uv pip install torchvisionRun the finetuning example:
# QLoRA (1x B300 @ ~120 GiB w offload, else ~216 GiB) axolotl train examples/qwen3.8-flash-next/qlora.yaml # Vision + text QLoRA (1x B300 @ ~110 GiB w offload, else ~207 GiB) axolotl train examples/qwen3.8-flash-next/vision-qlora.yaml# NVFP4 MoE-LoRA (1x B300 @ ~223 GiB) axolotl train examples/qwen3.8-flash-next/nvfp4-lora.yaml # bake the adapter back into a plain NVFP4 checkpoint. --lora-model-dir defaults to # output_dir, so pass it explicitly when the adapter lives elsewhere axolotl merge-lora examples/qwen3.8-flash-next/nvfp4-lora.yaml \ --lora-model-dir ./outputs/qwen3.8-flash-next-nvfp4-lora
Let us know how it goes. Happy finetuning! 🚀
Gated DeltaNet Linear Attention
36 of the 48 layers are Gated DeltaNet. Its projections are split rather than fused as in Qwen3-Next, so target them via:
lora_target_modules:
- linear_attn.in_proj_qkv
- linear_attn.in_proj_z
- linear_attn.in_proj_b
- linear_attn.in_proj_a
- linear_attn.out_projLimitations
| Feature | Status |
|---|---|
attn_implementation |
sdpa only. |
lora_target_linear |
Incompatible. It expands to the QSA indexer projection incorrectly. |
| LoRA kernels | Unsupported |
| Liger | Unsupported due to incompatible kernels |
sdpa_varlen |
Unsupported |
| Full finetuning | Untested. The bf16 weights alone are 329.6 GiB. |