Nex-N2.5
Nex-N2.5-mini is an open source model from Nex AGI built on the Qwen3.5-35B-A3B architecture (hybrid Gated DeltaNet + attention MoE with vision). Everything in examples/qwen3.5 applies here; only the chat template differs.
Getting started
Install Axolotl following the installation guide.
Install Cut Cross Entropy to reduce training VRAM usage.
Install FLA for sample packing support with the Gated DeltaNet linear attention layers:
uv pip uninstall causal-conv1d && uv pip install flash-linear-attention==0.4.1Run the finetuning example:
axolotl train examples/nex-n2.5/mini-qlora.yaml axolotl train examples/nex-n2.5/mini-vision-lora.yaml
Let us know how it goes. Happy finetuning! 🚀
Chat template
Use chat_template: tokenizer_default. The template renders a <think> block on every assistant turn. In text configs, empty <think> blocks are masked out and only the assistant content and <|im_end|> are trained. In multimodal configs (mini-vision-lora.yaml), masking is per turn, so the empty <think> block is trained along with each assistant reply.
TIPS
- For LoRA targets on the DeltaNet layers and routed/shared experts, see the Qwen3.5 README.
- Read more on how to load your own dataset at docs.
- The dataset format follows the OpenAI Messages format as seen here.
- For multimodal finetuning, set
processor_type: AutoProcessor,skip_prepare_dataset: true, andremove_unused_columns: falseas shown inmini-vision-lora.yaml.
Optimization Guides
Please check the Optimizations doc.