Nex-N2.5

Nex-N2.5-mini is an open source model from Nex AGI built on the Qwen3.5-35B-A3B architecture (hybrid Gated DeltaNet + attention MoE with vision). Everything in examples/qwen3.5 applies here; only the chat template differs.

Getting started

  1. Install Axolotl following the installation guide.

  2. Install Cut Cross Entropy to reduce training VRAM usage.

  3. Install FLA for sample packing support with the Gated DeltaNet linear attention layers:

    uv pip uninstall causal-conv1d && uv pip install flash-linear-attention==0.4.1
  4. Run the finetuning example:

    axolotl train examples/nex-n2.5/mini-qlora.yaml
    axolotl train examples/nex-n2.5/mini-vision-lora.yaml

Let us know how it goes. Happy finetuning! 🚀

Chat template

Use chat_template: tokenizer_default. The template renders a <think> block on every assistant turn. In text configs, empty <think> blocks are masked out and only the assistant content and <|im_end|> are trained. In multimodal configs (mini-vision-lora.yaml), masking is per turn, so the empty <think> block is trained along with each assistant reply.

TIPS

  • For LoRA targets on the DeltaNet layers and routed/shared experts, see the Qwen3.5 README.
  • Read more on how to load your own dataset at docs.
  • The dataset format follows the OpenAI Messages format as seen here.
  • For multimodal finetuning, set processor_type: AutoProcessor, skip_prepare_dataset: true, and remove_unused_columns: false as shown in mini-vision-lora.yaml.

Optimization Guides

Please check the Optimizations doc.