MiniCPM5

MiniCPM5-2B is an open source model from OpenBMB. It uses the Llama architecture, so Axolotl’s Llama optimizations (flash attention, sample packing, Cut Cross Entropy, Liger) apply.

This guide shows how to fine-tune it with Axolotl with multi-turn conversations and proper masking.

Getting started

  1. Install Axolotl following the installation guide.

  2. Install Cut Cross Entropy to reduce training VRAM usage.

  3. Run the finetuning example:

    axolotl train examples/minicpm5/lora-2b.yaml
    axolotl train examples/minicpm5/fft-2b.yaml

Let us know how it goes. Happy finetuning! 🚀

Turn terminator

The tokenizer’s eos_token is </s>, but the chat template ends each turn with <|im_end|>. Both configs set eot_tokens so <|im_end|> is trained and the model learns to stop:

chat_template: tokenizer_default
eot_tokens:
  - "<|im_end|>"

TIPS

  • The chat template adds an empty <think> block to assistant turns without reasoning_content; it is masked out of the loss. Verify with axolotl preprocess examples/minicpm5/lora-2b.yaml --debug.
  • To train on reasoning traces, put them in reasoning_content on the assistant message (docs).
  • Read more on how to load your own dataset at docs.
  • The dataset format follows the OpenAI Messages format as seen here.

Optimization Guides

Please check the Optimizations doc.