Rework is inevitable in large-scale LLM pre-training and fine-tuning. What it costs isn't.
Introducing xLLM, an efficient and flexible infrastructure for pre-training and fine-tuning dense and MoE LLMs. It keeps key training decisions changeable without giving up throughput.
Efficient, at 6,295 tokens/sec per GPU on K2-Horizon-MoVA-36B-A4B and 10,050 tokens/sec per GPU on Llama3-8B on H200s.
Flexible, because the tokenizer, data mixture, model architecture, and training stages can change without rebuilding the dataset or the system around them.
xLLM ships with the checkpoints, training logs, and recipes behind K2 Horizon:
github.com/ifm-ai/xllm
K2-Horizon-MoVA-36B-A4B:
huggingface.co/IFM/K2-Horizo…