AFAIK AC2 Is the only platform with support for full weight fine tuning of Kimi K3.
For short, simple tasks LoRA fine tuning may be similar. But for multi turn, agentic workloads, full weight training is strictly superior. dm if you’re interested in trying AC2
Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%.
At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to training frontier-scale open models.