Training MoVA produces fewer and smaller extreme gradient spikes in a 36B-A4B comparison against standard MoE, with gated attention disabled in both. Across 45,188 matched training steps after warmup and early restarts, MoVA triggered gradient clipping 26 times against 52 for standard MoE, a 50% reduction. Its 99.9th-percentile gradient norm was 0.42 against 1.24, and its largest spike reached 64.9 against 190.9.