China Telecom AI's Xing4.0-29B-A4B (29B MoE, 4B active) can start on a single consumer GPU. For operators, the bottleneck moves to serving: VRAM fit after quantization, KV cache growth, p95 latency, tenancy and chargeback.
clastiq.ai/insights/xing4-se…