Latest mlx-lm is out:
- New models: Kimi K2.5, Step3.5 flash, LongCat Flash lite thanks to
@kernelpool
- Support for distributed inference with mlx_lm.server thanks to
@angeloskath
- Much faster and more memory efficient DeepSeek v3 (and other MLA-based models)