Introducing celeris-1.
A general purpose language model which delivers near-GPT-5 level intelligence with 15x faster response times.
We’ve developed a new inference architecture that uses diffusion techniques instead of conventional autoregressive generation - unlocking dramatically better speed while maintaining frontier-level intelligence.
Celeris-1 delivers a p50 response latency of 157ms - around 15× faster than GPT-5-mini and 17× faster than GPT-5 - while scoring comparably 76% on MMLU-Pro, compared with 78% and 81%, respectively.
On tokens per second, Celeris-1 achieves a throughput of 1,280 tokens per second vs 144 for gemini-3.5 flash-light using a reconstructed version of
@ArtificialAnlys's tokens per second benchmark dataset.
The model is available starting today. Sign up at
celeris.ai