Two new small language models are live on ZeroGPU: embeddings, to power semantic search RAG, clustering, and deduplication.
This is the model layer that lets your app match things by meaning instead of by keyword: A user types "cancel my plan." Your help center article says "end your subscription." An embedding model ensures the right recommendation is returned.
More than 40% of these kidns of enterprise AI tasks can run for less by switching to our SLMs.