Moderation runs in front of everything your app serves.
Every message, every generation, every time. That's a lot of volume to be paying frontier prices for. So we built a new SLM for moderation, designed to run at the edge, using less compute to get more done.
We benchmarked our new zlm-v1-moderation-edge model head-to-head against OpenAI's omni-moderation-latest. Take a look at the results.
The takeaway? Save your frontier models for tasks that require deep reasoning. Use ZeroGPU for the repeatable work that runs a million times a day.
Full benchmark + docs linked in the comments ⬇️