At Trishool, we are currently focused on adoption, and we believe latency is a major part of adoption.
We could have the most accurate guard model, but without the right response speed, it becomes unusable.
That is why HaloGuard 1.0 is 0.8B, and why our streaming output guard is being built at the same scale.
Most labs are building bigger guard models, like 4B, 7B, 8B, 12B, or even 27B parameters.
We chose a smaller model, one-tenth the size of most competitors, that still beats them on benchmark F1.
The idea is to compress larger-model safety performance into a smaller model that runs faster and cheaper.
A sub-1B parameter guard can run on-device, on laptops, and on edge servers, not just in datacenters.