I asked
@grok if a normal LLM can't just do this.
Response below 👇
Grok, GPT, and Claude can look at an image and say what’s in it. They are general. They are also big, slow, and expensive if you run them on every frame of many cameras.
Score’s pitch is the opposite: tiny specialist models (often under 30MB) that do one job well — detect, track, count — fast enough for real cameras, including on CPU. They claim those models can beat or match frontier VLMs on specific vision tests while costing far less per frame.
Use an LLM if you want a description, a one-off answer, or messy reasoning about a scene.
Use Score if you want a camera system that runs all day without burning API credits.