We ran 6,120 tests over five weeks to see how far we could push Needle 2.
Here's a complete playbook for getting a 14MB function-calling LLM to production accuracy:
6
8
79
3,914
1. Descriptions
Give each tool one clear job. Describe what it does, when to use it, and how it differs from its neighbors.
Tool orthogonality and clear descriptions are the big easy win.
1
1
11
653
2. Enums
Use small, stable enums for categorical arguments. E.g. rooms, currencies, device states.
Needle compiles the schema into grammar, so enums and numeric ranges ensure argument validity.
1
1
9
322
3. Prompts
Keep the system prompt short.
Needle already reads the tool schema. Only write rules that your tools cannot express.
1
1
5
197
4. Naming
Needle prefers snake_case definitions over framework classes or other jargon.
If your tools are expressed in AnotherFormat, convert them to snake_case! 🐍
1
1
4
190
5. State
Needle is a small model with a limited context window.
If two adjacent requests are independent, make sure to Needle.reset() in between!
1
1
5
177
7. Fine-tune
Once you saturate out-of-the-box performance, fine-tune Needle on your own dataset!
Guide here: github.com/cactus-compute/ne…
Aug 16, 2026 · 10:00 PM UTC
1
1
3
251

