We ran 6,120 tests over five weeks to see how far we could push Needle 2. Here's a complete playbook for getting a 14MB function-calling LLM to production accuracy:
6
8
79
3,914
1. Descriptions Give each tool one clear job. Describe what it does, when to use it, and how it differs from its neighbors. Tool orthogonality and clear descriptions are the big easy win.
1
1
11
653
2. Enums Use small, stable enums for categorical arguments. E.g. rooms, currencies, device states. Needle compiles the schema into grammar, so enums and numeric ranges ensure argument validity.
1
1
9
322
3. Prompts Keep the system prompt short. Needle already reads the tool schema. Only write rules that your tools cannot express.
1
1
5
197
4. Naming Needle prefers snake_case definitions over framework classes or other jargon. If your tools are expressed in AnotherFormat, convert them to snake_case! 🐍
1
1
4
190
5. State Needle is a small model with a limited context window. If two adjacent requests are independent, make sure to Needle.reset() in between!
1
1
5
177
6. Measurement Measure tool retrieval, selection, arguments, and validation independently. A single accuracy score cannot tell you the cause of regression.
1
1
4
180
What tool environment are you currently working with? Drop a schema or use case below. We’ll benchmark the most interesting ones and share the results.
1
2
212
Sort replies: Relevant Recent Liked