You cannot pace a frontier you have not measured, and in our testing the open weights are moving closer to it 💡
An enterprise choosing a model for agent work has more good options this month than in August, and a louder argument around them. Three of the labs at the top of the Signal65 PINNACLE board,
@OpenAI,
@SpaceXAI and
@AnthropicAI, backed a call last week to pace frontier AI development, and the same week we scored eight new configurations, five of them open weights from
@Alibaba_Qwen,
@deepseek_ai and
@Zai_org.
Any agreement to pace capability starts from an independent measurement of where it stands, and in our testing the open weights landed closer to the hosted frontier than the commentary allows.
Both hosted additions land on the cost frontier as well.
➡️ Gemini 3.8 Flash has drawn online skepticism as an agent. In our testing it finished 99.6% of the multi-step enterprise jobs it started at high effort and 98.9% at medium, second and third on the board behind only GPT-6 Astra at max effort. On the IT Professional persona, the most agentic of the five, the high setting makes about half the weighted errors of Claude Fable 5.1. It takes a median 59 tool calls per task against 8 to 11 for Astra, at 54 to 65 cents per correct task.
➡️ Grok 4.6, our first look at xAI, lands fourth on the board at 47 cents per correct task with 0.0% fabrication in the retrieval suite. That is the same score as Astra at medium effort for half the price, and it puts Grok on the cost frontier alongside both Gemini settings.
But the open-weight additions matter most to a buyer running its own hardware, and they will matter more if the industry pushes away from the idea of slowing down.
➡️ Qwen3.8-2.4T-A95B is the strongest open-weight result we have scored, with 15% fewer weighted errors than Claude Opus 5 in our testing, and it did that at medium reasoning effort. The better comparison for a buyer is Qwen3.8-Flash-Next, a 180B model within 5% of its 2.4T sibling that runs on a desktop system like the NVIDIA DGX Spark.
➡️ DeepSeek-V4.1-Flash at maximum reasoning effort makes 1.5x fewer weighted errors than the previous DeepSeek V4 Flash and 1.8x fewer than its own high setting, so the generation gain is real. It pays with 20x the output tokens per correct answer and a median of more than 40 tool calls per task, the lowest token efficiency on the board, and whether that trade pays depends on what those tokens cost on your hardware.
➡️ GLM-5.3-Flash, 320B total and 18B active, lands between the full GLM-5.3 at maximum and low effort in our testing, so the small model gives up little to the large one, although both remain below the previous GLM-5.2.
A paced frontier, if one is ever agreed, will be negotiated over numbers like these, and in our testing the open weights are close enough behind the hosted models that there is less room to pace than the debate assumes. The score and price of a correct task for all 62 scored builds are live on the PINNACLE results views at
pinnacle.signal65.com