We are here to ensure our partners become the signal of innovation in the noise of the technology markets.

Pinned Tweet
Two platforms can hold the same number of agents and still give those agents very different response speed. Headcount is the number an enterprise sizes a node by, and the per-agent decode rate is what its users feel, which is why @Signal_65 PINNACLE publishes the whole concurrency sweep and not just the ceiling. MiniMax-M3 (@MiniMax_AI) from is the clearest case on the board. On one fully loaded 8-GPU node running an agentic workload, the @nvidia B300 and the @AMD Instinct MI355X both sustain 52 concurrent agents at the service level in our testing (every agent at 10 tokens per second or better) but the B300 is faster per agent at every rung on the way there. ➡️ 3.7x the per-agent decode rate at 8 agents, 67 against 18 tokens per second ➡️ 2.4x at 16 agents and 1.9x at 32 ➡️ 1.4x at the shared ceiling of 52, 14 against 10 tokens per second ➡️ Past the ceiling the two curves part ways. The B300 drops, with per-agent decode halving and first token past 9 seconds at 64 agents, while the MI355X declines gently with first token flat at 0.2 seconds through 56 For a buyer that means the same headcount with a more responsive agent on the B300, and for AMD it means a competitive performance with the B300 at the ceiling on a serving stack that is still being tuned. Cost per correct task on each platform is the next post. Full sweep and the live results views: stg-pinnaclesignal65-staging…
4
1
14
44,987
The @OpenAI GPT-6 models get better at agentic work when you turn the thinking effort up. @AnthropicAI Claude Opus 5.5 does not. @Signal_65 PINNACLE scored GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 at their default (medium) and maximum reasoning effort on the same multi-step enterprise jobs and the same 128K-token retrieval corpus, at one list price per model. ➡️ GPT-6 Sol at max effort: 2.5x fewer weighted errors than at medium (Model Score 3,998 to 9,853), 99.6% of multi-step jobs finished against 96.8%, fabrication down from 7.9% to 2.9%. It costs 1.8x as much per correct task, 29 cents vs 16, and it lands on the cost frontier above Grok 4.6, Claude Fable 5.1 and GPT-6 Astra at medium. ➡️ GPT-6 Luna at max effort: 2.4x fewer weighted errors (1,682 to 4,115), 99.6% of multi-step jobs finished against 88.6%, for 2x the cost, 3 cents against 1.5. That is above Claude Sonnet 5 and GPT-5.6 Sol at 27x and 33x less per correct task. ➡️ Claude Opus 5.5 at max effort: 1.8% more on the score (10,418 to 10,603) for 3.7x the cost, $1.45 against 40 cents, 4.8x the output tokens, and fabrication up from 2.8% to 4.9%. It already finished all of the multi-step jobs at medium in our testing. ➡️ The trade is the same on all three: max effort spends 3.3x to 4.8x the output tokens per correct answer, so token efficiency falls. On the GPT-6 models you get a better model for it. On Opus 5.5 you get the same model at a higher price, and medium is the setting to run. Every score and every price of a correct task is live at pinnacle.signal65.com
3
1
12
25,591
GPT-6 Luna deserves its own line. At max effort it is the first small hosted model on the PINNACLE board to finish agentic work at a frontier rate, and it does it for 3 cents a correct task. ➡️ 99.6% of the multi-step enterprise jobs finished end to end in our testing, up from 88.6% at medium, the same completion rate as GPT-6 Sol at max ➡️ A Model Score of 4,115, above Claude Sonnet 5 (3,782) and GPT-5.6 Sol (3,954), for 27x and 33x less per correct task, and above every open-weight build on the cost view except Qwen3.8-Flash-Next ➡️ Four rows leave the cost frontier because of it: GPT-6 Sol at medium, DeepSeek-V4-Flash-0731, Qwen3.8-27B and Qwen3.6-27B, so a hosted model now holds every frontier step below 5,000 but one ➡️ The cost of the jump is tokens, 30,624 output tokens per correct answer against 7,332 at medium, and retrieval is still where it shows its size: 92.8% on the 128K corpus against 97.3% for Sol, with fabrication at 11.4%, down from 14.4% but still second highest among the hosted models OpenAI does not publish a parameter count for Luna, and its $0.10 / $0.50 list price puts it in the small tier. Nothing priced like that has finished agentic jobs at this rate on our board before.
1
150
Signal65 retweeted
Even though there was no new silicon announcements for the @Snapdragon PC family, the addition of new @surface devices using X2 plus, new Googlebook system announcements, and Linux support for X2 Elite make for a pretty compelling mid-cycle refresh story.
5
3
21
1,101
Signal65 retweeted
Perplexity is bringing Portable Computer to @AMD Ryzen AI Max, including the Ryzen AI Halo developer platform, so local tasks and recurring workflows can run on the user's own hardware without drawing down Perplexity Computer credits. That connects the cost of an AI subscription directly to the silicon on the desk, and it gives users a concrete reason to pay for a large shared memory pool and a capable integrated GPU in a client system. I expect most PC AI to settle into this hybrid model, with routine agent work running locally and the cloud reserved for frontier-scale reasoning. Ryzen AI Max is one of the few client platforms today with enough memory to hold models large enough to make that split REALLY useful. The value of those savings will depend on how much of a typical agent workload the local models can absorb before handing off to the cloud, and are looking at multiple angles of this at @Signal_65 right now.
Portable Computer for Windows is now available on @AMD Ryzen AI Max Series processors. It makes it easy to run local AI agents that work with your connected apps and local files. Kick off tasks or schedule recurring work that runs entirely on your device.
2
2
20
1,238
Our Opus 5.5 benchmarking versus Fable 5.1: - >2x cheaper for correct task 🔥 - slightly more reliable ✅ but 100% - 4x more hallucinations 😱 but only 2.8%
Any enterprise working on @AnthropicAI Claude Opus for agent work gets more of the work done for less money by moving to yesterday's brand-new Opus 5.5, and in our testing the gap is...pretty large. Opus 5.5 gets more work done for more people, at a noticeably lower price than Anthropic offered before, as measured by @Signal_65 PINNACLE. ➡️ 2.3x fewer weighted errors than Claude Opus 5, and 100% of the multi-step jobs finished end to end against 95% ➡️ 34% less per correct task (!!), $0.40 against $0.60, on a list price 20% below Opus 5 with cache reads at 5% of input rather than 10% ➡️ Fabrication in retrieval answers down from 7.6% to 2.8% ➡️ It also passes Claude Fable 5.1 with 1.3x fewer weighted errors at 63% less per correct task Full comparison on the Signal65 PINNACLE results views at pinnacle.signal65.com
3
7
32
3,939
Any enterprise working on @AnthropicAI Claude Opus for agent work gets more of the work done for less money by moving to yesterday's brand-new Opus 5.5, and in our testing the gap is...pretty large. Opus 5.5 gets more work done for more people, at a noticeably lower price than Anthropic offered before, as measured by @Signal_65 PINNACLE. ➡️ 2.3x fewer weighted errors than Claude Opus 5, and 100% of the multi-step jobs finished end to end against 95% ➡️ 34% less per correct task (!!), $0.40 against $0.60, on a list price 20% below Opus 5 with cache reads at 5% of input rather than 10% ➡️ Fabrication in retrieval answers down from 7.6% to 2.8% ➡️ It also passes Claude Fable 5.1 with 1.3x fewer weighted errors at 63% less per correct task Full comparison on the Signal65 PINNACLE results views at pinnacle.signal65.com
1
3
7
43,437
Along with the new Opus model, yesterday saw the release of new GPT-6 options as well. GPT-6 Luna, in our PINNACLE testing it is the cheapest correct task of any hosted model at 1.5 cents. It is also the model that fabricates the most. @Signal_65 PINNACLE scores long-context retrieval on 369 documents at 128K tokens, lease contracts, sales field reports and HR evaluations generated from one ground truth, and counts every answer that cites something not in the documents. ➡️ GPT-6 Luna fabricates 14.4% of its retrieval answers, the highest rate on the board; the next highest is GPT-5.6 Sol at 10.3% ➡️ The models on the cost frontier above it sit at or near zero, Grok 4.6 and GPT-6 Astra at 0.0%, Claude Fable 5.1 at 0.7% ➡️ GPT-6 Sol, the mid-tier sibling, fabricates 7.9% at 16 cents per correct task, and Claude Opus 5.5 2.8% at 40 cents ➡️ Luna still makes 1.4x fewer weighted errors than GPT-5.6 Luna at a fifth of the price, so the generation gain is real; the fabrication rate is the cost of the price
1
129
Another standout data point from our testing: @OpenAI GPT-6 Sol finishes a correct task in 3,500 output tokens in our testing, the lowest of the 17 hosted builds on the @Signal_65 PINNACLE board. ➡️ 2.8x fewer output tokens per correct answer than Grok 4.6 or Claude Fable 5.1, and 22x fewer than Claude Haiku 4.5 ➡️ The three models launched this week all sit in the terse half of the board, Sol at 3,500, GPT-6 Luna at 7,300 and Claude Opus 5.5 at 7,500 ➡️ Sol scores level with GPT-5.6 Sol on 20% fewer output tokens and a list price half as high, which is how a correct task ends up at a sixth of the price ➡️ The verbose end of the board is where reasoning effort lives, Gemini 3.8 Flash at 20,000 to 24,000 and Muse Spark 1.3 at 37,000
2
133
Posted this over the weekend - worth another look. The improvement in the @AMD MI355X performance is notable from the engineering team. More optimization still to come!
Enterprises that want to utilize the largest open-weight models in an agentic fleet have a narrower choice of nodes and infrastructure that can support that level of quality. Kimi-K3 at 2.78T parameters is the largest model in the @Signal_65 PINNACLE GPU results, served from its MXFP4 weights on both current-generation platforms, and it is the one model where the comparison runs @AMD Instinct MI355X ahead of comp on both supported agent-count and speed. ➡️ 32 concurrent agents at the service level on one MI355X node against 20 on the B300 in our testing, 1.6x the headcount, with first token still in fractions of a second at the ceiling ➡️ From 16 agents up the MI355X is also the faster platform per agent, 20.4 against 15.1 tokens per second at 16 ➡️ The B300 holds a small edge at the lightest shared load, 31.2 against 26.4 tokens per second at 8 agents Full results: pinnacle.signal65.com/articl…
2
15
1,131
Signal65 retweeted
AMD has followed up its Advancing AI data release with a new blog and a white paper on EPYC Venice that attempts to put real detail behind the claims first made on stage in July. The 2.24x platform and 1.2x per-core advantages over NVIDIA Vera in SPECrate 2026 Integer now come with estimated scores, compiler versions, memory configurations, and all 14 subtests broken out. The generational data is impressive, with the 256-core EPYC 9996 landing between roughly 1.5x and 1.8x over the 192-core Turin flagship across Redis, NGINX, MySQL, MongoDB, and server-side Java, all measured by AMD on Venice silicon. @AMD also rebuilt its 100 kW rack model around SPEC CPU 2026, which moved Venice to 3.4x over Vera and nudged its own Turin figure down slightly. The core story is the same one AMD told at Advancing AI, now with some supporting detail. The agents-per-watt and performance-per-watt claims from the keynote have no backing data in this paper, interestingly. All of it is AMD internal testing or estimates, so independent results on shipping systems from both vendors are what will settle the debate. Blog: newsroom.amd.com/news/amd-ep… White paper: amd.com/content/dam/amd/en/d…
6
14
75
8,361
Signal65 retweeted
Also found part of my collection of GPUs where I had been storing at least one of everything going back to the original @NVIDIAGeForce days. I don’t know how much AI these can process, but maybe we should find out!
I was going through a storage unit where I have been keeping a lot of PC hardware stuff, not organizing it yet. Came across some items like this. Who knew I had been holding appreciating assets this whole time!
3
3
25
1,880
Enterprises that want to utilize the largest open-weight models in an agentic fleet have a narrower choice of nodes and infrastructure that can support that level of quality. Kimi-K3 at 2.78T parameters is the largest model in the @Signal_65 PINNACLE GPU results, served from its MXFP4 weights on both current-generation platforms, and it is the one model where the comparison runs @AMD Instinct MI355X ahead of comp on both supported agent-count and speed. ➡️ 32 concurrent agents at the service level on one MI355X node against 20 on the B300 in our testing, 1.6x the headcount, with first token still in fractions of a second at the ceiling ➡️ From 16 agents up the MI355X is also the faster platform per agent, 20.4 against 15.1 tokens per second at 16 ➡️ The B300 holds a small edge at the lightest shared load, 31.2 against 26.4 tokens per second at 8 agents Full results: pinnacle.signal65.com/articl…
1
14
40,951
Enterprises running agents on AMD Instinct today, or sizing a new Instinct fleet, now have an independent number for what the MI355X buys per node. @Signal_65 PINNACLE measured both generations on the same fully loaded 8-GPU node, the same agentic load and the same service level, every agent at 10 tokens per second or better with a first token inside 60 seconds. ➡️ A median 2x the concurrent agents of MI300X across the seven models measured on both in our testing, and up to 6.5x on MiniMax-M3 ➡️ The MI355X is also the faster platform at every agent count the two share, about 2x the tokens on Qwen3.5-35B-A3B and 1.8x on Qwen3.5-397B-A17B at the load where MI300X reaches its ceiling ➡️ The gain is largest on mixture-of-experts models, 2x to 2.4x against 1.3x to 1.8x on dense models, consistent with the step from 192 GB to 288 GB per GPU on agents that carry ~90K tokens of context ➡️ Every comparison runs at the same precision on both generations, and the MI355X results use mostly default configurations, so we expect the advantages to grow rather than shrink as AMD tuning lands Great work by the whole team at @AMD!! @AnushElangovan Full results: stg-pinnaclesignal65-staging…
16
43,964
Local AI on @Windows is not confined to one class of machine. The same platform, containment, and identity model now runs from a mainstream laptop up to a deskside AI supercomputer. From our report: ➡️ Mainstream, the Copilot+ PC install base at roughly $800 to $2,500 ➡️ High performance, the @surface RTX Spark Dev Box at up to 1 petaflop with 128GB unified memory, plus AMD Ryzen AI Max systems running 100B plus parameter models on a single box ➡️ Frontier, DGX Station for Windows on GB300 for models up to 1 trillion parameters, and workstations pooling 192 to 384GB across multiple NVIDIA RTX PRO 6000 Blackwell GPUs ➡️ A developer builds and governs an agent the same way on all three Consistency across tiers is what turns a set of devices into a platform. Full report: signal65.com/research/ai/unm…
5
49,482
“[NVIDIA] Vera finished the agentic task lifecycle 1.64x faster ... than the x86 processor…” @Signal_65 tested our Vera CPU servers against x86, under the demanding load patterns of AI agents with "every core on both processors running a sandbox..." and it delivered. Each agentic task starts and ends on the CPU: orchestration, sandboxing, tool calls. Vera keeps that loop fast at scale, and independent testing continues to highlight the power of a new CPU design point for AI.
🚨🚨NVIDIA Vera CPU Agentic Performance🚨🚨 If you are planning capacity for agents (and aren't we all??), the CPU decides how many sessions a rack can keep busy, and the GPU spends much of its day waiting on it. Over the past few weeks Signal65 worked with @nvidia to measure that work on the new Vera CPU, one sandbox per core, no model in the loop, timed from sandbox creation to teardown. In our testing Vera finished the agentic task lifecycle 1.64x faster per core than a leading x86 server processor across 69 Terminal-Bench 2 tasks, and 1.87x on the 20 that look most like an agent's daily work. Our own runs on the same hardware came in faster than the NVIDIA figures, so the headline number is conservative for this comparison. Full write-up and per-task data: signal65.com/research/ai/val…
8
9
94
5,960
Signal65 worked with $NVDA to time the agentic task lifecycle on Vera: 1.64x faster per core than a leading x86 server processor across 69 Terminal-Bench 2 tasks, and 1.87x on the 20 closest to daily agent work. An agent spends most of its clock outside the model. That time is CPU time. Full methodology and per-task data from Signal65 PINNACLE: signal65.com/research/ai/val…
🚨🚨NVIDIA Vera CPU Agentic Performance🚨🚨 If you are planning capacity for agents (and aren't we all??), the CPU decides how many sessions a rack can keep busy, and the GPU spends much of its day waiting on it. Over the past few weeks Signal65 worked with @nvidia to measure that work on the new Vera CPU, one sandbox per core, no model in the loop, timed from sandbox creation to teardown. In our testing Vera finished the agentic task lifecycle 1.64x faster per core than a leading x86 server processor across 69 Terminal-Bench 2 tasks, and 1.87x on the 20 that look most like an agent's daily work. Our own runs on the same hardware came in faster than the NVIDIA figures, so the headline number is conservative for this comparison. Full write-up and per-task data: signal65.com/research/ai/val…
1
2
10
3,127
🚨🚨NVIDIA Vera CPU Agentic Performance🚨🚨 If you are planning capacity for agents (and aren't we all??), the CPU decides how many sessions a rack can keep busy, and the GPU spends much of its day waiting on it. Over the past few weeks Signal65 worked with @nvidia to measure that work on the new Vera CPU, one sandbox per core, no model in the loop, timed from sandbox creation to teardown. In our testing Vera finished the agentic task lifecycle 1.64x faster per core than a leading x86 server processor across 69 Terminal-Bench 2 tasks, and 1.87x on the 20 that look most like an agent's daily work. Our own runs on the same hardware came in faster than the NVIDIA figures, so the headline number is conservative for this comparison. Full write-up and per-task data: signal65.com/research/ai/val…
2
2
14
9,730
Two weeks ago Claude Opus 5 was the top of the PINNACLE board but an open-weight model beats it on four of five enterprise personas🔓🔓 An enterprise that standardized on Claude Opus 5 this summer bought the highest score on the Signal65 PINNACLE board, but you can now download a model that beats it. (Data at pinnacle.signal65.com) Opus 5 led the board when PINNACLE launched on August 31st and was the Anthropic flagship until September 1st, and in our testing Qwen3.8-2.4T-A95B, an open-weight model from @Alibaba_Qwen scored at medium reasoning effort, makes 15% fewer weighted errors on the same multi-step enterprise work. At launch the best open-weight model trailed Opus 5 by 1.4x in weighted errors, and two weeks later the best open-weight model leads it, which is the distance the frontier is being asked to slow down from. Persona by persona, the open model leads four of the five. ➡️ Customer Operations is the widest margin, 1.56x fewer weighted errors. Qwen invented an answer on 1.3% of the unanswerable questions at 128K context in our testing, against 7.6% for Opus 5. Executive Decision Support is 1.22x, Knowledge Worker 1.10x and Data Analyst 1.05x. ➡️ IT Professional, the most heavily agentic persona at 55% weight on chained tool use, is the one Opus 5 holds, by 1%, which is inside the resolution of the test. Opus 5 also reaches a correct answer in 1.9x fewer output tokens, 9,668 against 17,944, so the hosted model still gets there with far less generated work. ➡️ The measurements underneath go the same way. Qwen finished 96.4% of the multi-step jobs it started end to end against 95.0% for Opus 5, and its retrieval fidelity at 128K context was 96.6% against 95.5%. But even the frontier has moved on since August, and GPT-6 Astra at maximum effort makes 3x fewer weighted errors than Qwen 3.8. The open weights are now where that frontier stood two weeks ago, and any plan to pace the frontier is a plan to let the distance close. Every score, persona by persona, is live on the PINNACLE results views at pinnacle.signal65.com
2
3
8
2,510
Signal65 retweeted
New Pinnacle Agentic benchmarks are live and the open weights are closer to the frontier than the pacing debate wants to admit. We just scored eight new models on PINNACLE. Five are open. Qwen3.8 made 15% fewer errors than Claude Opus 5 in our testing, and its 180B sibling runs on a desktop. If the gap is that small, the leverage in any pacing agreement sits with whoever can run their own hardware. That is the story for investors, not the safety letter. Maybe this explains the weekend headlines? Read the numbers before the commentary.
You cannot pace a frontier you have not measured, and in our testing the open weights are moving closer to it 💡 An enterprise choosing a model for agent work has more good options this month than in August, and a louder argument around them. Three of the labs at the top of the Signal65 PINNACLE board, @OpenAI, @SpaceXAI and @AnthropicAI, backed a call last week to pace frontier AI development, and the same week we scored eight new configurations, five of them open weights from @Alibaba_Qwen, @deepseek_ai and @Zai_org. Any agreement to pace capability starts from an independent measurement of where it stands, and in our testing the open weights landed closer to the hosted frontier than the commentary allows. Both hosted additions land on the cost frontier as well. ➡️ Gemini 3.8 Flash has drawn online skepticism as an agent. In our testing it finished 99.6% of the multi-step enterprise jobs it started at high effort and 98.9% at medium, second and third on the board behind only GPT-6 Astra at max effort. On the IT Professional persona, the most agentic of the five, the high setting makes about half the weighted errors of Claude Fable 5.1. It takes a median 59 tool calls per task against 8 to 11 for Astra, at 54 to 65 cents per correct task. ➡️ Grok 4.6, our first look at xAI, lands fourth on the board at 47 cents per correct task with 0.0% fabrication in the retrieval suite. That is the same score as Astra at medium effort for half the price, and it puts Grok on the cost frontier alongside both Gemini settings. But the open-weight additions matter most to a buyer running its own hardware, and they will matter more if the industry pushes away from the idea of slowing down. ➡️ Qwen3.8-2.4T-A95B is the strongest open-weight result we have scored, with 15% fewer weighted errors than Claude Opus 5 in our testing, and it did that at medium reasoning effort. The better comparison for a buyer is Qwen3.8-Flash-Next, a 180B model within 5% of its 2.4T sibling that runs on a desktop system like the NVIDIA DGX Spark. ➡️ DeepSeek-V4.1-Flash at maximum reasoning effort makes 1.5x fewer weighted errors than the previous DeepSeek V4 Flash and 1.8x fewer than its own high setting, so the generation gain is real. It pays with 20x the output tokens per correct answer and a median of more than 40 tool calls per task, the lowest token efficiency on the board, and whether that trade pays depends on what those tokens cost on your hardware. ➡️ GLM-5.3-Flash, 320B total and 18B active, lands between the full GLM-5.3 at maximum and low effort in our testing, so the small model gives up little to the large one, although both remain below the previous GLM-5.2. A paced frontier, if one is ever agreed, will be negotiated over numbers like these, and in our testing the open weights are close enough behind the hosted models that there is less room to pace than the debate assumes. The score and price of a correct task for all 62 scored builds are live on the PINNACLE results views at pinnacle.signal65.com
6
3
29
6,801