Pay attention to what @Signal_65 has been working on, the team is bringing a new and very pragmatic view to model and agent evaluation. Net - it’s all about user outcome economics. There is also a neutrality and transparency here that is sorely lacking in other evals
The interesting traits we find with these different models and design principles continues to open my eyes. It's clear @OpenAI uses a strategy to scale model intelligence based on token consumption, allowing them to move up and down our @Signal_65 PINNACLE benchmark and giving them flexibility based on demand cycles. Maybe @AnthropicAI feels more pressure to keep up with competition so is already putting its best foot forward with the new Opus 5.5 at default thinking? I'm going to be VERY curious to see if we see any more reported model capability regression in a month or two; something that the community has tracked previously.
1
3
648
Dan O’Brien retweeted
Any enterprise working on @AnthropicAI Claude Opus for agent work gets more of the work done for less money by moving to yesterday's brand-new Opus 5.5, and in our testing the gap is...pretty large. Opus 5.5 gets more work done for more people, at a noticeably lower price than Anthropic offered before, as measured by @Signal_65 PINNACLE. ➡️ 2.3x fewer weighted errors than Claude Opus 5, and 100% of the multi-step jobs finished end to end against 95% ➡️ 34% less per correct task (!!), $0.40 against $0.60, on a list price 20% below Opus 5 with cache reads at 5% of input rather than 10% ➡️ Fabrication in retrieval answers down from 7.6% to 2.8% ➡️ It also passes Claude Fable 5.1 with 1.3x fewer weighted errors at 63% less per correct task Full comparison on the Signal65 PINNACLE results views at pinnacle.signal65.com
1
3
7
43,649
Limited seats available on this 🚀 If you want to work with the best clients in the world, at the edge of technology innovation, with the best data & platform to elevate your research, and have fun collaborating with a smart, driven group of people, then let’s talk! careers@futurumgroup.com
Some of our best recent hires have been found here on X. So I’m going back to the well here. We are hiring Tech Analysts at Futurum. Looking for AI-pilled, tech-obsessed, data-loving, social media crazed analysts across Semiconductors, Infrastructure, Cyber, Developer/DevOps, and broadly across the AI sphere. Analysts that join us will be part of a team that works closely with the largest technology vendors, financial services firms, and enterprise users in the world. The work will include (but not limited to) -Product Research -Market Modeling -Testing and Validation -Evaluations and Assessments -Attending Vendor Events -Public speaking (TV, Podcasts, Radio) -Sharing your views on X -Lots of other fun stuff We are fast growing firm that is building technology, modeling important markets, providing market education, and we are always innovating across our businesses that include @TheFuturumGroup, @FuturumEquities, @ETRnews, @TechFieldDay, @TechstrongTV, @VisibleImpactUS, @TheSixFiveMedia, and @Signal_65 Experience from various backgrounds will be considered. Entrepreneurial, highly motivated, intellectually curious, and overly opinionated (with work to back it up) will receive highest consideration. Reply below or email website@futurumgroup.com and feel free to pass along or share with your AI-pilled super smart friends that may be perfect for this gig.
6
342
Dan O’Brien retweeted
The industry spent this week debating the AI pacing argument Anthropic CEO Dario Amodei set off with "We Must Pace the Frontier." We asked 50 large-enterprise buyers from our ETR Community: 70% say the labs should slow down, but only 20% read lab-proposed slowdowns as genuine safety measures. The detail that makes this worth publishing sits in a question nobody would think to look at. Among the 72% who want more regulation, 86% cite safety and misuse risk, but only one respondent cited slowing the pace so their own organization could keep up.
2
3
6
7,487
There’s been a lot of talk this week about slowing down or pacing AI development. A lot has been said from top leaders across the industry, government officials, and in the media. At @TheFuturumGroup we felt that the buyer and user were underrepresented in this industry dialogue, so we talked to 50 large US enterprise technology leaders from our @ETRnews community to see how the biggest AI consumers and spenders feel about the pacing debate. Key findings include: 1: the argument lands. 70% said the labs should slow down. 2: the messenger doesn't. Only 20% read lab-proposed slowdowns as genuine safety measures. 46% read them as competitive positioning. Buyers accept the argument and distrust the arguer. 3: they converge on a mechanism that neither the labs nor governments control. Independent third-party evaluation was the top answer for who should set the pace, at 40%. And it moves every commercial metric we tested — 70% more willing to deploy in production, 74% more confident building on that lab's models, 60% would pay a premium for independently certified models. Demand for third-party certification isn't a preference. It's a capability gap. Whoever builds the assurance layer for frontier models has a market waiting.
1
5
224
Dan O’Brien retweeted
Highest satisfaction in our universe. Lowest cost score of any tool we track. Users love it. Buyers argue with the bill. So far nobody has had to choose.
A year ago, Anthropic trailed OpenAI by 10 pts with the Global 2000. Today, it leads by 25. In the same survey, its midsize and small accounts fell 8 pts. The deepest budgets are deepening. The tail is pulling back. Preliminary October TSIS, still in the field.
2
7
27
61,383
If you haven’t seen a demo of what @TheFuturumGroup’s @ETRnews community studies are saying about your customer and prospect buying intent, I’d like to personally show you. It’s been really fun watching our clients react to what the data is saying about their companies. I’ve heard everything from “I knew it!” to “yes, we feel we are taking share vs (insert XYZ competitor)”; seeing gut instinct quantified has reinforced how predictive and tightly correlated the data is to reality. And now the ETR data powers the research and insights at the heart of everything we do to deliver value to our clients. @danielnewmanUV @Tiffani_Bova @ashimmy @SFoskett
A year ago, Anthropic trailed OpenAI by 10 pts with the Global 2000. Today, it leads by 25. In the same survey, its midsize and small accounts fell 8 pts. The deepest budgets are deepening. The tail is pulling back. Preliminary October TSIS, still in the field.
1
1
8
1,262
Reminder that everyone is talking their book on the AI slowdown/speed-up debate. OpenAI/Anthropic have the most to gain from slowdown/regulation. Google/Meta can afford to keep pushing forward because their core businesses benefit from building better models. And yes, even if we slow down recursive RL training, we need boatloads more compute for inference, and training of new model types (like for physical AI).
125
Dan O’Brien retweeted
💸💸Closed Model AI pricing is broken, and the vendors proved it with their own product lines 🤖🤖 Hosted AI crossed a threshold this month, in intelligence and in stability at once. In our testing the newest models from @OpenAI @Google and @AnthropicAI finish 98% or more of the multi-step enterprise jobs they start and rarely invent an answer, and they do it for roughly the same dollar per finished job the previous tier charged. @Signal_65 PINNACLE is our independent benchmark that scores AI models on real multi-step enterprise work and prices every correct task on the full API bill, or on the rented GPU node an open model runs on. The three snapshots connect each lab's models from the cheapest correct task upward. OpenAI put the new generation at the old price, and in our testing the ladder goes straight up between GPT-5.6 and GPT-6. ➡️ GPT-6 Astra at medium effort costs the same per correct task as GPT-5.6 Sol, about 94 cents, with 2.3x fewer weighted errors. A list price 2.5x higher was cancelled by 2.9x fewer input tokens per finished job. ➡️ Maximum effort on Astra buys another 1.8x fewer errors for 1.6x the cost and the top score on the board. The same dial on Sol bought nothing, so reasoning effort only pays on the new model. ➡️ GPT-5.6 Luna clears the open-weight baseline for 7 cents a task, the cheapest correct task from any hosted model we have tested. The Anthropic ladder has the new generation at the top and a broken rung at the bottom in our testing. ➡️ Claude Haiku 4.5 lists at half the price of Claude Sonnet 5 and costs more per correct task, $1.09 against $0.62, because it finished 29% of the jobs and every failed attempt is billed. A router that sends easy work to the small model sends it to the dearest rung here. ➡️ Sonnet 5 is the value rung of the line. Claude Opus 5 costs 2.2x Sonnet for 1.2x fewer weighted errors. ➡️ Claude Fable 5.1 is the Anthropic entry in the new generation and the one that did not arrive at the old price, 1.7x fewer errors than Opus for 1.8x the cost and the most expensive correct task from any hosted model on the board at $2.46. Google shows the whole curve, four open-weight Gemma 4 builds priced on a rented B300 node and two settings of hosted Gemini 3.8 Flash, and in our testing the ladder turns in the middle. ➡️ The Gemma builds get cheaper per correct task as they get better, because job completion climbs from 17% to 54% faster than the token bill grows. Below the turn the better model is also the cheaper one and there is no trade-off to make. ➡️ Gemini 3.8 Flash at medium effort makes 11x fewer weighted errors than the best Gemma for 1.5x the cost, 54 cents against 35. Inside the Google catalog the hosted Flash beats running the open weights yourself at the same money. ➡️ The high effort setting adds 16% fewer weighted errors for 20% more. Above the turn you are paying for reasoning, not completion. The score and the price of a correct task for every model we have tested, hosted and open, are live on the PINNACLE results views at pinnacle.signal65.com
5
2
12
37,252
AI-native delivery is the future, and the future is now! The modern analyst subscription is a time-bound license to all of our IP, embedded into the way our clients work with no access restrictions.
Another week, another launch. Announcing the Futurum Intelligence API & MCP Connector! The real AI risk to any industry is thinking the same way we used to do things will continue to make sense in an era of almost infinite intelligence. We have been full steam ahead to meet the AI moment, and our next big launch is here today with the release of our Futurum Intelligence API & MCP connector. Essentially, we bring our entire catalog of data, research, insights, and media to you. Your workflow in Claude, Copilot, ChatGPT, or Gemini. More than 10 million proprietary data points across our catalog of market models, buyer intent, decision making, and research accessible to you in your workflow with zero friction. I couldn't be prouder of our team and the work that has been put in to build this novel capability for our clients across the vendor landscape, end users, and financial services firms. The future of how market intelligence and technology research is consumed has arrived and I'm so excited that Futurum was able to be one of the first to make it widely available to our customers.
3
589
Dan O’Brien retweeted
AI leaderboards are lying to you. The question was never which model is smartest. It is which one finishes the job on your data, at a price you can defend. Signal65 launched PINNACLE today to answer it.
Article

AI Leaderboards Are Lying To You. Meet Pinnacle by Signal65

Why Signal65's PINNACLE becomes the standard benchmark for agentic AI, and what enterprise buyers already told ETR. Enterprises do not buy tokens. They buy correct, completed work. I have said for two

13
7
50
29,474
Awesome to see this launch, and reset market expectations about what success looks like for agentic AI. Sometimes the tech industry gets so into performance that we forget that “it just works” is what most users need, and what really allows for adoption to scale.
1
1
6
824
Dan O’Brien retweeted
That’s a wrap on AI Unleashed 2026. Thank you to all our speakers, partners and participants, including @salesforce Chair & CEO @Benioff, @Snowflake CEO @RamaswmySridhar, @MarvellTech CEO Matt Murphy and @Applied4Tech CEO Gary Dickerson. Missed the live sessions? Watch on demand: sixfivemedia.com/summit
3
7
26,317
Always great to discuss @TheFuturumGroup's take on @nvidia earnings on @SchwabNetwork. We discussed the new growth floor set by the FY28 guidance, supply vs. demand, and NVIDIA's masterful use of the balance sheet to secure HBM capacity needed to accelerate revenue growth.
Alex Roy of @salesboxai and @danOBtech of @TheFuturumGroup discuss $NVDA post-earnings. Roy sees "very promising times ahead" for the company, and O'Brien says the report was a "reset of the floor higher" as the company projected 70% revenue growth in its next fiscal year. For more market news, tune in at: SchwabNetwork.com/?CID=SM:Tw…
1
1
6
1,726
For the “circular financing” crowd, doesn’t the net result of the “bubble bursting” scenario leave $NVDA with a level of vertical integration that they wouldn’t never be able to achieve through M&A given regulatory hurdles?
1
93
Humans vs AI With AI increasing average intelligence, the supply of above-average intelligence decreases, meaning that unless demand for above-average intelligence decreases, the value of above-average intelligence rises. AI favors the expert vs the generalist.
62