Building machines with fundamental human qualities Makers of the strongest spreadsheet agent @tryshortcutai

San Francisco, CA
Fundamental Research Labs retweeted
Opus 5 is a shocking step up from Fable 5 and 5.6 Sol for spreadsheets Almost couldn’t believe our evals at first A step function increase in intelligence, while as efficient as Sol Feels like another Opus 4.5 moment for knowledge work
29
65
1,380
112,826
Fundamental Research Labs retweeted
I challenged the MSFT Excel World Champion to a battle. AI beat humans at Chess, then Go. But those are games We built an agent to surpass humans on the most important app in the history of work Meet Shortcut: The Excel AI Agent Comment SATYA and I'll send you free credits
172
98
274
194,228
Fundamental Research Labs retweeted
@nicochristie just beat the Excel world champion at financial modeling by using @tryshortcutai 18 months ago, we demoed to investors an agent that's 2x SOTA on OSWorld, and "superhuman" on its Spreadsheet category, scoring ~80% when every other agent was <20% At the time our agent was able to create pivot tables and that was considered impressive We have truly gone a long way, super proud of the team🔥
I challenged the MSFT Excel World Champion to a battle. AI beat humans at Chess, then Go. But those are games We built an agent to surpass humans on the most important app in the history of work Meet Shortcut: The Excel AI Agent Comment SATYA and I'll send you free credits
2
4
18
2,755
nico
1
7
7,088
Fundamental Research Labs retweeted
What a week! The first GPT model worth being served by default in Shortcut. Let's benchmark Grok 4.5 tomorrow! 😁
we evaluated GPT 5.6-Sol against Fable and Opus using the Shortcut production harness on our own internal set of benchmarks, results were very surprising ShortcutBench V1: swiss-army knife spreadsheet tasks ShortcutBench V2: gnarly, tough, long spreadsheet tasks Sol costs half as much as opus and had similar if not higher accuracy. It also took way less turns and was much faster. Historically, we never served GPT models as default because they are significantly worse at formatting (trust-breaking) than Opus class models. This is also no longer true in our taste tests, although fable is still the best at formatting well done @OpenAI
2
4
1,116
Fundamental Research Labs retweeted
we evaluated GPT 5.6-Sol against Fable and Opus using the Shortcut production harness on our own internal set of benchmarks, results were very surprising ShortcutBench V1: swiss-army knife spreadsheet tasks ShortcutBench V2: gnarly, tough, long spreadsheet tasks Sol costs half as much as opus and had similar if not higher accuracy. It also took way less turns and was much faster. Historically, we never served GPT models as default because they are significantly worse at formatting (trust-breaking) than Opus class models. This is also no longer true in our taste tests, although fable is still the best at formatting well done @OpenAI
7
11
138
15,000
Fundamental Research Labs retweeted
Tianhang finally on twitter. He was a founding member at Qwen and head of RL at 01.ai There's so much discussions about models from the Chinese labs, but very few people on twitter actually worked there and in the US now
Hey guys, I'm leading LLM research for fundamental research labs. I'll be posting more.
2
6
97
36,379
Fundamental Research Labs retweeted
Tianhang is finally on twitter He was formerly a founding team member and head of RL at Qwen, and leads model training here at FRL / shortcut He will be posting more on twitter now Probably the most cracked person i've ever met. Excited for you to see his work soon!
Hey guys, I'm leading LLM research for fundamental research labs. I'll be posting more.
63
122
3,480
844,399
Fundamental Research Labs retweeted
Hey guys, I'm leading LLM research for fundamental research labs. I'll be posting more.
257
141
4,583
2,023,313
Fundamental Research Labs retweeted
Shortcut can now spin up 1000's of agents to parallel search, but we built the feature entirely around the filesystem - enabling some pretty cool new things Here I fill out a huge table of VCs we tracked for a fundraise and another table that has subtables, merged cells, etc.
8
15
92
25,503
Fundamental Research Labs retweeted
/Autonomous is now default on in Shortcut. @BrainsAndTennis built it before Codex released /goal and has gotten pretty awesome reviews. You need runtime observability and verifiable goals for long term autonomy, and our solution to compacting is to never ever do it at all
I think it's safe to say Excel is solved. The sign of an Excel rookie used to be manually using your mouse... now it's doing ANYTHING at all without AI. If you use Excel do yourself a favor and watch all 3 mins. It's a new world. Solving tasks in 20 mins that take teams days.
2
18
6,189
Fundamental Research Labs retweeted
Fable 5 is the best model today, but for everyday spreadsheet task, you won't see a huge difference You'll get much more gain with the right harness & environment Codex + Mog is dramatically better than Codex's default spreadsheet skill, and almost as good as the best agents
Mog v0.9 is published! Mog is the best spreadsheet for agents, and available wherever your agents are Use it in Codex, Claude Code, etc. Just tell them to use npm mog-sdk/sdk, that's all Use it with Shortcut agents, Cloudfare workers, etc Edit xlsx in VSCode, Cursor, etc
3
5
29
20,787
Fundamental Research Labs retweeted
I think it's safe to say Excel is solved. The sign of an Excel rookie used to be manually using your mouse... now it's doing ANYTHING at all without AI. If you use Excel do yourself a favor and watch all 3 mins. It's a new world. Solving tasks in 20 mins that take teams days.
18
40
713
354,398
Fundamental Research Labs retweeted
I'm rebuilding Excel Excel was designed for floppy disks and fax machines, we deserve something better now Run xlsx headless, API for agents + UX for humans, colab, embeddable Mog is the new way to Excel Free, Open source, Building in public: github.com/fundamental-resea…
6
9
79
33,498
Fundamental Research Labs retweeted
For the second time now, @tryshortcutai was independently ranked as the top Excel AI agent in the world. We specifically build for the top analysts, and this matches what we see on the ground. It's firmly us, then Claude, or your firm forces Copilot upon you.
6
10
62
52,216