Advancing multimodal AI & computer-use

San Francisco, CA
Pinned Tweet
How good are agents actually at CAD? Today we release CADBench, evaluating 10 Models on 105 realistic tasks in Fusion 360 Fable 5 and Gemini 3.7 Flash are #1 with only 24.6% pass rate
27
61
540
147,425
Our Co-Founders @PascalBtd and @atoof_sh sat down on the Europe on Edge podcast and talked about the current state of multimodal AI and the company's mission. If this mission resonates with you, have a look at our open job listings (links are below).
1
2
11
220
New SOTA on CADBench: Gemini 3.8 Flash. 30% pass rate on 105 expert-authored CAD tasks in Autodesk Fusion, up from 24.6% for Gemini 3.7 Flash. Muse Spark 1.3 also improves its mean verifier score by 54% over 1.2, though full task completion remains a challenge.
2
3
11
1,170
Muse Spark 1.3 improves despite pass rate staying at 0%: • Mean verifier score: 13.44 → 20.66 (+54%) • Average runtime: 60.8 → 23.8 minutes (61% shorter) • Valid tool calls: 76.8% → 93.0% Better intermediate results and faster runs, but full task completion remains unsolved.
1
3
86
Seldon retweeted
Can computer-use agents use complicated graphical software? You’re watching GPT-5.6 Sol operate Autodesk Fusion for 250 turns as it attempts a real mechanical-design task. The run is part of a benchmark aimed at the question "how good are agents at CAD?" @seldon_tech's CADBench contains 105 mechanical-design tasks tested across 10 frontier models, and the results show how early reliable, autonomous CAD work remains: More than two-thirds of the tasks are still unsolved. That gap matters as agents make their way into the professional tools used to design the physical world. Getting around Autodesk Fusion is one thing, but producing a correct, editable model an engineer can pick up and use is much harder. CADBench gives us a way to see and measure that frontier. Great work from the Seldon team! Explore CADBench: seldon.global/blog/cadbench
1
2
7
506
We're hosting another event in San Francisco coming monday, this time a research talk with @samsja19 as guest speaker Drop by if you're in town! Sign up link is in the comments
1
1
15
3,619
How good are agents actually at CAD? Today we release CADBench, evaluating 10 Models on 105 realistic tasks in Fusion 360 Fable 5 and Gemini 3.7 Flash are #1 with only 24.6% pass rate
27
61
540
147,425
Gemini 3.7 Flash delivers performance comparable to Fable 5 at a fraction of the cost We see that many models finish the tasks prematurely, resulting in lower scores and costs
1
37
3,999
Infinibench-v0: part of the benchmark was created by a continuously exploring agent, rewarded for identifying new failure modes in VLMs. It produces non-trivial, out-of-distribution problems — trivially easy for humans, yet SOTA VLMs fail — and diagnoses failure modes that were previously unknown.
1
11
226
The 12 skills:
1
10
188
Introducing VGI-Bench: a multimodal, holistic benchmark probing 12 distinct visual and audio-visual skills. 550 human-curated questions, designed to mitigate the common mistakes in today's video benchmarks and expose pragmatic failures of state-of-the-art models. Best model: 64.73%. Humans: 84.5%.
10
14
62
8,817
Our dataset consists of 398 public videos and programmatically generated videos, the latter to avoid models having been exposed to the test data at train time, spanning just under 100 hours.
1
9
193
Every question in our final dataset must pass two gates, in order: -> Not solved in pass^3 by a blind, text-only model -> Not solved in pass^3 by an older baseline VLM (Gemini 2.5 Flash-Lite) This ensure that every questions tests actual visual understanding
1
11
244
VLMs are today's de-facto standard for visual reasoning. They are widely applied in embodied systems like humanoid robots and self-driving cars, and serve as the visual backbones for computer-use and design agents. This demands diagnostic benchmarks that show exactly where a model can be trusted and what remains difficult.
1
1
18
1,233