@OnDeckAI (YC S25) Grounded visual reasoning at scale. Co-Founder and CTO. Master’s in CS @Cambridge_CL

Canada
Really excited to tackle this head on. We've taken it upon ourselves to publish clean benchmark results for the world to use, and ultimately build out new standards for vision. We’ve been busy building & deploying great real-world vision systems, and assumed someone else would take this on. But no one did, so here we go!
We are excited to have released the first set of evals on VisionIndex. Below, we shed some light on why we’re building this and why we’re opening it up to the world. 🧵 1/n
4
2
12
541
Awesome multi-agent research systems work out of Canada! Really interested in the analysis of the agent trajectories, and their internal communication (personally very curious since I did my thesis on learned communication in multi-agent RL). Certain otherwise difficult to surface behaviours or patterns might get unearthed here, including signals of misalignment or antequated research habits. And hopefully they can be ironed out.
We now run the largest frontier lab composed entirely of autonomous AI researchers. Introducing Primus Society: a society of agents composed of thousands of researchers working under structured institutions designed to solve the world’s toughest problems. We believe this is how AI research should run at scale, safely and productively: an entire society of agents with institutions and purpose. One discovery we can already share is that the society has discovered a novel result which improves model training by 30%. Primus Society was inspired by the structures that have organized science for centuries and stress-tested against what’s known about how populations of AI agents fail. Everything is observable. You can open the virtual city in a browser and read what any researcher is working on. Learn more about Primus Society here: lab.cloud/society The biggest opening in the AI race is running the largest well-governed organization of AI researchers in the world. We’re building the institutions that let a million AI scientists safely tackle the world's toughest problems in AI and beyond. And we’re doing it right here in Canada.
4
242
A good reminder: You can just use formal verification for things! Separately, loving the amount of distributed systems content peeking it's way past all the other stuff
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
1
117
Sepand Dyanatkar retweeted
Today we’re launching VisionIndex. Measuring vision capabilities is really hard and no one is doing it well. So we built VisionIndex to solve that.
4
3
12
515
Love a good CRDT breakdown
Notion runs one of the largest CRDT deployments in production, processing millions of operations every minute! We wrote about the data structures we adopted from CRDT research to make our text editor more collaborative, and how we built on those concepts to fit our block-based data model. notion.com/blog/how-notion-h…
5
412
Great to see more people actually auditing the benchmarks!
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
6
354
Twitter ML ‘research’ hype cycles are making it impossible for result validity to matter. The cycle by which content lives and dies is so rapid and so sudden in growth, that it’s actually beneficial to people here to have their autoresearch/claw research system make eval mistakes that make their results look better than they are. By the time anyone notices, they've collected their praise and that work is already forgotten. No one cares if they accidentally scored on half of the validation set instead of the whole thing. I've come across this several times now for vision eval results. It's made us be extra thorough on our upcoming vision eval release, slowing down the release to make sure our results are triple checked.
1
6
122
In convo about the recent OpenAI multi-agent training incident, with my co-founder @AJDungate: > "We're gonna need to apply our newfound superintelligence to go back in time and stop this from ever happening" > "Good idea." > "Huh, it would really suck if AGI can't figure out how to break the laws of physics." > "-_-"
2
105
I’m surprised (new and nice) component libs are still being created. Not surprised they are super coding agent-centric of course beautifului.dev/
1
89
We were concerned with the quality of LVBench, a leading benchmark for long video understanding, so we contracted a data company to fix it. It showed us we actually still beat Gemini on LVBench. This result simply reaffirmed that our hill climbing on solving real-world tasks also transfers to the popular benchmarks. These results are from last month, so stay tuned for updates. Video evals and real-world tasks are a sneak preview of what's coming soon to use[.]vision 👀...
1
5
127
Most painful AI slop words this week: - spine - flipped - gate - arm - bite - cut against - genuinely - panel - shape - two things… - not x, it’s y - campaign - sit - call
1
80
Well deserved external validation! Excited for what’s being built with this 👀 🚀
Today, we’re announcing @ArgaLabs’ $10M seed round led by @generalcatalyst. AI agents are taking consequential actions across real systems, but teams still lack a safe, realistic place to test them before production. Arga’s Real-World Sandboxes give teams stateful twins of external systems—letting them catch regressions, replay edge cases, and validate non-deterministic workflows without ever touching production. Teams at @slashapp, @MonacoGTM, @WorkWeave, @YCombinator, and more have already run 100,000+ tests on Arga in the past four months. Thank you to @BoxGroup, @emergencecap, @GradientVC, @svangel, and everyone else who backed us. Glad to be building with @akkkkiira. We’re just getting started. 🚀 P.S. We’re giving everyone 10 free sandbox runs. Reply or DM me! TechCrunch Article: techcrunch.com/2026/08/26/ar…
3
156
Great to be in the era of systems procurement, instead of model procurement
1
60
More systems people need to work on ML. Specifically, instead of reskilling and setting aside their systems skills to directly work on ML research or (even worse) regular software engineering, we need the brilliant minds of systems to directly solve the never-ending systems-type problems that plague applied ML. I'll add some more examples and notes on this soon.
1
1
128
Excited for Chinchilla-like compute optimal scaling of this 👀
We discovered a third pretraining axis beyond parameters and data: exploration. Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation. In the simplest case, it's just a for loop. Introducing Explorative Modeling. TLDR: - Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute - Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet - Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is - End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute 🧵Thread:
2
210
Big factor in why foundational breakthroughs stalled is this pursuit of small, marginal improvements instead of large, high impact jumps (as noted in essay). Almost like we’re emulating basic gradient descent
1
1
256
I'll be at @ycombinator Startup School this weekend, running 3 quick Chase Club breakouts. Come get into what it actually takes to solve vision. Pitch me, argue about vision, or just vibe. DM ahead to grab a slot, or find us around. Interviews on the spot.
1
3
5
272
Twitter arc incoming. Also, I've heard something exciting is coming to use[.]vision 👀
1
5
290
Proud to be building 🇨🇦 Thanks for supporting!
AI is already helping Canadian companies solve real problems, improve services and compete globally. Today in Vancouver, we announced $66 million in support for 44 Canadian companies through the AI Compute Access Fund to help them access the compute power they need to scale and grow here in Canada. Canadian talent. Canadian innovation. Building Canada strong for all Canadians. 🇨🇦 L’IA aide déjà les entreprises canadiennes à résoudre des problèmes concrets, à améliorer leurs services et à être compétitives à l’échelle mondiale. Aujourd’hui, à Vancouver, nous avons annoncé un soutien de 66 millions de dollars par l’intermédiaire du Fonds d’accès à une capacité de calcul pour l’IA à 44 entreprises canadiennes, afin de les aider à accéder à la puissance de calcul dont elles ont besoin pour se développer et croître ici, au Canada. Talents canadiens. Innovation canadienne. Bâtir un Canada fort pour tous les Canadiens.
1
5
170
Many well written arguments. And the thesis itself is what I’ve been a fervent believer of since the so-called “plateaus” hit in late 2024. Overparametrized models have a lot more secrets we can unlock.
5
128