We are happy to announce the release of the Phi 3.5 family of models! Check them out: 👉Phi 3.5 mini: huggingface.co/microsoft/Phi… 👉MoE: huggingface.co/microsoft/Phi… 👉Phi 3.5 vision: huggingface.co/microsoft/Phi…
13
38
3,000
Victor Fragoso 💻 retweeted
Our paper, CIPHERGRID, was accepted at #NeurIPS2026! We test whether models can take familiar abilities—visual reasoning, rule inference, and planning—and combine them in a setting where both the representation and rules have to be figured out from scratch. This will be my second time attending #NeurIPS, after giving a tutorial on Positional Encoding last year. If you are attending and interested in discussing architecture, abstract reasoning evaluation, and human-AI interaction, let’s connect! It was an absolute pleasure to work with @VMFragoso and @saiphcita
1
5
9
177
Frontier LRMs can navigate grids; but can they infer the rules first? Our accepted NeurIPS E&D paper introduces CIPHERGRID, a benchmark for decoding an unfamiliar symbolic system and carrying it through sequential action. Humans crush the baseline; no model reaches 50%.
2
3
212
Victor Fragoso 💻 retweeted
Our paper CIPHERGRID was accepted to NEURIPS! Can frontier models solve puzzles when rules & language are latent? We test whether Large Reasoning Models can infer an unknown symbolic system from multimodal data & plan sequentially Kudos Chris Curtis +@VMFragoso ! #NeurIPS2026
3
5
11
658
Victor Fragoso 💻 retweeted
I’m hiring PhD students and postdocs! With recent funding from NSF, DOE, ARPA-E, and NIH, my research group is expanding. (1/10)
17
146
550
91,158
Victor Fragoso 💻 retweeted
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia? These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach. The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date. The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant. At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems. So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models. Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field. If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section. Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models. I'm interested in your thoughts.
91
388
2,362
673,602
Victor Fragoso 💻 retweeted
We are happy to announce the list of accepted NeurIPS 2026 workshops: blog.neurips.cc/2026/08/10/a… The workshops will take place on: - Fri Dec 11 and Sat Dec 12, 2026 – Sydney - Sat Dec 12 and Sun Dec 13, 2026 – Paris and Atlanta We want to thank everyone who submitted a proposal and look forward to hosting the accepted workshops!
6
38
346
119,998
Victor Fragoso 💻 retweeted
Impactful research has always been about asking the right questions. AI alone isn’t going to make you a good researcher. It’s better judgement to ask the right questions that will.
4
5
71
7,632
Victor Fragoso 💻 retweeted
28
520
813
16,235
Victor Fragoso 💻 retweeted
computer vision used to be applied signal processing. the dark art was feature engineering: turning intuition for optics & geometry into filters now it’s applied machine learning. the new dark art is architecture & datasets: turning intuition for filters & optimization into models
3
8
224
19,297
Victor Fragoso 💻 retweeted
The human brain is strikingly modular, with distinct networks for language, formal reasoning, social reasoning, and physical reasoning. Is this a fundamental principle of how intelligent systems are built, or an accident of biological evolution? In our latest preprint, we find that a similar modular organization emerges in Large Language Models, another class of intelligent system. Brains and LLMs are shaped by entirely different kinds of optimization (biological evolution vs. gradient descent). That they arrive at the same modular design anyway suggests modularity may be a fundamental property of intelligent systems. 🌐 Web: pengrui-han.github.io/LLM_Mo… 📄 Paper: pengrui-han.github.io/LLM_Mo… 💻 Code & data: github.com/Pengrui-Han/LLM_M… Using circuit analyses across 46 tasks spanning four cognitive domains, we find: 1️⃣ Tasks that draw on the same network in humans recruit overlapping units in LLMs, while tasks drawing on different networks recruit distinct units. 2️⃣ These units are causally linked to model behavior. Ablating the units critical for one domain impairs performance in that domain (−26% accuracy) but barely touches the others (−2.5%). This project has been in the works for a while :) Huge thanks to my advisors @jacobandreas @ev_fedorenko @devarda_a, and to @Nancy_Kanwisher for valuable conceptual input and feedback throughout. #MIT
48
429
2,329
228,062
Victor Fragoso 💻 retweeted
NeurIPS 2026 is recruiting Ethics Reviewers! If you have experience critically evaluating potential risks and harms in machine learning research, and can provide thoughtful feedback on broader impacts, please read our Call for Reviewers (blog.neurips.cc/2026/05/12/n…) and consider volunteering as an ethics reviewer. You can directly volunteer by filling the form here: forms.office.com/r/iXbzJ7dtx… Please share with your qualified and interested colleagues. As the scale of the conference continues to grow, so we are always seeking to grow our pool of ethics reviewers.
2
14
62
18,556
Victor Fragoso 💻 retweeted
It has been a blast sharing the excitement, science, art, and community of CVPR with all of you. The Publicity Chairs are officially signing off. Seattle, you’re up next! 🌲 See you at #CVPR2027! 👋 @deblinaforAI @anfurnari @CSProfKGD @YVinker @_vztu
3
35
302
19,508
Victor Fragoso 💻 retweeted
Announcing the official #CVPR2026 Best Paper Award Winner! 🏆 Congratulations to the authors for their landmark contributions to the field!👏👏
4
64
488
159,847
Victor Fragoso 💻 retweeted
This year, the NeurIPS 2026 Position Paper Track made the decision to require that all papers be substantially human-written, with AI used for only copy-editing or similar peripheral changes to the main text! For more details, please check our blogpost: blog.neurips.cc/2026/06/02/a…
16
61
406
149,706
Victor Fragoso 💻 retweeted
Releasing Echo-2 HQ, an improved model that delivers greater detail and sharper results. You can zoom in super close and discover remarkable appearance fidelity. Available via API and in the app. Try it out! Check out some of the scenes below👇
4
26
97
15,008
Victor Fragoso 💻 retweeted
As AI systems become increasingly capable, a fundamental question emerges: Can AI still learn from humans? Join us at the ReLearn Workshop @CVPR (June 3, Denver) to explore these questions with an outstanding lineup of speakers.
3
14
95
14,393
Victor Fragoso 💻 retweeted
We added Quake3-style multiplayer to our 3D world generator and it changes the game, literally 🔥🔥🔥 Runs directly on our model output: fast movement, arena combat, and every single match plays out on a generated world. Stay tuned for public release and play yourself!
4
23
89
16,287
Victor Fragoso 💻 retweeted
Excited to share Qwen-VLA paper, our exploration of generalist Vision-Language-Action models. It extends Qwen’s multimodal backbone from visual understanding and reasoning to continuous action generation and trajectory prediction. Paper: arxiv.org/pdf/2605.30280
7
113
592
78,357
Victor Fragoso 💻 retweeted
There have been a lot of fresh discussions lately around "bitter lessons". Continuing our tradition of community-building at CVPR, our workshop is back! This year's theme: Bitter Lessons in Computer Vision. Join us on Jun 3rd at 8:45 AM in Room 3A-3D at #CVPR2026 🔗 Website: sites.google.com/view/bitter… We have an incredible lineup to share their "bitter lessons": Bill Freeman, Alyosha Efros, @georgiagkioxari, @jon_barron, @vincesitzmann, @BharathHarihar3, @ShenlongWang, David Forsyth, @dimadamen @ev4n3sce, @CVPR
2
26
205
64,931