šŸ¤– Building Ambient Intelligence @SFResearch šŸ‘ØšŸ½ā€šŸ’» Former Intern @Apple MLR, @AIatMeta, @GoogleResearch, @mbzuai šŸ’ƒšŸ» Dancer in free time

NYC
[CVPR 2026] FOFPred has been accepted to #CVPR2026 (Findings)! We build a diffusion-based model that predicts Future Optical Flow from a single image guided by natural language instructions. Checkout code, model ckpt, & live demo at: fofpred.github.io
4
6
28
3,681
Kanchana Ranasinghe retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,370
5,666
49,218
6,047,906
Kanchana Ranasinghe retweeted
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface šŸ¤—
207
866
9,743
1,054,098
Kanchana Ranasinghe retweeted
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
4,049
8,314
76,586
39,904,683
Super cool!!
Astra made me a sonar app that emits undetectable audio to scroll up/down on your computer. It uses the doppler effect to determine where your hand placement is. You can even double tap in the air to change scroll directions!
1
2
344
Kanchana Ranasinghe retweeted
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia? These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach. The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date. The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant. At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on ā€œoldā€ problems that have a long history. This history is based on assumptions about how the ā€œvision problemā€ will be ā€œsolvedā€. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems. So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models. Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field. If we want there to be a ā€œfieldā€ of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section. Concretely, I think papers should include a new section analogous to ā€œRelated Workā€ where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models. I'm interested in your thoughts.
91
390
2,366
676,240
Kanchana Ranasinghe retweeted
Our analysis suggests that AI has so far created around 1m new jobs in America. We explain how the technology has created a hiring boom economist.com/finance-and-ec…
70
330
1,351
1,100,952
Very interesting take, and makes a lot of sense seeing recent improvements in general models.
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/wri… I think it's an important change in the trajectory of robotics!
191
Kanchana Ranasinghe retweeted
New blog post expanding on our thoughts around submission policies: blog.iclr.cc/2026/09/02/subm…
1
26
135
93,746
Kanchana Ranasinghe retweeted
Are you ready for Claudeforce?
47
80
933
238,493
Kanchana Ranasinghe retweeted
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: pollen-robotics.com/microduc… (video with sound on šŸ”Š)
590
916
8,191
2,898,798
Kanchana Ranasinghe retweeted
After so many demos, models, pilots, robots are still struggling to land real deployments with real customers. Until now. Today we’re excited to share that our robots have successfully crossed the ROI threshold, and Din Tai Fung, one of the highest revenue per location restaurant chain in the US, is rolling out Dyna robots across its extensive restaurant network. This brings our rollouts across hotels, logistics, data centers, and many other use cases to a fleet that reaches hundreds of robots by the first half of 2027. And we’re just getting started. It’s been a wild year, and today we’re double clicking on the battlefield stories and sharing a few learnings about scaling robot deployments. We are just scratching the surface. Read the full blog post: dyna.co/news/scaling-custome…
54
141
1,088
341,749
Kanchana Ranasinghe retweeted
Honored to be chosen for the TIME 100 influential people in AI. A year ago, I was fighting tooth and nail to convince people to pay attention to AI detection. We had made huge technical breakthroughs, but it felt like nobody cared. Looking toward the future, it’s a completely different story. As AI becomes more powerful and autonomous, I believe that human touch will be the one of the last remaining scarce resources, and Pangram has an incredible opportunity to help the world place value in humanity. I am grateful to be able to work with such an incredible team, and can't wait to show you all what's next.
Max Spero is one of the hundred most influential people in artificial intelligence.
108
26
991
60,216
Kanchana Ranasinghe retweeted
Claudeforce is here. āš”ļø The #1 AI (Claude) now runs natively on the #1 CRM (Salesforce). Through the new AIforce harness + Headless 360, Claude gets direct, governed access to Data 360, Tableau, Slack, and your entire Salesforce workflow—without ever leaving the chat. What this changes today: • Instant Grounded Intelligence — Ask complex enterprise questions and get real-time, trusted answers • App & Agent Builder — Deploy custom workflows, agents, and secure apps on the fly • Action-Oriented — Trigger live enterprise actions and get work done, not just summaries • Ironclad Governance — Zero Data Retention, full trust boundary, Salesforce-certified by default Probabilistic models alone can’t run a company. Deterministic systems alone can’t reason. Claudeforce fuses Claude’s reasoning with Salesforce’s trusted data and controls. The AI is the interface. This is how every business will run. See it at Dreamforce. #DF26
219
439
3,091
1,201,932
Kanchana Ranasinghe retweeted
As foretold in the prophecy
Did you know Elon Musk was weirdly "predicted" in a Wernher von Braun book nearly 80 years ago? In his 1948 novel *Project Mars*, von Braun describes a future Martian government led by a figure called the "Elon." It's not a person's name, but a title for the head of government. Still, considering Elon Musk's lifelong obsession with Mars, this has to be one of the strangest coincidences in space history!
4,303
7,418
57,680
7,258,244
Kanchana Ranasinghe retweeted
There are cathedrals everywhere with those with eyes to see.
Introducing terminal-code: VS Code inside the terminal - VS Code compatible CLI - works over ssh - syncs with your terminal theme terminal-code.com
2
3
40
10,735
Kanchana Ranasinghe retweeted
Don’t code alone. Slack Code is live. Humans and agents. Same channel. Same work. Launching today with agents from @AnthropicAI, @github, @Cognition, and @vercel. This is real multiplayer coding. See it at @Dreamforce #DF26
305
402
4,287
2,273,028
Amazing how fast video gen has evolved!
We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000. Starring @N3onOnYT, @stylebender, @RampageJackson, and @MKIATPIS It's 100% open-sourced on Higgsfield: all prompts and assets are public now. Made with Seedance on Higgsfield. Watch full film below šŸ‘‡
128
Kanchana Ranasinghe retweeted
For the socialists dreaming of the Scandinavian model, that means *cutting* taxes on the rich in half. While hiking them *3 times* on everybody else until they pay half their salary šŸ™Œ
152
1,123
5,186
197,827
Kanchana Ranasinghe retweeted
Catching skin cancer early is a home robotics problem. Melanoma is highly treatable when detected early, yet today’s screening process depends heavily on patients noticing tiny changes across their entire skin surface. This requires patients to solve a near-impossible visual-memory and registration problem. I built OpenDerm, an open-source 4-DOF robot that captures high-resolution images of the skin and uses them to reconstruct and track the skin surface in 3D over time. The best way to make skin screening truly routine is to bring it into the home. OpenDerm shows that inexpensive robotic skin imaging is possible, but the path to scale is not a dedicated screening robot in every household—it is to make skin screening one of the many useful things a general-purpose home robot can do. Read more about why I built OpenDerm and how it works here: Blog: marionlepert.github.io/blog/… Project: openderm.github.io/
416
977
9,064
1,505,754