AgentSTAR is an agentic method for both reconstruction and tracking of articulated objects from monocular videos.
We can now track through rapid motion, severe occlusion and thin objects! There’s no spoon but we can still track it 🥄
Wrote done my thoughts on what's fundamentally new for vision when using agents --- obviously, apart from some hard problems that have now been nearly solved :)
AgentSTAR is an agentic method for both reconstruction and tracking of articulated objects from monocular videos.
We can now track through rapid motion, severe occlusion and thin objects! There’s no spoon but we can still track it 🥄
I particularly like this example, where we show our reconstruction alongside the ground-truth hand poses -- it gives a much better sense of the 3D tracking accuracy
In my mind, this line of thinking was popularised by Real2Code and SceneScript, but can now finally be taken to the extreme, where code is not only the representation but also what gets the tracking done
I'll be presenting our work at #ECCV2026 in Malmö! 🇸🇪 Find me at the poster and demo sessions:
Poster
📅 Thu 10 Sep, 10:30
📍 ExHall #272
Demo
📅 Sat 12 Sep, 8:00
📍 ExHall
Also at NeuSLAM Workshop 🤖:
Tue 8 Sep, 14:00, Malmö Arena Hotel
(1/3) I am happy to share our work on Triangle Splatting SLAM!
We show the first RGB-D SLAM to use differentiable triangles as a 3D map representation.
Our method enables online mesh-based deformations and collision checking via on-the-fly Delaunay triangulation.
Just announced: #CVPR2026 Best Demo Award for KV-Tracker. Congrats to Marwan and co-authors. And thanks to @deepfry_n, @shinjeong99 for massive help running the demo. That's 4 CVPR Best Demo Awards/Honourable Mentions in 3 years for our lab at Imperial! Proud of our demo culture.
How can we run reconstruction models like π³ and Depth Anything 3 in real-time?
We present KV-Tracker, a training-free approach, for real-time tracking of scenes and objects. Achieving up to 30 FPS!
With @alzugarayign, @makezur, @XinKong_IC and @AjdDavison
Introducing 4D Primitive-Mâché (4DPM), a new method for replayable 4D reconstruction from monocular videos.
We split dynamic scenes into 3D primitives and recover their motion. 4DPM can infer object positions even after they leave view.
Joint work with @marwan_ptr@AjdDavison
Introducing 4D Primitive-Mâché (4DPM), a new method for replayable 4D reconstruction from monocular videos.
We split dynamic scenes into 3D primitives and recover their motion. 4DPM can infer object positions even after they leave view.
Joint work with @marwan_ptr@AjdDavison
How can we run reconstruction models like π³ and Depth Anything 3 in real-time?
We present KV-Tracker, a training-free approach, for real-time tracking of scenes and objects. Achieving up to 30 FPS!
With @alzugarayign, @makezur, @XinKong_IC and @AjdDavison
Introducing 4D Primitive-Mâché (4DPM), a new method for replayable 4D reconstruction from monocular videos.
We split dynamic scenes into 3D primitives and recover their motion. 4DPM can infer object positions even after they leave view.
Joint work with @marwan_ptr@AjdDavison
4DPM handles multiple moving entities with distinct trajectories.
The result is a complete, replayable 4D reconstruction of all observed scene elements at every timestamp.
Our method cuts 3D primitives out of the outputs of a feed-forward reconstruction model (π³) and then glues them across time.
Each primitive’s motion is represented compactly as a single SE(3) pose, inferred from estimated 2D correspondences via an optimisation pipeline.