Ph.D. @Berkeley_AI. research intern @physical_int. ex @GoogleDeepMind, @MSFTResearch. vision, generative model, robotics.

Pinned Tweet
๐—ข๐—ป๐—ฒ ๐—บ๐—ฒ๐—บ๐—ผ๐—ฟ๐˜† ๐—ฐ๐—ฎ๐—ปโ€™๐˜ ๐—ฟ๐˜‚๐—น๐—ฒ ๐˜๐—ต๐—ฒ๐—บ ๐—ฎ๐—น๐—น. We present ๐—Ÿ๐—ผ๐—š๐—ฒ๐—ฅ, a new ๐—ต๐˜†๐—ฏ๐—ฟ๐—ถ๐—ฑ ๐—บ๐—ฒ๐—บ๐—ผ๐—ฟ๐˜† architecture for long-context geometric reconstruction. LoGeR enables stable reconstruction over up to ๐Ÿญ๐Ÿฌ๐—ธ ๐—ณ๐—ฟ๐—ฎ๐—บ๐—ฒ๐˜€ / ๐—ธ๐—ถ๐—น๐—ผ๐—บ๐—ฒ๐˜๐—ฒ๐—ฟ ๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ, with ๐—น๐—ถ๐—ป๐—ฒ๐—ฎ๐—ฟ-๐˜๐—ถ๐—บ๐—ฒ ๐˜€๐—ฐ๐—ฎ๐—น๐—ถ๐—ป๐—ด in sequence length, ๐—ณ๐˜‚๐—น๐—น๐˜† ๐—ณ๐—ฒ๐—ฒ๐—ฑ๐—ณ๐—ผ๐—ฟ๐˜„๐—ฎ๐—ฟ๐—ฑ inference, and ๐—ป๐—ผ ๐—ฝ๐—ผ๐˜€๐˜-๐—ผ๐—ฝ๐˜๐—ถ๐—บ๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป. Yet it matches or surpasses strong optimization-based pipelines. (1/5) @GoogleDeepMind @Berkeley_AI
62
437
3,371
563,716
Junyi Zhang retweeted
Our paper VisGym has been accepted to NeurIPS 2026! Itโ€™s been crazy to see how the field has developed since we have released this paper. Hope that VisGym can continue to serve as a diagnostic playground to identify failure pattern for new frontier multimodal agents, and serve for the development of this field.
We release VisGym: 17 environments for training & evaluating VLM interaction! โœจ Interaction is still the bottleneck of VLMs: Success rates: Gemini-3 26.0% | GPT-5 12.6%. Kudos to co-lead @zwcolin @junyi42 and the amazing team! @berkeley_ai
3
4
62
6,013
Junyi Zhang retweeted
I didn't get what GPT-6 meant for robotics until I actually tried it. GPT-6 Astra just does physical ICL out of the box. we drop a recording of a human doing a novel task into ๐—ฐ๐—ผ๐—ฑ๐—ฒ๐˜… app. Prompt it to drive a robot arm the same way. It just works on the first pass!
55
144
1,482
239,818
Junyi Zhang retweeted
Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8ร— NVIDIA B200 GPUs. Checkpoints + training/inference code + Technical Blog โฌ‡๏ธ (1/6)
70
171
1,217
440,334
"The latent world representation is a program, not a vector" Finding the right representation to be inverse-graphics is really important. Really cool work!
Today, weโ€™re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. ๐Ÿงต
4
3
49
7,856
Junyi Zhang retweeted
The term "continual learning" has become overloaded if you see it as an ML problem. One classic thread is about memorization: regularization-based continual learning methods, such as EWC, MAS, and SI, estimate which parameters mattered for previous tasks and resist changing them too much. One modern thread is about adaptation: test-time training and inference-time learning methods, such as TTT, adapt part of the model on the incoming test stream before making predictions. These are sometimes discussed as separate threads. But in modern scalable architectures, I think they are better seen as complementary constraints: a model that learns quickly at test time also benefits from a mechanism for deciding what not to forget. In our #ECCV2026 paper, we study this in large-scale 4D reconstruction: how to build fast spatial memory that can adapt over long observation streams while reducing collapse and forgetting. Instead of using fully plastic test-time updates, we stabilize fast-weight adaptation with an elastic prior that balances adaptation and memory. Key ideas: - Elastic Test-Time Training: Fisher-weighted consolidation for fast-weight updates - EMA anchor weights that provide a moving reference for stability - Chunk-by-chunk inference for long 3D/4D observation streams We show that this scales across large 3D/4D pretraining settings, including both LRM-style and LVSM-style models, and improves reconstruction across benchmarks including Stereo4D, NVIDIA, and DL3DV-140. We release model checkpoints across different design choices: resolution, post-training curriculum, and whether the model uses an explicit 4DGS intermediate representation. - Homepage: fast-spatial-memory.github.iโ€ฆ - Paper: arxiv.org/abs/2604.07350 - Code: github.com/Mars-tin/fast-spaโ€ฆ - Models: huggingface.co/marstin/fast-โ€ฆ This work is co-led with @Xueyang_Y, contributed by @zhnhoy5 @YuncongYY, and advised by @SLED_AI @gan_chuang.
3
24
115
34,044
Junyi Zhang retweeted
Great work using offline agentic exploration to develop robot skills!
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
1
10
68
23,155
Junyi Zhang retweeted
Learning from task-agnostic, explorative experience!
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
3
17
2,980
Junyi Zhang retweeted
Excited to share T-Rex: Tactile-Reactive Dexterous Manipulation ๐Ÿฆ–๐Ÿค– Touch is fundamental to human dexterity, yet most Vision-Language-Action (VLA) models either ignore tactile feedback or lack the ability to react to high-frequency contact signals. In this work, we tackle both the data and architectural challenges of tactile-reactive dexterous manipulation. ๐Ÿฆ– A 100-hour tactile-synchronized dexterous manipulation dataset with 7,700+ trajectories, 22 motor primitives, and 200+ everyday objects. ๐Ÿฆ– A tactile-reactive MoT architecture with spatial-temporal tactile encoding and asynchronous high-frequency tactile refinement. ๐Ÿฆ– A scalable training recipe combining 22,889 hours of human egocentric pretraining with tactile-grounded robot mid-training. Across 12 real-world contact-rich manipulation tasks, T-Rex achieves over 30% higher average success rate than the strongest baseline. We are fully open-sourcing the dataset, models, teleoperation stack, training code, and inference pipeline. ๐ŸŒ Project: tactile-rex.github.io/ ๐Ÿ“„ Paper: arxiv.org/abs/2606.17055 ๐Ÿ’ป Code: github.com/ZhuoyangLiu2005/Tโ€ฆ ๐Ÿค— Dataset: huggingface.co/datasets/zekaโ€ฆ ๐Ÿงต Thread โ†“
25
72
273
105,044
Junyi Zhang retweeted
The most inspiring thing I took from this paper: there's far more to squeeze from simulation than sim-to-real training of task-specific policies. RATs shows a coding agent can self-propose tasks, self-construct scenes in sim, and acquire skills that transfer to real-world deployment. It's promising to imagine handing coding agents a bunch of simulation clusters on top of ENPIRE to enable Sim-and-Real Co-research, where agents massively learn skills and try ideas in sim while continuously grounding them in the real world. Then robot skill acquisition can really scaling like everything else in the deep learning era.
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
5
9
97
11,846
Junyi Zhang retweeted
3 of 3: Kids can learn how to generalize via play (vs rote repetition of goal tasks) to learn skills that are useful for the future; we think agentic robotics should do so as well. We revisit curiosity-based intrinsic learning for agentic robotics:
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
2
3
37
4,983
Junyi Zhang retweeted
While ENPIRE w/ @nvidia @_wenlixiao @DrJimFan enables coding agents to explore algorithms and improve policies for a given real-world task, RATs asks: what can agents learn before a human specifies the task? Through curiosity-driven play, agents propose tasks, hill-climb toward solutions, and accumulate reusable, transferable skills. When a human later requests a new task, the agents retrieve and compose these skills to solve it. RATs explores an analogue of pre-training for embodied coding agents: broad skill acquisition through play, which accelerates task-specific problem solving with the skills acquired. Looking forward to the agentic future of robotics! See the detailed tweet from @junyi42!
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
2
8
39
7,743
Junyi Zhang retweeted
Excited to release our new work: Playful Agentic Robot Learning w/ @junyi42! Instead of relying on test-time scaling, we found that a "pretraining stage" through curious play enables robots to discover general skills before any tasks are assigned. ๐ŸŒ playful-rats.github.io
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
1
4
51
8,238
Junyi Zhang retweeted
What excites me most isnโ€™t just that we built an agentic coding system for robots and ran it in the real world. Itโ€™s that the agentic system learned a generalizable prior during "Play-Time", and then reused it to adapt across multiple downstream tasks.
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
2
6
34
4,786
Junyi Zhang retweeted
Frontier coding agents have shown they can run on real robots when given defined tasks. ๐Ÿค– Now we show these agents can learn the physical world like childrenโ€”no task required: give them curiosity and self-play, and real robotic skills emerge on their own.โœจ ๐Ÿ€๐‘๐€๐“๐ฌ are Robotics Agent Teams: embodied coding agents that learn through self-directed play before any downstream task is given. playful-rats.github.io/
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
1
12
1,914
Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with @jiaxin_ge_
10
64
350
96,461
These play-learned skills generalize across different simulations and directly transfer to the real world. Directly using the skill library learned in LIBERO, we get: RoboSuite (cross-environment): +8.9pp Real-world tasks: +8.8pp
1
6
1,577
๐‘๐€๐“๐ฌ is a first step toward ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ : ๐ŸŒplayful-rats.github.io We see a future where the next step for agentic robots isn't just stronger test-time harness, but a play stage where they set their own goals, fail, and build up skills long before we hand them a task. Huge thanks to the team: @lukehanjun (co-first) @letian_fu, Zihan Yang, Yaowei Liu, Raj Saravanan (core contributors), @istoica05 @akanazawa @JiahuiLei1998 @HavenFeng @trevordarrell and many others!
1
11
1,515