Recording paper notes by @chongzzzhang who is interested in robot learning

busy with paperwork recently
arxiv.org/pdf/2609.08224 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints two view wapoints -> triangulation -> better spatial representation for VLM to output after finetuning
5
673
TANGO : Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model arxiv.org/pdf/2609.09158 VLA motion generation + sonic tracker. Dataset includes whole body motions. A small step towards WBC + VLN.
14
870
arxiv.org/pdf/2609.08209 Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation A dataset with paired human demo and robot simulation data. Good idea for scene understanding in manipulation.
5
678
C's Robotics Paper Notes retweeted
arxiv.org/pdf/2609.08511v1 PGMT is probably the best tracking paper I've seen. It can walk on stairs or a 40cm box while tracking ref, without modifying the flat ground reference. No reference-terrain adaptation, cheap training. The model arch design is the key and is very smart.
3
3
63
3,883
C's Robotics Paper Notes retweeted
New paper release: Learning Agile Perceptive Traversal of Sparse 3D Structures for Humanoids Paper arxiv.org/pdf/2608.29769 Videos nemantor.github.io/sparse-3d… We study how to do learning and sim2real for 3d traversal, such as monkey bars and overhanging obstacles.
12
40
267
82,645
arxiv.org/pdf/2608.18234v1 gigabrain wbc 0.5 1. Pair reference with terrain primitives to afford contacts 2. Enabling motion tracking on terrains/object contacts via such data 3. Use next state and latent prediction, plus GMM check to filter out unsafe OOD commands.
1
4
38
2,610
back to life. first paper to update here is adept. adept-dexterity.github.io/ 1) pretrain multiple skills 2) BC-RL fnetuning 3) Distill a visual student Achieves skillful precise manipulation via sim2real
1
25
2,174
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation arxiv.org/pdf/2603.03279 1. RL based retargetting 2. multi-modal command student distillation and finetuning so that it can switch between goal reaching vs reference tracking, mocap vs depth
7
30
2,368
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations arxiv.org/pdf/2603.03243 for cross-embodiement transfer: 1. use look-at points in 3D space instead of direct head states 2. mask out arms
7
35
2,171
Watch Your Step: Learning Semantically-Guided Locomotion in Cluttered Environment arxiv.org/pdf/2603.02657 this works shows you can train a policy plus using a semantic map to avoid stepping on valuable things. this is why I believe locomotion should all be mapping based.
3
28
1,873
arxiv.org/pdf/2602.21723 LESSMIMIC: Long-Horizon Humanoid Interaction with Unified Distance Field Representations use distance field (DF) as a representation for HOI. each link's traj can be defined by DF + grad of DF + vel_norm + vel_tangent. (tbc)
1
4
48
2,470
Training is multi stage. They first have a mimic policy and distill that to a base policy. Then they use AIP (AMP for interactions) to make the policy generalize instead of memorizing kinematic references. DF needs mocap, so they also distill this into vision policies.
2
459
arxiv.org/pdf/2603.01126v1 Pro-HOI: Perceptive Root-guided Humanoid-Object Interaction trained with mimic + contact commands, but deployed with planner to replace the reference.
4
55
2,623
arxiv.org/pdf/2602.18164 The grandtour dataset, with tons of odometry methods benchmarked in tons of environments.
2
29
1,883
hero-humanoid.github.io/stat… Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation For humanoid EE goal reaching, on top of WBC pose tracking, a neural model is trained to further correct small biases near the goal.
1
9
53
2,661
VIGOR: Visual Goal-In-Context Inference for Unified Humanoid Fall Safety arxiv.org/pdf/2602.16511 visual humanoid fall recovery 1) flat terrain human sparse reference 2) terrain-aware ref adjustment 3) keypoint tracking teacher 4) distilled to visual student without keypoint obs
3
17
1,419
arxiv.org/pdf/2602.15733v1 meshmimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction kinda like CRISP + OmniRetarget, but: 1) uses human edge and excludes depth edge in image to extract contacts, which is smart; 2) use polygonal primitives for clean scenes
11
78
4,717
arxiv.org/abs/2602.11758 Humanoid Agile Object Interaction Control via Dynamics-Aware World Model predict object states from proprioceptive obs, use it to transform object template pointcloud, and embed the transformed pcl as policy inputs.
2
6
31
2,161
Kids today are not citing wococo :(
1
1
575
apex-humanoid.github.io/ reward engineering + state machine engineering + nice progress-based rewards for reference-free humanoid agile climbing learning.
1
11
64
10,913