Most 3D scene graph methods rely on complex pipelines that break down without test-time ground-truth labels. GraphWrit3R strips away those multi-stage dependencies: Feeds raw point clouds or 3D Gaussian splats directly into a single model Outputs structured JSON scene graphs without intermediate bounding box steps Runs entirely end-to-end with zero reliance on ground-truth annotations during inference
2
1
9
446
By supporting both point clouds and 3D Gaussian splats, the architecture handles unstructured 3D representations directly. It maps physical geometry to spatial and functional object relationships without relying on heuristic post-processing. This gives robotics and embodied AI agents a clean route from raw sensor data to structured scene understanding.
1
2
88
The offline benchmark results completely back up this new architecture. In two step planning evaluations, X Planner scored a massive 0.9011 BERTScore F1. It directly outperformed established models like Qwen and Doubao in both semantic text matching and judge rated plan quality, proving its superior reasoning capabilities.
1
2
67
To make this framework incredibly fast, the team developed a method called Staircase Decoding. Rather than waiting on standard autoregressive generation, the architecture relays hidden states across staggered Transformer depths. Latent plan states process in parallel through the upper layers, resulting in massive computational efficiency without sacrificing logic.
1
3
122
Long horizon manipulation has always been a massive bottleneck for end to end robotic models. X Square Robot just changed that with X Planner, a 9 billion parameter event structured task planning front end. It perfectly bridges the gap between high level instructions and executable control actions. By leveraging a custom dataset of 1500 training episodes, it is already hitting top tier benchmark scores and making long term physical reasoning a reality.
1
2
8
778
The secret to this performance is how X Planner structures events. Instead of spitting out a black box chunk of actions, every single step in the plan is meaningful in language, observable in video, and physically realizable by the hardware. This approach completely opens up the middle ground of robotic reasoning and makes debugging transparent.
1
1
2
122
It also started doing in-context learning, solving new tasks with only a few examples given in the prompt. Separately, the generator stumbled onto Fibonacci and other number sequences on its own, it didn't know they existed.
1
2
134
They tested this across a genuinely wide spread of data: DCLM web text, GitHub Python, C source, Metamath proofs, CIFAR-10 images, speech commands, music, even human DNA. Self-play's scaling rate came out comparable to models trained directly on that real data, in some cases even a bit faster. DNA was the outlier, scaling much slower than real DNA-specific models.
1
2
142
Paper: arxiv.org/pdf/2609.30063v1 this website is just so good for you to visualise and understand this brilliant method : lcrh.github.io/self-play/
2
134
What if we trained a LLM with zero real data? No internet text, no images, nothing. Just two neural nets playing a game against each other from scratch. It still gets predictably better at predicting real text, images, and music it's never seen. The setup is just so elegant a "generator" writes tiny programs for a Universal Turing machine(BrainFuck'd) running them produces byte sequences. A "learner" tries to predict those bytes. The generator gets rewarded for writing programs that sit right at the edge of what the learner can currently handle.
2
1
24
1,314
As they scale up compute on this loop, loss on real held-out data like text, images, speech, DNA, and melodies drops in a clean power law, comparable to training on the real thing directly. The model never touched any of it.
1
1
1
165
CWM is compared against Dreamer and Momentum Prediction, two different ways of learning the world model. In the default setting, they’re comparable but once visual distractors and natural video backgrounds are introduced, CWM substantially outperforms both.
1
2
217
Do world models really need to predict pixels to understand the future? Traditional world models learn the future by reconstructing observations. This forces the model to learn everything in the observation ,  including visual details that are irrelevant or unpredictable. Contrastive World Models (CWM) asks a simple question: can we model the future directly in latent space, without reconstructing pixels? CWM replaces pixel reconstruction with a Deep InfoMax objective
2
4
40
2,364
CWM maximizes mutual information between the latent state and local features of future observations. The idea is to keep the information that helps predict the future, while discarding irrelevant visual details. The pipeline becomes: Current observation → Useful representation → Identify the correct future No decoder, no pixel reconstruction and minimal assumptions.
1
3
206
Because it builds on the multimodal pretraining behind the broader FLUX 3 family, it deeply understands the physical world. The model learns critical representations of motion, contact, object behavior, and cause and effect directly from large scale video data. It predicts future video frames and actions together, meaning it anticipates exactly how a scene will visually evolve before making a move.
1
4
178
Beyond robotics, FLUX 3 Action showed promising early results when we trained task-specific policies for playing games and controlling a simulated drone, suggesting the same approach could apply wherever a model needs to understand a visual environment and choose what to do next (e.g. simulations, games, and computer use).
1
4
118