3D gets another shot
The first arrival tried to create demand by putting new displays on our faces. The second may create demand by putting agents inside the tools.
I lived through the first one.
Around 2016 and 2017, the industry had a fairly coherent thesis. Oculus and Vive had shipped. Mobile VR was everywhere. ARKit and ARCore were about to put spatial computing onto hundreds of millions of phones.
If computing was becoming spatial, the world would need an absurd amount of 3D. Products, environments, avatars, materials, animations, simulations. Everything around us would eventually need some digital representation.
So the chain looked obvious:
VR + AR → demand for 3D content → demand for 3D tools → 3D becomes a foundational computing primitive.
At @scapic, we believed this deeply. We were unusually heavy Blender users for 2017/18, back when Blender still felt like software that required a minor apprenticeship before it would let you make anything useful. I also spent time around glTF, Khronos and the broader effort to make 3D portable enough for the web.
That work mattered because the entire thesis depended on 3D becoming infrastructure. Geometry had to move between applications. Materials and animations had to survive. Assets had to become small enough to download. Renderers had to roughly agree on what an object looked like.
The web needed something closer to JPEG or MP4 for 3D. glTF was part of that answer.
It felt like we were laying tracks just before the train arrived.
The train was much smaller than expected.
The great 3D ice age
The strange thing about the next decade is that 3D became extremely important without becoming everyday software.
Games grew. VFX grew. CAD remained fundamental. Robotics needed simulation. NVIDIA pushed deeper into graphics, simulation and digital worlds.
3D was everywhere underneath technology.
Very few normal people created 3D scenes.
The reason becomes obvious once you look at the authoring stack. A useful 3D asset can involve geometry, topology, UVs, textures, materials, normals, rigging, animation, lighting, collisions, LODs and export settings.
Sometimes one shoe was a small software project.
Then came the interfaces: modeling, sculpting, texture painting, shader graphs, rigging, animation, lighting, simulation, rendering.
Every layer had its own vocabulary and dozens of degrees of freedom.
We spent the first wave focused heavily on demand: better headsets, more devices, larger installed bases, better distribution.
The supply problem was just as important.
Even if every household owned a headset, somebody still had to make all the 3D.
Formats like glTF made the finished object easier to distribute. Someone still had to create the finished object.
That became the great 3D ice age.
Blender survived it
Which brings me back to Blender.
Looking at Blender today after using it heavily almost a decade ago is slightly absurd. It just kept compounding.
Modeling. Sculpting. Animation. Geometry nodes. Simulation. Rendering. Materials. Cameras. Asset systems. Compositing. Python.
Blender has become something closer to an operating system for manipulating 3D state.
And all the complexity that made Blender intimidating suddenly becomes interesting in an agent world.
Agents like deterministic tools with inspectable state.
A Blender scene has objects, geometry, transforms, materials, cameras, lights, constraints, animation and physics. Much of it can be inspected and manipulated programmatically.
That changes the product question.
The next iteration of Blender-like software may have less to do with simplifying every menu for a human.
It may be about humans no longer having to personally operate every menu.
The old loop was:
human learns tool → manipulates scene → renders → inspects → repeats.
The new loop can become:
human states intent → agent inspects scene → manipulates tools → renders → evaluates → human directs the next move.
“Make the lamp 20% shorter. Keep the silhouette. Move the key light left. Try a 50mm lens. Give me three camera positions. Keep the second.”
Every sentence hides several expert operations.
That is the point.
3D becomes playful again.
Generation is the smaller idea
This is why I find “text-to-3D” slightly too narrow as the framing.
Generating a mesh is one step.
A real workflow looks more like:
prompt → geometry → cleanup → topology → UVs → materials → rigging → placement → lighting → simulation → rendering → export.
Some of those steps benefit from generative models. Many benefit from deterministic software.
Put them together and something more interesting happens.
A model can propose. A scene graph can hold state. A geometry engine can transform. A physics engine can test. A renderer can observe. An agent can orchestrate the loop.
The useful unit becomes the workflow.
This starts extending into physical products too.
Imagine saying:
“Design a flat-pack stool. 18mm plywood. CNC manufacturable. Support 120kg. No visible front fasteners. Fit inside this carton.”
An agent can generate variants, test constraints, run simulations, check clearances, render options, alter geometry and prepare manufacturing files.
Atoms still have their own clock. Materials, factories, tolerances and supply chains remain stubbornly physical.
But the design loop can start behaving more like software.
Run. Inspect. Modify. Test. Repeat.
That compression alone is enormous.
3D as infrastructure for AI
There is another angle I think is even more underexplored.
3D software may become infrastructure for generative models themselves.
Take video.
Video models have extraordinary visual capacity with comparatively weak deterministic control. Language is still an imprecise way to specify camera position, geometry, blocking, focal length and motion.
A 3D environment gives you explicit control over all of those.
So a 3D scene can become the control rig for a video model.
An agent builds the scene, positions subjects, selects the lens, creates the camera path, renders reference frames, depth or normals, then hands that structure into a generative model.
Camera direction starts becoming software.
The same applies to world models.
Pixels are observations.
A scene graph is state.
Inside a programmable world you can know what an object is, where it is, how it relates to another object, what collided, where the camera moved and what changed after an action.
I increasingly think of the scene graph as something like the DOM for worlds.
A screenshot tells an agent what a webpage looked like. The DOM tells it what exists and what can change.
A render shows what a world looked like. A scene graph tells the agent what the world contains and what can be manipulated.
That feels like an important substrate for spatial reasoning, simulation, training and agent evaluation.
This is also why I found agentic 3D like Mixar interesting.
Mixar is essentially exploring this interface directly: Blender underneath, an agent operating inside the environment, and generative tooling composed with deterministic 3D operations.
Good to see them launch:
github.com/Mixar-AI/mixar-ap….
The interesting piece for me is the architecture rather than the chat interface. Agent instructions can become scripts operating against Blender's Python API. The application routes those operations into Blender's execution environment and tracks what changed in the scene.
That feels closer to an agent having a 3D computer than an AI assistant sitting beside a creative application.
And open source becomes particularly important in this world.
Computer use gives an agent a mouse.
Open software gives it the machine room.
Mixar is one example, but I expect the category to get much larger: agent-native 3D environments spanning Blender-like tools, CAD, simulation, manufacturing and entirely new interfaces.
The architecture will rhyme:
structured world state + deterministic tools + generative models + simulation + an agent operating the loop.
Looking back, I think we were directionally right during the first 3D boom.
3D did become foundational infrastructure.
We were early about the interface that would make it accessible.
We thought VR and AR would force people to learn 3D because everyone would suddenly need spatial content.
Instead, Blender kept compounding. glTF made 3D easier to move. USD pushed scene composition forward. GPUs got absurdly fast. Physics improved. Generative models arrived.
And now agents can operate computers.
All that infrastructure left behind by the first wave suddenly has a new interface.
The first arrival of 3D put virtual worlds in front of humans.
The second may put agents inside those worlds.
We spent the first boom teaching humans how to operate 3D software.
This one may be about teaching software how to operate 3D.
Then the rest of us get to play.