3D Was Never an Interface

For most of the last decade, three dimensions were sold as somewhere to go. The headset, the metaverse, the immersive showroom, the spatial computer. Every one of those propositions rested on the same assumption: that the value of a three-dimensional model lay in a human being inside it, looking around. Tens of billions of dollars were committed to that assumption. It was wrong, and it was wrong in an instructive way.
The assumption made 3D a cost centre. Someone had to build the model. Someone had to wear the device. Someone had to be present, and paying attention, at the moment of value creation. A representation that only pays out when observed is not infrastructure. It is media.
What is happening now is a reclassification. Three dimensions are being moved out of the display column and into the data column, and the shift is being funded at a scale that makes it hard to dismiss as a cycle.
The capital has already moved
Fei-Fei Li, who built ImageNet and therefore has some standing on what constitutes a substrate, published an argument in late 2025 that has become the reference text.1 Her structural claim is precise: where language models learn the statistical structure of text, world models learn the statistical structure of space and time, including how light falls on a surface and how objects respond to force. She calls the target spatial intelligence, and frames it as the frontier beyond language.
This is not a lone position. Yann LeCun left Meta and raised $1.03 billion at a $3.5 billion pre-money valuation in March 2026 for AMI Labs, on the thesis that AI should learn from reality rather than only from language. World Labs closed its own billion-dollar round in February 2026, with Nvidia and AMD participating and Autodesk writing a $200 million strategic check.2 That last detail is the one worth pausing on. A company whose entire business is the geometry of physical things bought a seat at a frontier model lab. It did not buy a rendering feature.
Nvidia has been making the same argument in industrial language for longer. Jensen Huang's framing is that factories are ceasing to be static physical assets and becoming living systems that are designed, simulated and operated as virtual twins, and that everything that moves will eventually be embodied by AI. Strip the vendor vocabulary and the claim is modest: if a system is going to act in physical space, it needs a representation of physical space that supports consequence, not just appearance.

Geometry as evidence
The useful inversion is this. In the interface era, 3D was the output. A model was built so that it could be seen. In the current reclassification, 3D is the input. It is the representation in which a system can reason about physical consequence at all.
This matters because geometry has a property that text does not. Text about a space is testimony. It can be optimistic, stale, or self-serving, and nothing about the format resists that. Geometry is evidence. A point cloud is a falsifiable claim about where matter was at a particular moment. It does not have a point of view. When the reported state of the world and the measured state of the world disagree, the disagreement is legible, and it is legible without anyone deciding to look.
There is also a supply argument that has become impossible to ignore. Text has been scraped. The public corpus of written language is a known and largely exhausted quantity, which is why so much of the current research effort goes into synthetic data and inference-time compute rather than more reading. Physical space has not been collected. Most of the economy operates in rooms, yards, floors, corridors, and shelves that have never been instrumented in any form a model could learn from. The built environment is the largest untrained dataset in the world, and it is legible only in three dimensions.
The architects got there first
The current wave presents itself as new. It is not, and the amnesia is worth correcting, because the earlier attempts failed for a reason that has now been removed.
Nicholas Negroponte published The Architecture Machine in 1970, five years before Soft Architecture Machines. He was an architect at MIT running a lab dedicated to machines that would design with you rather than for you, and that lab became the Media Lab. Gordon Pask had already written "The Architectural Relevance of Cybernetics" in 1969, arguing that a building should be understood as a conversational system that adapts to its occupants. Cedric Price and John Frazer went furthest with the Generator, designed between 1976 and 1979: a reconfigurable structure whose controlling program would register that nothing had changed for too long and propose alterations on its own initiative. A building that got bored.
Christopher Alexander, in Notes on the Synthesis of Form and later A Pattern Language, attempted to make spatial logic composable, and succeeded so well that software borrowed the idea wholesale and turned it into design patterns. Herbert Simon, who won both a Turing Award and a Nobel Prize, wrote The Sciences of the Artificial in 1969 and remains the only figure who properly belongs to both halves of this subject.
None of these people lacked the idea. They lacked the compute, the sensing, and the data. All three now exist. The intellectual debt has simply gone unacknowledged, which is why so much current work rediscovers, at considerable expense, conclusions that were published in architectural journals fifty years ago.
Space was the first cognition
There is a deeper reason the spatial turn is not a modality expansion but a return to the substrate.
J.J. Gibson's ecological account of perception, published in 1979, proposed that organisms do not perceive objects and then infer uses. They perceive affordances directly: possibilities for action offered by the environment. An affordance is simultaneously a fact about geometry and a fact about cognition, which makes it the natural unit for any system meant to reason about space and behaviour together.
The neuroscience has since caught up. Edward Tolman proposed cognitive maps in 1948, over the objections of behaviourism. O'Keefe and the Mosers took the 2014 Nobel for place cells and grid cells, and subsequent work has shown that the same hippocampal machinery maps abstract conceptual spaces, not only physical ones. Human beings appear to run general reasoning on hardware that evolved for navigation. The residue is visible in ordinary language, which is spatial almost everywhere it is abstract: closer to the truth, above my pay grade, behind schedule, a field of study, a high point.
Language, on this account, is the late arrival. Treating spatial capability as something to be bolted onto a language model has the order of evolution exactly backwards.
The part that is still unearned
The honest assessment of the current moment is that generation is running well ahead of judgement.
The demonstrations are genuinely impressive. Text prompts produce explorable environments; real-time engines hold a world stable as you move through it; simulators run millions of physical trials at zero marginal cost. What almost none of this addresses is the harder half of the problem. A system that produces a thousand plausible configurations of a space is worth very little without a function that can say which one works, and why, and under what conditions it stops working.
Evaluation, not generation, is the binding constraint. This has always been true of design practice, where the brief precedes the sketch and the constraint precedes the form, and where the discipline of the work lies in the criteria rather than the output. Evaluative pressure is what drives generation. Generation that runs without it produces volume, not value.
Which is why the real proving grounds are unglamorous. The places where this technology will be judged are warehouses, shop floors, hospital wards, substations, construction sites, and distribution centres: environments where being wrong is expensive, where the ground truth is geometric, and where nobody is going to put on a headset. These are precisely the parts of the economy that language models cannot see.
The original digital twin makes the point better than any current marketing does. NASA kept a full physical duplicate of the spacecraft on the ground, so that a fix could be attempted somewhere cheap before it was attempted somewhere fatal. Michael Grieves formalised the concept for product lifecycle management in 2002. The purpose was never visualisation. It was rehearsal.
Architectural drawings have always been instructions, a set of commands issued to people who would execute them later, in a different place, at a different time. What is changing is not the nature of the document. It is who reads it.
Footnotes
-
Substack (Dr. Fei-Fei Li), From Words to Worlds: Spatial Intelligence is AI's Next Frontier (2025-11-10) — primary source. "The pursuit of visual and spatial intelligence has been the North Star guiding me since I entered the field. It's why I spent years building ImageNet, the first large-scale visual learning and benchmarking dataset and one of three key elements enabling the birth of modern AI, along with neural network algorithms and modern compute like graphics processing units (GPUs)." ↩
-
Artificial intelligence model developer World Labs Inc. today announced that it has raised $1 billion in funding. The capital was provided by a consortium that included Nvidia Corp., Advanced Micro Devices Inc., Autodesk Inc. and several others. The engineering software maker invested $200 million. SiliconANGLE, World Labs closes $1B investment backed by Nvidia, AMD and Autodesk (2026-02-18). This factual passage directly confirms that World Labs raised $1 billion in February 2026 with participation from Nvidia, AMD, and a $200 million investment from Autodesk, directly substantiating the author's claim regarding the timing, size, and corporate participants of the funding round. ↩