Mapping the Landscape of Spatial Intelligence
Where do 3DGS, spatial computing, robotics, and world models fit in spatial intelligence?

In 1946, Jorge Luis Borges wrote a short fable about an Empire obsessed with cartography. Its guild of mapmakers, dissatisfied with every map that was made before, finally created one at the scale of the Empire itself, with every detail displayed in perfect correspondence. Later generations, however, judged this colossal map to be completely useless and eventually abandoned it.
In the Deserts of the West, still today, there are Tattered Ruins of that Map, inhabited by Animals and Beggars; in all the Land there is no other Relic of the Disciplines of Geography.— On Exactitude in Science. Jorge Luis Borges, Collected Fictions, translated by Andrew Hurley.
I initially read this fable in Jean Baudrillard's Simulacra and Simulation and it has lingered in my head ever since. If the map was the Empire's grandest technology, we are now building things of a very similar kind: Gaussian splats of streets and towns... Digital replicas spanning the spectrum from loose abstraction to point-for-point matching with the physical world.
From a product perspective, the map failed because it was a copy that served no one (except perhaps the Emperor) and no purpose beyond its own exactitude.
So one might ask: if perfect fidelity per se is insufficient, what makes a 'map' useful? What should we do differently to ensure that the spatial technologies we are building today help us solve real problems, rather than becoming decayed silicon ruins?
I believe the answer lies in Spatial Intelligence.
The people building today's grandest spatial technologies begin with a shared recognition of the limitations of language-based AI. John Hanke, Founder and CEO of Niantic Spatial, calls AI "trapped inside the screen," deeply knowledgeable about text but largely ignorant of the physical world. Similarly, Fei-Fei Li, CEO of World Labs, describes current LLMs as "wordsmiths in the dark; eloquent but inexperienced, knowledgeable but ungrounded."
These observations alongside Borges' fable point to limitations on both sides. The Empire’s map was a perfect spatial representation of the world but lacked the intelligence capable of interpreting for utility. Today’s language-based AI has intelligence but is weakly grounded in the spatial world. Spatial intelligence begins to bring these two capabilities together.
Spatial Intelligence represents the frontier beyond language—the capability that links imagination, perception and action, and opens possibilities for machines to truly enhance human life, from healthcare to creativity, from scientific discovery to everyday assistance.— Fei-Fei Li
But what does building it actually involve? Much of the recent attention has shifted towards world models, physical AI, as well as smart glasses and robotics. Yet how these pieces relate to each other has not been clearly illustrated. This essay proposes a framework for understanding these different technologies and use cases as part of one interconnected landscape of spatial intelligence. Inspired by and written for my fellow XR & AI builders, I also want to demonstrate where our work sits in this landscape, and how it can contribute to the field.
The First Dimension
The framework is organised around two key questions:
- What relationship does a spatial representation have to a particular physical reality?
- What roles do humans and intelligent systems play in closing the perception–action loop?
The first dimension asks what a spatial representation refers to in physical reality. Along this dimension, there are three modes:
- Indexed: The spatial representation is anchored to a particular existing place, scene, or object. It may be as abstract as a floor plan or as detailed as a photorealistic reconstruction. It may be overlaid directly onto the physical world or reconstructed digitally with high fidelity.
- Variant: Spatial representations become malleable and editable. Conditions, structures, behaviours, or possible outcomes can be altered to explore counterfactual versions of a particular physical reality.
- Invented: The spatial representation is no longer anchored to any particular physical reality. It can be generated from language, artistic references, or learned world knowledge.
A filmmaker might use XGRIDS PortalCam to capture an existing location and reconstruct it as a navigable 3D Gaussian Splat for creative productions (Indexed). In Variant mode, World Labs' real-to-sim pipeline reconstructs a physical robot, its environment, objects, and task demonstrations as an aligned simulation, then systematically varies object configurations, physics, and other parameters to create thousands of simulated variations from one real task. Similarly, the Waymo World Model, built on Google DeepMind's general-purpose world model Genie 3, can take recorded driving environments and alter road layouts, road users’ behaviours, traffic signals, weather, or time of day for counterfactual simulation. And with World Labs’ Marble, creators can use text or concept images as prompts to generate imagined worlds with no specific real-world counterpart (Invented).
The Second Dimension
The second dimension addresses the role humans and intelligent systems play in the perception-action loop. Software agents already raise the question of who determines what should happen next. Physical agents add another layer: who carries out that decision in the physical world? Along this dimension, there are three modes:
- Representational: Spatial intelligence is used to represent or simulate an environment for viewing and exploring. It may be manipulated and inform subsequent decisions. Any physical action happens outside the system. The loop stays open.
- Human-mediated: Within this system, the person may act directly, or communicate goals and constraints for an embodied system to execute. The person owns the decision authority over whether and what to act on. The loop closes through a person.
- Autonomous: The intelligent system determines and executes the physical action through an embodied machine, closing the perception–action loop itself. This is where NVIDIA's definition of physical AI sits: enabling systems such as robots and self-driving cars (e.g., the Waymo Driver) to "perceive, understand, reason, and perform or orchestrate complex actions in the physical world."
While the autonomous mode reflects the current trending discourse around robotics, I find the human-mediated mode the most interesting as it explores a new direction: how might people understand, direct and coordinate action with physical AI? This is where XR interfaces and devices can play a crucial role.
XR can situate spatial intelligence directly within the physical environment, creating a shared spatial interface through which machine understanding becomes legible to the human and human intent legible to intelligent systems. This takes two forms:
- The person acts directly. Smart glasses, for instance, can perceive and interpret the environment from a first-person perspective. The indexed spatial understanding can provide guidance to the wearer (identifying an object, anchoring information in space, or reasoning about the next step in a physical task).
- The person directs an embodied system. Johannes Tscharn visualizes robot data such as LiDAR through Snap Spectacles and demonstrates directing a Unitree Go2 by pointing to a spatial target in AR. VectAR, created by Pavlo Tkachenko and Stijn Spanhove, similarly connected Snap Spectacles with an Anki Vector robot, aligning the robot’s SLAM map and live sensor data with the human view to create an AR air-hockey game.
The Interconnected Field
Taken together, the two dimensions form the landscape of spatial intelligence as an interconnected field of representations, simulations, human interaction, and embodied systems.
The modes describe different relationships between spatial representations, physical reality, humans, and intelligent systems. A system may occupy one position, span several, or move between them over time. For example, the boundary between Human-mediated and Autonomous is especially fluid: an embodied system may act independently under normal conditions, then return decision authority to a person when it encounters uncertainty, requires approval, or receives new spatial direction.
World Labs' R2S2R engine provides another example. A physical robot, environment, and task provide the real-world grounding; these are reconstructed in simulation, where object configurations, physics, and other conditions can be varied at scale. Policies learned across these variations are then transferred onto physical hardware, where the robot can act autonomously. This pipeline shows how representations can become environments for generating possibilities, learning behaviours, and ultimately informing physical action.
Seen this way, world models, robotics, and XR all contribute different capabilities to shaping the landscape. World models provide ways to represent, transform, simulate, and generate worlds. Robotics and physical AI extend spatial intelligence into non-human physical agency. XR and spatial computing give people perceptual and interactive access to both virtual worlds and these systems, allowing them to perceive, navigate, manipulate, and direct spatial intelligence while remaining the primary actor.
This is how we can avoid the Empire's regret: by viewing spatial intelligence as the larger infrastructure connecting reality, representation, imagination, human intent, and physical action.
Building Spatial Intelligence Together
As new systems, interfaces, and forms of embodiment emerge, this landscape will keep changing and we will gain new ways to perceive, model, and shape the environments we live in. I hope this framework can help thinkers and builders navigate the landscape of spatial intelligence to see where your work sits, what it connects to, and where there might still be room to contribute.
We cannot afford to build in silos. People who are constructing this landscape must be connected. This is why we are building Physical I/O: a community for the people building spatial intelligence spanning the matrix: The ones building foundation world models, the ones building applications and experiences on top of them, the ones using them to train robots and physical systems, and the ones designing the interfaces for humans to direct and interact with these systems.
References
- kwarc.info — Borges, On Exactitude in Science
- worldlabs.ai — 3D as code
- worldlabs.ai — A Functional Taxonomy of World Models
- pavlo-stijn.dev — VectAR
- NVIDIA — Generative Physical AI
- waymo.com — The Waymo World Model
- Fei-Fei Li — From Words to Worlds
- nianticspatial.com — AI and the Real World
- worldlabs.ai — Real-to-Sim-to-Real
- sylvanerd.substack.com — How AI and XR Are Complementary
- Johannes Tscharn on X