I wrote this in 2020, during my first months at AfterNow, as internal concept work toward a Spatial Operating System for augmented reality. The multimodal interaction diagram in Layer 2 is dated 13 March 2020. It is shared here as a portfolio artifact — early product thinking about a category that did not yet have settled vocabulary — rather than a finished specification. The closing section marks where the thinking was still open.
Augmented reality has been called the fourth wave of computing, after the PC, the internet, and mobile. If that holds, it needs an operating system built for three dimensions rather than a 2D desktop pushed into space. This paper breaks that system into seven layers, borrowing its structure from Benjamin Bratton's The Stack: User, Interaction, Interface, Application, Spatial OS, Hardware, and AR Cloud. Along the way it proposes three levels of user sovereignty that apply equally to human and non-human agents, models human-technology exchange as a closed feedback loop across a perceptual gap, and argues that spatializing content makes ownership and permission a first-class design problem rather than an afterthought.
Why a Spatial OS
Augmented reality grants us the ability to immerse ourselves inside the data we are consuming and manipulating. The familiar 2D interfaces of personal computers and mobile devices can be expanded into a third dimension, affording a more spatial perspective and an embodied user experience. That shift is not a change of screen size. It changes what an operating system is responsible for.
A desktop OS assumes a fixed rectangle, a seated user, and a single point of focus. A spatial OS can assume none of those. Content has a position and an owner. The user moves. Other people are present, sometimes in the same room and sometimes on the other side of the world, and what they can see is a design decision rather than a given.
The term "system" here means a collection of components that combine to perform an overarching function. For this breakdown I took inspiration from The Stack: On Software and Sovereignty by Benjamin H. Bratton, which draws an analogy between the layers that build up our global information architecture and the layers representing the physical world itself — six layers descending from User to Earth. The same holistic approach applies well to a spatial operating system: each layer takes a specific form, performs a targeted action, and serves a particular purpose.
The Spatial OS stack
The layers below can be viewed independently, but they are completely interdependent — the system as a whole relies on their interplay.
| Layer | Description |
|---|---|
| 1. User | Sovereignty levels and user experience — agency, roles and capabilities |
| 2. Interaction | H2H, H2M and M2M collaboration — communicating intent and response |
| 3. Interface | Multimodal input/output exchange — device sensors and actuators |
| 4. Application | Front-end user programs — task-specific software applications |
| 5. Spatial OS | Back-end operating system — the software central nervous system |
| 6. Hardware | Physical device hardware — power, connectivity, wearability |
| 7. AR Cloud | Multiuser AR ecosystem — privacy, ownership and authorizations |
Layers 3 through 6 together constitute the device. Layer 2 is the gap the device has to cross to reach the user; layer 7 is the shared space it sits inside.
1. User Layer
Free agents
When engaged, the user is the most outward-facing part of the system. "User" here means any agent occupying the top of the stack — human or not. Unlike the software, hardware and network components that make up the technology, the user is not a permanent fixture of the system. Within a single session the same headset can be worn by different people in turn: the hardware persists while the users switch. The user is a free agent who can come and go.
Human or not
A user does not necessarily have to be human. With the Internet of Things, our increasingly networked world produces a rising number of interactions with — and between — non-human agents. On the web or in games this kind of entity is a bot or NPC. Outside computing we encounter autonomous machine systems performing dedicated tasks: retail assistants, self-driving vehicles. Autonomous agents are programmed to respond to specific stimuli in preordained ways; where a higher level of autonomy is wanted — closer to a human who can learn, adapt and improvise — more powerful systems can be developed through data analysis and machine learning.
The future workplace might contain advanced versions of these agents, mechanical or virtual, functioning like non-human co-workers: on the same network but not part of your personal AR system, dedicated to their own tasks until one of you initiates contact.
Sovereignty levels
Depending on the type of interaction and the user's demand for control, there are different extents to which they can influence outcomes.
At the top of the chain is the original creator of new content and controller of the interaction. This agent — human or not — can express itself freely, learn through observation, adapt based on reasoning and improvise accordingly. One level down, a person may be asked to provide feedback on someone else's work, usually as a rating or a comment, without the ability to respond outside those confines; the non-human equivalent is an agent with a limited set of predetermined responses. And where a user only wants to consume a data stream — a video, a music file, a live concert — a consumer-level role is enough; the non-human equivalent is an input sensor, whose sole task is gathering data.
| Role | I/O capabilities | Non-human equivalent |
|---|---|---|
| Controller | Observes and responds freely (creator) | Autonomous AI agent |
| Responder | Observes and responds within constraints (commenter) | NPC / bot / autonomous machine system |
| Consumer | Observes only (client / observer) | Input sensor |
Defining the roles this way — by I/O capability rather than by job title — is what lets the same table cover a person, a bot and a sensor without special-casing any of them.
2. Interaction Layer
Where the terminology lands
I have stuck with the term AR, though the umbrella term Extended Reality covers the ground more accurately: all real-and-virtual combined environments and human-machine interactions generated by computer technology and wearables. Within the scope of a Spatial OS we should consider the full range of interactivity and immersion — from the way people naturally interact and share experiences outside the system entirely, through the levels of augmented and mixed reality that use the physical world as a canvas for digital content, to fully immersive virtual experiences that transport the user somewhere else.
Interaction as a feedback loop
Any goal-oriented action is based on an observed discrepancy between the current and the desired state of the world. There are three types of interaction to consider — human-to-human (H2H), human-to-machine (H2M), and machine-to-machine (M2M) — and in every case both parties engage in an iterative loop that alternates between observing input and crafting a response. The most relevant here is H2M. It is in that fuzzy gap between the user's brain and the system's software where uncertainty lives, because intentions and interpretations have to be matched.
For the exchange to succeed, the user first has to perceive and understand the current state of the system, work out the change they want, and find a way to express that desire as input. The system then observes the behaviour, interprets the intent, determines how it differs from the current situation, and applies the necessary changes as output. Then it is the user's turn again. The cycle continues until the user is either satisfied or abandons the task — because they ran out of time, lost interest, got frustrated, or realised it was impossible given the device's limitations.
Interpretation and decision making
The cognitive processes that inform our understanding of the world and our decision making have digital counterparts in the diagram under the terms sensor fusion and data fission. Sensor fusion describes the collection and interpretation of all input signals. Data fission refers to the decision process involved in selecting the best output channels to convey a result. Depending on the task, the user might benefit from visual, auditory, or haptic information — often a combination.
Multimodality
The centre of the diagram breaks down the modalities available to each side. Which one works best depends on the task, and they are often complementary or interchangeable — clicking a button versus gaze plus voice command.
User input
- Speech is central to natural H2H communication, and its application to H2M has only recently gained ground. The range of supported commands keeps improving, but systems are still a long way from intuitively understanding natural language, so users often need training to learn which commands exist. Visual cues that appear on gaze — surfacing the phrase to say — help close that gap.
- Body movement is the second aspect of natural communication. Traditionally H2M interaction is facilitated through tangible input devices: keyboard, mouse, controller, touchscreen. With computer vision, users can express intent through body movement, hand gestures and eye tracking instead.
- Physiological responses are a less intentional aspect of human behaviour — subconscious functions that nonetheless offer valuable cues about physical and cognitive state. Beyond facial expressions, this includes heart rate, respiration rate and skin conductance, measurable with biosensors and interpretable as signs of stress. Few systems use this well, though wearables that track temperature, heart rate and activity are making the signal routinely available.
System output
- Visual cues are the most common form of conveyance: the appearance, alteration or movement of an object confirming that the intended action was executed.
- Auditory cues tend to be more momentary — a quick effect signifying confirmation, or a repeated sound signalling an event that needs attention.
- Haptic feedback differs from both in requiring physical stimulation. It can be tactile (touch, felt in the skin) or kinetic (motion, felt in the muscles). Tactile is by far the more common of the two: the clicks and vibrations of a hand-held controller, or the airwaves generated by ultrasonic haptics. Kinetic force feedback is harder and riskier — applied badly it can injure — but applied well it is powerful, simulating the physical presence of a virtual object by restricting the user's movement.
3. Interface Layer
The interface layer is the device front-end: where sensors and actuators exchange information with the user. Like a person's reliance on sight, hearing and touch, the system uses an array of sensors to collect information about the world around it, and actuators to communicate back.
Sensors
- Cameras enable most computer-vision processes and are essential for spatial mapping. They also support optical 6-DOF tracking for head and controller positioning, and hand-gesture recognition. Infrared cameras enable eye tracking.
- Microphones gather vocal input; devices usually carry an array of them to filter out background noise.
- Controller input tracks button presses and hand movements. Touchscreens became popular because they do both. Hand-held controllers free users from a tabletop surface; making input devices wearable and untethered adds further degrees of freedom.
- Inertial measurement units are essential for interpreting device movement and orientation. An accelerometer is an inertial-frame sensor for translation — when a device is stationary the measured acceleration is upwards and equal to gravity — and is of limited use in isolation, but valuable as part of sensor fusion. A gyroscope senses angular velocity relative to itself using the Coriolis effect; it oscillates at high frequency, which makes it one of the most power-hungry motion sensors and leaves it vulnerable to interference from a vibration motor or speaker on the same device. A magnetometer generates a 3D vector pointing to the strongest magnetic field and can act as a compass absent external influence; it too is most useful fused with the others.
Actuators
- Displays convey visual information, and AR can be achieved two ways. See-through optics project images onto a translucent surface, so the user keeps seeing the physical world with their own eyes while the system overlays digital content. Pass-through video obscures the user's view entirely with a display, captures the real world with cameras, and composites digital content into that feed before showing it. Each carries its own trade-offs in field of view, latency, and how convincingly virtual content can occlude the real.
- Speakers communicate auditory information, and split along the same line. Built-in speakers are the audio equivalent of see-through optics — housed near the ears, additive to the real world, leaving the user able to talk to people nearby. Headphones are the pass-through equivalent, mostly replacing environmental sound with generated audio, which suits VR better than AR.
- Haptic actuators likewise come in two varieties: tactile, typically a small vibration motor embedded in a hand-held controller or an ultrasonic array producing sensation in mid-air; and kinetic, applying force to restrict movement, as in a force-feedback glove.
4. Application Layer
System output can be expressed at two levels: by the operating system directly, or by special-purpose programs called applications. Applications run only at the user's request and perform optional tasks not crucial to the system's own functioning — word processing, browsing, checking the weather. I picture the operating system as a general commerce platform, and applications as dedicated storefronts where only particular goods and services are exchanged.
Many subdivisions are possible, but broadly applications fall into categories such as productivity, entertainment, education and utility. Many conventional applications stand to benefit from visualization in 3D space rather than merely being ported into it — the question worth asking of any candidate app is not whether it can be spatialized, but what the third dimension actually buys the user.
5. Spatial OS Layer
Function
The Spatial Operating System is the engine that brings the technology to life. If the hardware resembles the body, the OS is the brain and central nervous system. It has three main functions: manage hardware resources — processor, memory, storage, connected devices; establish a user interface; and execute and provide services for application software.
The evolution of mobile computing
Conventional 2D implementations exist as Windows and macOS on personal computers, iOS and Android on mobile. Mobile made our digital experiences far more portable, but the content itself stayed confined to a small hand-held 2D window. Wearable AR and VR headsets push this a step further, immersing the user in content while leaving the hands free to interact with it in three dimensions — at the cost of a trade-off between performance (tethered, reliant on external compute) and portability (untethered, battery powered). The near-term shift that looks inevitable is from smartphones to smart glasses.
The open challenge
Given the space available and the user's tendency to move around in it, the key challenge for a Spatial OS is deciding how and where to present its own content — including applications. On a desktop this question is settled by the edges of the screen. In a room it is not settled at all, and it is the question the closing section returns to.
6. Hardware Layer
Where the interface layer is user-facing, the hardware layer covers the backbone: memory, processing power, battery capacity, connectivity — and, critically for a device worn on the head, portability and comfort. Comfort is a function of form factor, padding and weight, and it sets a hard ceiling on session length that no amount of software quality can raise. Battery life is the equivalent constraint for standalone headsets. We want the user to get away from their desk and move through physical space, which makes wearability a requirement rather than a feature.
7. AR Cloud Layer
A networked world
The AR Cloud is the network arrangement that lets users connect not only to the internet but to their environments and to each other. Like nodes in a traditional network, the number of connections each system needs depends on how many access points are wanted. Some users are perfectly content working in isolation — a person in a library with an internet connection and no interest in what anyone else is doing. More commonly, users want to interact with others as part of a team, an organization, or a social group.
Space ownership
Spatializing content makes ownership a design problem. Depending on the setting, users experience different extents to which they can consume — and be subjected to — the digital content around them.
At the most personal level, each user has a bubble they exercise full control over. Whatever the environment, they should always own their personal space, with the option to see only their own content if they want to shield themselves entirely. Like the preset levels on noise-cancelling headphones, the user should control the size of that bubble — three metres, ten metres, infinite. Others may send an invitation to share part of their experience, and the user can always decline.
Private spaces are shared among selected peers: teams, companies, social groups. Users join a shared reality much as they walk into an office building or sign into a chat program. Again, content in these spaces cannot be forced into the personal bubble without consent — entering private mode is how the user signals readiness to interact. Content there is shielded from strangers, who need an invitation first, permanent or time-limited.
Public spaces contain content available to anyone with a headset. Some may carry age restrictions or vary by viewer, as targeted advertising does. And once again, any user inside a public space can filter all of it out using their personal bubble.
| Ownership | Level of control over digital content, depending on the setting |
|---|---|
| Personal | An owner's personal bubble — visibility cannot be controlled by others |
| Private | A digital space for selected peers to share — others need to be invited |
| Public | Content accessible to every user in proximity — some restrictions may apply |
Sharing levels
Once people are on the same network, the question becomes what kind of shared experience they are having. Thinking about how communication is set up between IP addresses turns out to map cleanly onto how experiences can be shared, with the receiver holding their own sharing settings at their end.
| Cast level | Type of shared experience within the AR cloud |
|---|---|
| Unicast | For a single controller to perceive and act upon, with or without the help of non-human agents |
| Multicast | For selected individuals as either controllers or consumers — a team meeting |
| Broadcast | For unlimited individuals as either controllers or consumers — e-sports |
Cross-referencing this table with the sovereignty levels in Table 2 gives the actual permission matrix: who can act, and how many can see them do it. Those two questions are separable, and treating them as one is where most collaborative systems get into trouble.
Open questions — where this was heading
The layers above were worked through. What follows was not finished, and I am leaving it visible rather than tidying it away, because the unresolved parts are the interesting ones.
How should an OS represent itself in a room?
A well-rounded AR experience needs an OS representation — not just one application with an interesting interface, but the whole system cohering intuitively. Its main task is to convey essential system information (time, battery, connectivity, updates) while providing a fast way to find and open applications. Every device manufacturer was solving this independently, each with their own merits and flaws. The evaluative questions I wanted to ask of each: How is the main menu summoned? Where does it appear, near or far? How many shortcuts are shown by default? Where do you find everything else? A systematic comparison across the major platforms of the day remained to be done.
A geometric OS
It would be interesting to depart from conventional 2D interfaces altogether and shape the OS as one of the Platonic solids — the five regular convex polyhedra: tetrahedron, cube, octahedron, dodecahedron, icosahedron. The user could summon it with a simple gesture, take hold of the shape in one hand, rotate it, and make a selection with the other. Starting with the cube, each of its six faces could correspond to a category of commands: system settings and status; productivity; entertainment; education; commerce; utility. If six faces proved too few, an octahedron would allow a seventh face for currently-open applications — an app switcher — and an eighth, oriented downward when summoned, for log out, sleep and power down.
Whether a graspable solid is genuinely better than a menu, or merely more novel, is exactly the question that needed testing rather than asserting.
The menu itself — graspable or flat — may also prove transitional. With capable voice AI assistance, the user can simply ask for what they want instead of looking through a visual menu for the right application; navigation gives way to conversation. If so, the solid becomes the browsing view for the moments when you do not yet know what to ask for.
System-wide versus application-specific control
Which commands belong to the OS and which to the application is a settled question on desktop and unsettled in space. Compared with a screen-based interface, an AR user has far more room to distribute content, which may remove much of the need to switch between applications competing for scarce screen real estate. If windows no longer compete for space, it is not obvious that the app switcher — one of the load-bearing metaphors of 2D computing — needs to survive at all. That question was open when I stopped writing, and largely still is.