This white paper was written as exploratory concept work for Second Studio, an early-stage VR startup that is no longer operating. It is shared here as a portfolio artifact showing early product thinking around multi-user VR permissions and roles, rather than a finished or shipped specification.
A shared virtual environment is more complex than a single-user one: several people occupy the same space at once, and not all of them should be able to do the same things. This paper proposes a framework for that problem. It first builds an analogy between a natural ecosystem and a cooperative VR room — environment as the abiotic layer, data as producer, users as consumers — and from it derives three user roles: Passive Observer, Active Observer, and Controller. It then structures how those roles interact through three session patterns borrowed from networking — unicast, broadcast, and multicast — and applies the resulting model to three use cases in medical imaging, psychological treatment, and EEG-based brain activity monitoring.
Goal
To define user roles and structure their interactions within multi-user content creation platforms.
Introduction
Second Studio's platform will enable multiple users to coexist and cooperate in a shared VR environment. This increase in complexity requires a new framework that provides a clear definition of roles and dependencies within this room-sized ecosystem. The second challenge lies in defining how different users can interact with the environment and each other, depending on their assigned roles.
Challenge 1: Defining the VR ecosystem
This paper aims to establish a new framework for understanding, by first creating an analogy between the VR experience and a natural ecosphere, and second by defining the various user scenarios and the different levels of involvement that these encompass.
Background
I have composed the following analogy to compare a natural ecosystem with the cooperative VR environment that we are facilitating. I selected a farm theme for this ecosystem, since its focus lies on the cultivation of environmental resources (data), rather than the consumption of other inhabitants.
Framework
Each ecosystem consists of abiotic and biotic components. Abiotic components consist of the basic elements: earth, air, water, light. In our analogy this describes the VR room and its features. Biotic components describe the living parts of the ecosystem, which can be subdivided into producers (plants) and consumers (animals). In our analogy, data is plantlike, still part of the environment, and plays the role of producer as the source of information. Similarly, users are the consumers of information, with various levels of hierarchy among them. All users are observers by default, but depending on their role they may also freely navigate and manipulate the environment and its contents.
As shown in the figure, I propose the following user designations:
- Passive ObserverCan only perceive the environment and the events within it.
- Active ObserverCan navigate VR space, but is unable to interact with it.
- ControllerCan navigate VR space and manipulate its contents.
Each of these roles will encompass a different set of interface options.
Challenge 2: Structuring interaction between users and the environment
Background
The best way to envision these VR scenarios is to compare them with real-life examples: a VR session with one-on-one training behaves very differently from classroom teaching. Depending on which users are involved, you will need to provide an appropriate workplace. The solution is to form a clear understanding of how to assign different roles to each user.
- Dataset — data file loaded and represented in 3D space
- Controller — observes, navigates, manipulates (own POV)
- Active Observer — observes, navigates (own POV)
- Passive Observer — observes (controller POV or static camera)
Observer display options
- Passive observer — watches the controller interact with VR, either on a conventional screen (computer or tablet) or using an HMD. When wearing an HMD, the controller's view should be displayed on a theatre screen to avoid motion sickness. Additional information can be displayed next to the controller's view, for example the object and tool that have been selected.
- Active observer (requires HMD) — has the ability to autonomously explore the VR space and watch the controller at work. Can potentially see which tool the controller is operating through a HUD.
Datacasting scenarios
The following session patterns will apply, depending on the application:
- Unicast — single controller, single observer (e.g., doctor and patient).
- Broadcast — single controller with multiple observers (e.g., teacher and classroom).
- Multicast — multiple controllers, with or without observers (e.g., cooperative design).
The framework in summary
User roles
- Controller — manipulates VR space using HMD and interaction devices.
- Active Observer — can move in VR space, but is unable to interact with objects.
- Passive Observer — perceives the controller's view and manipulations.
Observer interface options
- Conventional — observer watches the controller interact with VR on a conventional screen (computer or tablet).
- HMD — observer enters VR space and watches a theatre screen, with side panels containing additional information (e.g., object and tool selected).
Possible interface scenarios
- Unicast — single controller and single observer (e.g., doctor's office).
- Broadcast — single controller with multiple observers (e.g., classroom).
- Multicast — multiple controllers working together, with or without observers (e.g., co-working).
Example scenarios
The following use cases are described using the STAR structure: Situation, Task, Activity, Result.
Application 1: Medical imaging — manipulating 3D data
STwo users can be in the same room or in two different locations. Medical data is loaded into the environment, and both users are free to manipulate it.
TThe goal is to analyze a 3D model (e.g., an MRI) in order to diagnose or discuss medical conditions.
AThe users can manipulate the 3D model — for example, slice it and view it from different angles. Utilizing voxel data, the interface allows them to access different scales, letting them literally dive in and identify potential health issues. Picture a scene in which one doctor is manipulating a section plane at the meso scale, while another doctor observes the data at the micro scale, exploring the 3D scan as if positioned on the section plane surface. Users can apply drawing tools to note areas of interest, or extract them for closer analysis.
RBy cooperating and combining their expertise, the users can explore the data and come to a diagnosis together.
Development requirements
- Generate 3D voxel models from scan data files (possible with third-party software).
- Enable multiple controllers to view the 3D model simultaneously.
- Utilize voxel data to shrink or grow the observer across scale levels: macro, meso, micro, nano.
- Enable users to work at different scale levels simultaneously (one at meso, one at macro).
- Apply a virtual toolset to manipulate data areas (e.g., highlight, extract).
User scenarios
- 1a — Two controllers. These users can be two doctors, or one instructor and one student. This application can be used for training, or as an assessment tool by appointing the student to the primary role and having the instructor act more as an observer.
- 1b — One controller and one observer (active or passive). This can describe a personal consultation between a doctor, dentist, or chiropractor and a patient, or between an instructor and a student. When applied to educational use, the system can also function as an assessment tool by appointing the student to the primary role and having the instructor observe.
- 1c — One controller and multiple passive observers. This can be applied in a classroom setting, where one instructor controls the VR space and multiple students follow along on their screens.
Application 2: Psychology — stress and phobia treatment
The positioning here is a "Netflix" of psychological content: prescribing specific experiences to help patients deal with anxiety or stress. We would not create the content ourselves, but offer it.
SA patient needs help from a psychologist in dealing with stress or trauma.
TTo allow a patient to familiarize themselves with stress stimuli in a controlled and comfortable environment.
AThe instructor interacts with the fear stimulus first (e.g., a spider) to convey safety, talking the patient through the experience and guiding them. This is an opportunity for biosensors embedded in the controller to measure changes in the patient's stress level.
RProlonged use will familiarize the patient with the stimulus, countering psychological and physiological stress responses.
Development requirements
- Investigate the most common fears (e.g., spiders, heights, public spaces) so that we can generate a small set of safe environments offering variable stress levels suitable for a wide range of potential users.
- Create a multi-user interface that allows a psychologist to demonstrate safety and guide the patient through the experience.
User scenarios
- 2a — Two controllers. These users can be a psychologist and a patient who is ready to interact with the fear stimulus.
- 2b — One controller and one observer (active or passive). These users can be a psychologist and a patient who is not quite ready to interact with the fear stimulus. As a mere observer, the patient might feel safer for the time being.
- 2c — One controller and multiple passive observers. This can be applied in a classroom setting, where one instructor controls the VR space and multiple students or patients follow along on their screens.
Application 3: EEG — brain activity monitoring
Patient and instructor both observe the patient's brain activity, represented as a fuzzy voxel cloud, and can jump between scale levels as proposed in the medical imaging example.
SA patient and instructor share a VR session to review the patient's real-time brain activity, captured via EEG and represented as a volumetric voxel cloud in the shared space.
THelp the instructor guide the patient through mental exercises while both observe how specific tasks affect brain activity, in real time and across multiple scales.
AThe instructor assigns short mental tasks (e.g., word association, simple calculations) while biosensors embedded in the controller feed live EEG data into the shared voxel cloud. Both users can jump between macro and micro scale levels to view activity across the whole brain or zoom into specific regions.
RThe patient gains a visual, real-time understanding of how their own mental activity changes, and the instructor can tailor exercises to observed patterns — useful for neurofeedback training or cognitive research.
A better analogy: the arena
Returning to this paper years later, the ecosystem analogy in Challenge 1 still does one job well and another badly, and it took me a while to see why. Ecology is excellent at describing what kinds of things exist in a world — substrate, producers, consumers. That is why the abiotic/biotic split and data-as-producer hold up. But ecology has no vocabulary for how much each agent may change the world. Trophic levels rank who eats whom, not who may rearrange the furniture. I was asking one analogy to do two jobs, and the seam shows at the bottom rung: I leaned on microbes to represent immobility, when motility is precisely the thing many microbes have.
A shared space with a single event in it works far better for the second job. Picture a concert in an arena, with the stage in the centre rather than at one end.
The show — music and light — is the shared content. It is spatial, multimodal and ambient: it fills the volume and reaches everyone at once, though differently depending on where they stand. The performers are the only agents who can alter it. The floor audience moves wherever it likes and changes nothing. The stands hold an audience with their own view of the same show, fixed to the seat they chose from a finite set. All four are present simultaneously, by the architecture of the building rather than by a rule — which is what the farm could never manage, since penned and free-ranging animals are not in the same field as the farmer.
And the arena covers something the original paper treated separately. A concert is also broadcast, and the viewer at home is not a fourth role — it is the second flavour of Passive Observer this paper already named: controller POV or static camera. The person in the stands is the static camera, holding their own fixed viewpoint. The person at home is the controller POV, watching a viewpoint someone else selected.
| At the concert | Role | Viewpoint | Presence |
|---|---|---|---|
| On stage | Controller | own, free, plus authority | in the space |
| On the floor | Active Observer | own, free | in the space |
| In the stands | Passive Observer | own, fixed | in the space |
| Watching the stream | Passive Observer | chosen by someone else | remote |
Laying it out this way exposes something I had left implicit. The ladder is not one variable with three settings; it is two separate increments:
fixed position → free position → free position plus authority over shared state
Stands to floor is purely positional. Floor to stage is authority. That distinction matters in a real system, because the two permissions are separately grantable: you can let someone roam without letting them touch anything, which is the entire reason the middle role exists.
It also produces a genuinely counterintuitive result. The viewer at home often sees more than the person in the stands — close-ups, camera changes, replays, the mix straight off the desk. Lowest permission, highest information. Anyone designing an observer experience should not assume that more freedom yields a better one; an active observer who has wandered behind the data may understand it less well than a passive observer following a well-chosen view.
A fourth axis: the return channel
There is one more thing the arena separates that the original paper never did, and it only becomes visible once the broadcast audience is in the picture. It is not what you can perceive. It is whether you can be perceived.
The floor crowd is felt: a band demonstrably plays differently to a live room. The stands are still seen and heard — they cannot move, but forty thousand people roaring is unmistakably present to the performer. The remote viewer is perceptually absent. Same show, same content, and from the stage they do not exist. In current terms, they are attending with camera and microphone off.
This has a name. Benford and Fahlén's spatial model of interaction splits presence into aura, focus and nimbus — nimbus being how much of you is projected to others, how observable you are. Awareness between two participants is a function of one's focus and the other's nimbus. That is exactly the pair of properties the arena pulls apart, and the vocabulary is older than this paper by two decades; I arrived at the distinction independently, but I did not invent it.
One precision matters here. The floor's influence is not authority. They cannot change a note. What they change is the person who changes the notes. Influence through the return channel is mediated by the controller; authority is direct. That is why crowd energy does not promote the floor to Controller — and it is the cleanest available answer to why the middle role is genuinely a middle role rather than a weak Controller.
The completed framework
Put together, the properties that were tangled through the original paper separate into two groups. Three describe a participant; two describe the session they are in. Keeping them apart is what the earlier version failed to do — Table 2 mixed roles with what were really session-level properties.
| Axis | Question it answers | Values |
|---|---|---|
| Authority | What may you change? | nothing · within constraints · freely |
| Vantage | Who decides where you see from? | you, freely · you, from a fixed set · someone else |
| Nimbus | How perceptible are you to others? | absent · visible · visible and influential |
| Axis | Question it answers | Values |
|---|---|---|
| Reach | How many share this? | unicast · multicast · broadcast |
| Ownership | Whose space is it? | personal · private · public |
The three roles this paper named in 2017 are a diagonal through the participant axes, not a scale. Controller is full authority, free vantage, high nimbus. Active Observer is no authority, free vantage, visible. Passive Observer is no authority, fixed or assigned vantage, and nimbus that varies from present to absent depending on whether they are in the room or watching a stream. Once the axes are separate, combinations the roles never covered become expressible — and most of the interesting ones are exactly the combinations nobody has designed for.
Where each piece came from
These axes were not invented here; they accumulated. Setting them against their sources shows which questions kept recurring:
| Axis | First appears | As |
|---|---|---|
| Authority | This paper, 2017 | Controller / Active Observer / Passive Observer; reframed in 2020 as Controller / Responder / Consumer, ordered by response bandwidth rather than position |
| Vantage | This paper, 2017 | Observer display options — "controller POV or static camera" |
| Nimbus | Microsoft Research, c. 2022 | Built as the Embodiment axis of the 48-configuration matrix below, and latent in the brief itself: the requirement that participants be "equal presence, represented, and heard" is a statement about nimbus. Named as such only here, in 2026. |
| Reach | This paper, 2017 | Unicast / multicast / broadcast, carried unchanged into the 2020 stack |
| Ownership | Spatial OS, 2020 | Personal / private / public, in the AR Cloud layer — carried into the Microsoft matrix as Personal / Communal / Virtual |
The chart this became at Microsoft
Four years after the Spatial OS paper, consulting for Microsoft Research on hybrid meetings, I built a matrix of every configuration a hybrid meeting could take. It had three axes — Embodiment (levels of agency based on representation), Spaces (levels of control based on ownership), and Interfaces (levels of immersion) — with each of the 48 resulting nodes coloured by how feasible it was to build.
I did not notice at the time that I had reached for my own vocabulary. Spaces is the ownership ladder from the AR Cloud layer of the 2020 paper, relabelled from personal/private/public to Personal/Communal/Virtual. Embodiment is the axis this 2026 note calls nimbus — how much of you is projected to the people you are with. And Interfaces is the display-options question from 2017, generalised: not screen or headset, but a graded scale of immersion.
The original was a lattice you walked around in VR. ShapesXR deleted the model when their billing changed and I was not on a paid plan; what survives is a screen capture, from which the axes were recovered. Rebuilt here as three panels, because on a page you cannot walk around a cube.
Two things fall out of it that were not visible in the lattice. The Avatar column is feasible almost everywhere — across every interface and every kind of space. The Physical column is difficult everywhere except face-to-face, which is close to a tautology once stated but was not obvious before: being bodily present is not something an interface can grant you. Between those two, the whole design space collapses onto a single question — if you cannot be there, how well can you be represented?
That is the nimbus axis, discovered a second time and under a different name.
All 48 configurations as a table
| Space | Interface | Screen | Robot | Physical | Avatar (AR/VR) |
|---|---|---|---|---|---|
| Personal | Face-to-face | Difficult | Difficult | Feasible | Difficult |
| Personal | PC tablet | Feasible | Feasible | Difficult | Feasible |
| Personal | HoloLens (AR) | Difficult | Partial | Difficult | Feasible |
| Personal | Quest (VR) | Difficult | Feasible | Difficult | Feasible |
| Communal | Face-to-face | Difficult | Difficult | Feasible | Difficult |
| Communal | PC tablet | Feasible | Feasible | Difficult | Feasible |
| Communal | HoloLens (AR) | Difficult | Partial | Difficult | Feasible |
| Communal | Quest (VR) | Difficult | Feasible | Difficult | Feasible |
| Virtual | Face-to-face | Difficult | Difficult | Difficult | Difficult |
| Virtual | PC tablet | Difficult | Difficult | Difficult | Feasible |
| Virtual | HoloLens (AR) | Difficult | Difficult | Difficult | Feasible |
| Virtual | Quest (VR) | Difficult | Difficult | Difficult | Feasible |
Two places the analogy would mislead
At a concert the controller authors the content live, whereas in a shared virtual environment the dataset generally pre-exists and the controller acts upon it. What carries across is not where the content comes from, but who may alter what everyone else is experiencing.
More importantly: at a concert, the remote viewer's absence is physics. In a virtual environment it is a setting. The return channel exists, and somebody chose to leave it off by default. That single substitution — a design decision wearing the costume of a physical constraint — is, I think, the reason hybrid meetings go wrong. The asymmetry looks like a property of the medium, so nobody revisits it.
Open questions
Who holds authority over the view? Someone chooses what the home audience sees, and it is not the band. The vision mixer has authority over presentation without any authority over content: they cannot change a note, but they decide what a million people look at. None of the three roles account for that. It sits sideways to the ladder — less than a Controller over the data, more than an Observer over the framing. In a shared session that role is real: the facilitator who repositions everyone's view, the guide who directs attention, whoever defines the fixed camera positions a passive observer chooses between.
Is nimbus grantable per participant, or only per session? Broadcast bundles "many receivers" with "no return path," but those are separable, and unbundling them is where equitable hybrid participation would have to start.
This paper identified one missing middle role in 2017. Nine years later the arena surfaces two more — the presentation authority, and the participant whose presence is switched off by a default nobody chose deliberately. The recurring problem, across all of it, is the middle.