White paper · Virtual reality · Multi-user systems

VR Ecosystem and User Scenarios

A framework for defining user roles, permissions, and interaction in a shared virtual environment

Context

This white paper was written as exploratory concept work for Second Studio, an early-stage VR startup that is no longer operating. It is shared here as a portfolio artifact showing early product thinking around multi-user VR permissions and roles, rather than a finished or shipped specification.

Summary

A shared virtual environment is more complex than a single-user one: several people occupy the same space at once, and not all of them should be able to do the same things. This paper proposes a framework for that problem. It first builds an analogy between a natural ecosystem and a cooperative VR room — environment as the abiotic layer, data as producer, users as consumers — and from it derives three user roles: Passive Observer, Active Observer, and Controller. It then structures how those roles interact through three session patterns borrowed from networking — unicast, broadcast, and multicast — and applies the resulting model to three use cases in medical imaging, psychological treatment, and EEG-based brain activity monitoring.

Goal

To define user roles and structure their interactions within multi-user content creation platforms.

Introduction

Second Studio's platform will enable multiple users to coexist and cooperate in a shared VR environment. This increase in complexity requires a new framework that provides a clear definition of roles and dependencies within this room-sized ecosystem. The second challenge lies in defining how different users can interact with the environment and each other, depending on their assigned roles.

Challenge 1: Defining the VR ecosystem

This paper aims to establish a new framework for understanding, by first creating an analogy between the VR experience and a natural ecosphere, and second by defining the various user scenarios and the different levels of involvement that these encompass.

Background

I have composed the following analogy to compare a natural ecosystem with the cooperative VR environment that we are facilitating. I selected a farm theme for this ecosystem, since its focus lies on the cultivation of environmental resources (data), rather than the consumption of other inhabitants.

Framework

Each ecosystem consists of abiotic and biotic components. Abiotic components consist of the basic elements: earth, air, water, light. In our analogy this describes the VR room and its features. Biotic components describe the living parts of the ecosystem, which can be subdivided into producers (plants) and consumers (animals). In our analogy, data is plantlike, still part of the environment, and plays the role of producer as the source of information. Similarly, users are the consumers of information, with various levels of hierarchy among them. All users are observers by default, but depending on their role they may also freely navigate and manipulate the environment and its contents.

Diagram mapping a natural ecosystem onto a shared virtual environment. Ecosystem splits into Abiotic and Biotic. Abiotic leads to Basic Elements — earth, wind, water, light — which map to the VR Environment, whose function is to host. Biotic splits into Producer and Consumer. Producer leads to Plantlife, mapping to Data, whose function is to supply. Consumer leads to three tiers: Microbes map to Passive Observer, which observes; Herbivores map to Active Observer, which observes and navigates; Farmer maps to Controller, which observes, navigates and manipulates.
Figure 1. The VR ecosystem analogy. Components of a natural ecosystem are mapped onto their equivalents in a shared virtual environment, yielding a hierarchy of user roles defined by what each is permitted to do.

As shown in the figure, I propose the following user designations:

Each of these roles will encompass a different set of interface options.

Challenge 2: Structuring interaction between users and the environment

Background

The best way to envision these VR scenarios is to compare them with real-life examples: a VR session with one-on-one training behaves very differently from classroom teaching. Depending on which users are involved, you will need to provide an appropriate workplace. The solution is to form a clear understanding of how to assign different roles to each user.

Isometric render of a virtual room. A red dataset volume sits at the centre. Two green figures with controllers stand close to it, several blue figures are positioned around it inside the room, and yellow figures sit outside the room's walls. Colour indicates role: red for the dataset, green for controllers, blue for active observers, yellow for passive observers.
Figure 2. A VR room with various users observing the same dataset from different scale levels.

Observer display options

Datacasting scenarios

The following session patterns will apply, depending on the application:

The framework in summary

User roles

Observer interface options

Possible interface scenarios

Example scenarios

The following use cases are described using the STAR structure: Situation, Task, Activity, Result.

Application 1: Medical imaging — manipulating 3D data

STwo users can be in the same room or in two different locations. Medical data is loaded into the environment, and both users are free to manipulate it.

TThe goal is to analyze a 3D model (e.g., an MRI) in order to diagnose or discuss medical conditions.

AThe users can manipulate the 3D model — for example, slice it and view it from different angles. Utilizing voxel data, the interface allows them to access different scales, letting them literally dive in and identify potential health issues. Picture a scene in which one doctor is manipulating a section plane at the meso scale, while another doctor observes the data at the micro scale, exploring the 3D scan as if positioned on the section plane surface. Users can apply drawing tools to note areas of interest, or extract them for closer analysis.

RBy cooperating and combining their expertise, the users can explore the data and come to a diagnosis together.

Development requirements

User scenarios

Application 2: Psychology — stress and phobia treatment

The positioning here is a "Netflix" of psychological content: prescribing specific experiences to help patients deal with anxiety or stress. We would not create the content ourselves, but offer it.

SA patient needs help from a psychologist in dealing with stress or trauma.

TTo allow a patient to familiarize themselves with stress stimuli in a controlled and comfortable environment.

AThe instructor interacts with the fear stimulus first (e.g., a spider) to convey safety, talking the patient through the experience and guiding them. This is an opportunity for biosensors embedded in the controller to measure changes in the patient's stress level.

RProlonged use will familiarize the patient with the stimulus, countering psychological and physiological stress responses.

Development requirements

User scenarios

Application 3: EEG — brain activity monitoring

Patient and instructor both observe the patient's brain activity, represented as a fuzzy voxel cloud, and can jump between scale levels as proposed in the medical imaging example.

SA patient and instructor share a VR session to review the patient's real-time brain activity, captured via EEG and represented as a volumetric voxel cloud in the shared space.

THelp the instructor guide the patient through mental exercises while both observe how specific tasks affect brain activity, in real time and across multiple scales.

AThe instructor assigns short mental tasks (e.g., word association, simple calculations) while biosensors embedded in the controller feed live EEG data into the shared voxel cloud. Both users can jump between macro and micro scale levels to view activity across the whole brain or zoom into specific regions.

RThe patient gains a visual, real-time understanding of how their own mental activity changes, and the instructor can tailor exercises to observed patterns — useful for neurofeedback training or cognitive research.

Note added 2026 — not part of the original document

A better analogy: the arena

Returning to this paper years later, the ecosystem analogy in Challenge 1 still does one job well and another badly, and it took me a while to see why. Ecology is excellent at describing what kinds of things exist in a world — substrate, producers, consumers. That is why the abiotic/biotic split and data-as-producer hold up. But ecology has no vocabulary for how much each agent may change the world. Trophic levels rank who eats whom, not who may rearrange the furniture. I was asking one analogy to do two jobs, and the seam shows at the bottom rung: I leaned on microbes to represent immobility, when motility is precisely the thing many microbes have.

A shared space with a single event in it works far better for the second job. Picture a concert in an arena, with the stage in the centre rather than at one end.

The arena analogy — roles around a centre stage Plan view of an arena with a stage at its centre. The circular stage holds three performers marked green as Controllers — the only agents who alter the show. Music and light radiate outward from the stage as a red field crossing every ring, representing the shared content. Around the stage lies the floor, an open blue ring holding figures scattered at arbitrary positions: Active Observers, free to move anywhere but unable to change anything. Beyond a barrier is the seating bowl, two concentric rings of evenly spaced amber seats: Passive Observers, each with their own viewpoint but fixed to the seat they chose, one of which is highlighted as occupied. Outside the arena a broadcast feed runs to a screen showing a close-up of a performer, watched by a further amber figure: also a Passive Observer, but with the viewpoint chosen by a director rather than by themselves. Because the stage is central rather than at one end, the content has no privileged side and is viewable through a full 360 degrees — a volume rather than a surface. Two arrows run inward from the audience back to the stage, showing the return channel: a thick blue arrow from the floor, whose energy the performers feel, and a thinner dashed amber arrow from the stands, who are still seen and heard. From the remote viewer no such arrow exists; the line back to the arena is cut and marked no return path. Return channel: the floor is felt, the stands are seen and heard, the remote viewer is perceptually absent — like a call with camera and microphone off. stage floor stands The stage is central, not at one end: the show has no privileged side and is viewable through 360°. Stage — the show music and light: the shared content, radiating outward Performers — Controller the only agents who alter the show Floor — Active Observer own viewpoint, free to move, changes nothing Stands — Passive Observer own viewpoint, fixed to the seat they chose Livestream — Passive Observer same permission, but the viewpoint is chosen by a director, not by the viewer broadcast feed no return path
Figure 3. The arena analogy. A centre stage has no privileged side, so the show is a volume rather than a surface — which is the same difference that separates a shared virtual environment from a screen.

The show — music and light — is the shared content. It is spatial, multimodal and ambient: it fills the volume and reaches everyone at once, though differently depending on where they stand. The performers are the only agents who can alter it. The floor audience moves wherever it likes and changes nothing. The stands hold an audience with their own view of the same show, fixed to the seat they chose from a finite set. All four are present simultaneously, by the architecture of the building rather than by a rule — which is what the farm could never manage, since penned and free-ranging animals are not in the same field as the farmer.

And the arena covers something the original paper treated separately. A concert is also broadcast, and the viewer at home is not a fourth role — it is the second flavour of Passive Observer this paper already named: controller POV or static camera. The person in the stands is the static camera, holding their own fixed viewpoint. The person at home is the controller POV, watching a viewpoint someone else selected.

Table 5. One analogy covering roles, display options and presence.
At the concertRoleViewpointPresence
On stageControllerown, free, plus authorityin the space
On the floorActive Observerown, freein the space
In the standsPassive Observerown, fixedin the space
Watching the streamPassive Observerchosen by someone elseremote

Laying it out this way exposes something I had left implicit. The ladder is not one variable with three settings; it is two separate increments:

fixed position → free position → free position plus authority over shared state

Stands to floor is purely positional. Floor to stage is authority. That distinction matters in a real system, because the two permissions are separately grantable: you can let someone roam without letting them touch anything, which is the entire reason the middle role exists.

It also produces a genuinely counterintuitive result. The viewer at home often sees more than the person in the stands — close-ups, camera changes, replays, the mix straight off the desk. Lowest permission, highest information. Anyone designing an observer experience should not assume that more freedom yields a better one; an active observer who has wandered behind the data may understand it less well than a passive observer following a well-chosen view.

A fourth axis: the return channel

There is one more thing the arena separates that the original paper never did, and it only becomes visible once the broadcast audience is in the picture. It is not what you can perceive. It is whether you can be perceived.

The floor crowd is felt: a band demonstrably plays differently to a live room. The stands are still seen and heard — they cannot move, but forty thousand people roaring is unmistakably present to the performer. The remote viewer is perceptually absent. Same show, same content, and from the stage they do not exist. In current terms, they are attending with camera and microphone off.

This has a name. Benford and Fahlén's spatial model of interaction splits presence into aura, focus and nimbus — nimbus being how much of you is projected to others, how observable you are. Awareness between two participants is a function of one's focus and the other's nimbus. That is exactly the pair of properties the arena pulls apart, and the vocabulary is older than this paper by two decades; I arrived at the distinction independently, but I did not invent it.

One precision matters here. The floor's influence is not authority. They cannot change a note. What they change is the person who changes the notes. Influence through the return channel is mediated by the controller; authority is direct. That is why crowd energy does not promote the floor to Controller — and it is the cleanest available answer to why the middle role is genuinely a middle role rather than a weak Controller.

The completed framework

Put together, the properties that were tangled through the original paper separate into two groups. Three describe a participant; two describe the session they are in. Keeping them apart is what the earlier version failed to do — Table 2 mixed roles with what were really session-level properties.

Table 6. Participant axes — what a single agent may do.
AxisQuestion it answersValues
AuthorityWhat may you change?nothing · within constraints · freely
VantageWho decides where you see from?you, freely · you, from a fixed set · someone else
NimbusHow perceptible are you to others?absent · visible · visible and influential
Table 7. Session axes — properties of the shared space itself.
AxisQuestion it answersValues
ReachHow many share this?unicast · multicast · broadcast
OwnershipWhose space is it?personal · private · public

The three roles this paper named in 2017 are a diagonal through the participant axes, not a scale. Controller is full authority, free vantage, high nimbus. Active Observer is no authority, free vantage, visible. Passive Observer is no authority, fixed or assigned vantage, and nimbus that varies from present to absent depending on whether they are in the room or watching a stream. Once the axes are separate, combinations the roles never covered become expressible — and most of the interesting ones are exactly the combinations nobody has designed for.

Where each piece came from

These axes were not invented here; they accumulated. Setting them against their sources shows which questions kept recurring:

Table 8. The lineage of each axis.
AxisFirst appearsAs
AuthorityThis paper, 2017Controller / Active Observer / Passive Observer; reframed in 2020 as Controller / Responder / Consumer, ordered by response bandwidth rather than position
VantageThis paper, 2017Observer display options — "controller POV or static camera"
NimbusMicrosoft Research, c. 2022Built as the Embodiment axis of the 48-configuration matrix below, and latent in the brief itself: the requirement that participants be "equal presence, represented, and heard" is a statement about nimbus. Named as such only here, in 2026.
ReachThis paper, 2017Unicast / multicast / broadcast, carried unchanged into the 2020 stack
OwnershipSpatial OS, 2020Personal / private / public, in the AR Cloud layer — carried into the Microsoft matrix as Personal / Communal / Virtual

The chart this became at Microsoft

Four years after the Spatial OS paper, consulting for Microsoft Research on hybrid meetings, I built a matrix of every configuration a hybrid meeting could take. It had three axes — Embodiment (levels of agency based on representation), Spaces (levels of control based on ownership), and Interfaces (levels of immersion) — with each of the 48 resulting nodes coloured by how feasible it was to build.

I did not notice at the time that I had reached for my own vocabulary. Spaces is the ownership ladder from the AR Cloud layer of the 2020 paper, relabelled from personal/private/public to Personal/Communal/Virtual. Embodiment is the axis this 2026 note calls nimbus — how much of you is projected to the people you are with. And Interfaces is the display-options question from 2017, generalised: not screen or headset, but a graded scale of immersion.

The original was a lattice you walked around in VR. ShapesXR deleted the model when their billing changed and I was not on a paid plan; what survives is a screen capture, from which the axes were recovered. Rebuilt here as three panels, because on a page you cannot walk around a cube.

Feasibility of 48 hybrid meeting configurations Three 4x4 grids side by side, one for each Space: Personal, Communal and Virtual. In each grid the rows are Interfaces in increasing immersion — Face-to-face, PC tablet, HoloLens (AR), Quest (VR) — and the columns are Embodiment in increasing agency — Screen, Robot, Physical, Avatar (AR/VR). Each of the 48 cells is shaded by implementation feasibility on a single-hue scale where darker means more feasible: pale open circle for Difficult, mid-tone half circle with a 135-degree hatch for Partial, dark filled circle with a 45-degree hatch for Feasible. Seventeen cells are feasible, two partial, twenty-nine difficult. The feasible cells cluster in the Avatar column across every interface and space, and along the PC tablet row in Personal and Communal spaces; the Virtual space panel is almost entirely difficult apart from its Avatar column, and the Physical embodiment column is difficult everywhere except face-to-face. Personal Face-to-face PC tablet HoloLens (AR) Quest (VR) Screen Robot Physical Avatar (AR/VR) Communal Screen Robot Physical Avatar (AR/VR) Virtual Screen Robot Physical Avatar (AR/VR) Spaces control / ownership Interfaces Embodiment — agency based on representation Feasibility Feasible Partial Difficult 48 configurations · 17 feasible · 2 partial · 29 difficult
Figure 4. Feasibility of 48 hybrid meeting configurations, Microsoft Research, c. 2022. The original encoded feasibility as green, amber and red; that pairing fails on green-versus-amber for red-green colour deficiency, so it is redrawn here on a single hue stepped light to dark, with hatching and a glyph carrying the value a second and third way.

Two things fall out of it that were not visible in the lattice. The Avatar column is feasible almost everywhere — across every interface and every kind of space. The Physical column is difficult everywhere except face-to-face, which is close to a tautology once stated but was not obvious before: being bodily present is not something an interface can grant you. Between those two, the whole design space collapses onto a single question — if you cannot be there, how well can you be represented?

That is the nimbus axis, discovered a second time and under a different name.

All 48 configurations as a table
Table 9. The full matrix. Columns are Embodiment; rows are grouped by Space, then Interface.
SpaceInterfaceScreenRobotPhysicalAvatar (AR/VR)
PersonalFace-to-faceDifficultDifficultFeasibleDifficult
PersonalPC tabletFeasibleFeasibleDifficultFeasible
PersonalHoloLens (AR)DifficultPartialDifficultFeasible
PersonalQuest (VR)DifficultFeasibleDifficultFeasible
CommunalFace-to-faceDifficultDifficultFeasibleDifficult
CommunalPC tabletFeasibleFeasibleDifficultFeasible
CommunalHoloLens (AR)DifficultPartialDifficultFeasible
CommunalQuest (VR)DifficultFeasibleDifficultFeasible
VirtualFace-to-faceDifficultDifficultDifficultDifficult
VirtualPC tabletDifficultDifficultDifficultFeasible
VirtualHoloLens (AR)DifficultDifficultDifficultFeasible
VirtualQuest (VR)DifficultDifficultDifficultFeasible

Two places the analogy would mislead

At a concert the controller authors the content live, whereas in a shared virtual environment the dataset generally pre-exists and the controller acts upon it. What carries across is not where the content comes from, but who may alter what everyone else is experiencing.

More importantly: at a concert, the remote viewer's absence is physics. In a virtual environment it is a setting. The return channel exists, and somebody chose to leave it off by default. That single substitution — a design decision wearing the costume of a physical constraint — is, I think, the reason hybrid meetings go wrong. The asymmetry looks like a property of the medium, so nobody revisits it.

Open questions

Who holds authority over the view? Someone chooses what the home audience sees, and it is not the band. The vision mixer has authority over presentation without any authority over content: they cannot change a note, but they decide what a million people look at. None of the three roles account for that. It sits sideways to the ladder — less than a Controller over the data, more than an Observer over the framing. In a shared session that role is real: the facilitator who repositions everyone's view, the guide who directs attention, whoever defines the fixed camera positions a passive observer chooses between.

Is nimbus grantable per participant, or only per session? Broadcast bundles "many receivers" with "no return path," but those are separable, and unbundling them is where equitable hybrid participation would have to start.

This paper identified one missing middle role in 2017. Nine years later the arena surfaces two more — the presentation authority, and the participant whose presence is switched off by a default nobody chose deliberately. The recurring problem, across all of it, is the middle.