30 Years of Game UX: From Watching Players to Predicting Them

Thirty years took game research from observation rooms to synthetic players. The tools changed. The hardest question did not: which friction belongs to the game?

July 8, 2026 · 18 min read · Game UX, UX History, UX Research

A banking app should never make me fight for a transfer. A game absolutely should make me fight for a sword.

A boss fight should resist you. A puzzle should withhold its answer. An inventory can force a decision under pressure. Remove every obstacle and you do not get perfect UX. You get a game with nothing left to play.

That gives game UX a strangely surgical job. We remove the friction created by unreadable icons, broken navigation and unclear feedback without cutting into the challenge, tension and discovery the game was built around.

The tools for finding that boundary have changed dramatically. We have gone from watching players through two-way mirrors to tracking millions of sessions, measuring physiological signals and experimenting with synthetic players.

But the central question has survived all of it:

Which friction belongs to the game, and which friction did the interface create?

The 1990s: mirrors, clipboards and thinking out loud

Before game user research became a recognised discipline, testing was largely organised around a different question: does the game work?

QA could find a broken trigger, a crash or an enemy stuck inside a wall. It could not necessarily explain why a perfectly functional inventory made new players feel lost.

By the late 1990s, formal user-research groups were bringing the usability lab into game development. A typical setup placed a player in front of the game, cameras around the room and researchers behind a two-way mirror. The player was often asked to think aloud, explaining each decision while trying to play.

Try narrating your cognitive process during Doom. Fun for the researcher. Slightly less natural for the person holding the shotgun.

A schematic reconstruction of a 1990s usability lab: moderator behind the glass, player at the CRT and a two-way mirror between them.

Think-aloud testing gave researchers access to something telemetry still struggles to provide: the player's interpretation of a moment. Not only where they clicked, but what they believed would happen when they clicked it.

The method had obvious limits. The room was artificial. The equipment was expensive. Speaking changes how people perform a task. But observation made one uncomfortable fact visible: a system can behave exactly as designed and still fail the player.

Jakob Nielsen's discount-usability work helped popularise a more practical lesson: small qualitative rounds can reveal recurring failures early. Five participants are not a universal sample size, but you rarely need hundreds of players to discover that nobody can find the crafting menu. You need the right participants, clear tasks and another round after the fix.

The 2000s and 2010s: when observation became data

As game production grew, user research became more specialised. Observation rooms did not disappear, but researchers gained new ways to examine what players could not easily put into words.

The question expanded. It was no longer only "can the player use this?" Researchers also wanted to know where attention went, when tension rose and whether difficulty felt exciting or simply confusing.

Players could now be observed through more than words and behaviour. Researchers experimented with skin conductance, heart rate and, in some studies, EEG.

These signals opened another window into the experience, but not a transparent one. A spike in skin conductance may show that arousal changed. It cannot tell us whether the player felt fear, delight, frustration or the sudden need to sneeze.

The useful question is not "what does the biometric signal mean?" in isolation. It is "what happened at this moment, and what else did the player do?"

A repeated spike combined with hesitation, a failed interaction and later abandonment gives the team somewhere worth investigating. Biometrics do not deliver a UX verdict. They place a marker on the timeline.

At the same time, telemetry changed the scale of observation. Researchers no longer had to rely only on what happened inside a scheduled session. Live games could reveal where players died, which routes they ignored, when they abandoned a tutorial and which systems they never touched.

A heatmap can show that hundreds of players died near the same doorway. It cannot tell us whether the cause was poor signposting, an unfair encounter or a perfectly intentional difficulty spike. Scale did not remove the need for interpretation. It made interpretation more urgent.

An illustrative death heatmap, not a real dataset. The cluster identifies a moment to investigate, not its cause.

Left 4 Dead's AI Director is my favourite inversion of this idea. Instead of collecting player data for a report, the game estimates player intensity while the session is still happening, then adjusts the population of threats to shape peaks and valleys in the pacing.

The system observes the player and changes the experience before a researcher ever opens a report.

The 2020s: remote testing and synthetic players

The pandemic pushed many playtests out of controlled labs and into players' homes. Researchers lost some control over hardware, observation and test conditions. In return, they gained a clearer view of the environment games are actually shipped into: distracting rooms, unstable connections, personal setups and nobody sitting nearby with a clipboard.

A player behaves differently on their own couch. That does not make remote testing inherently better, but it reveals kinds of friction a controlled lab can hide.

Then the observer stopped being the only thing that changed. The player became synthetic too.

LLM-based personas can answer interview questions and inspect interface copy, while tool-using agents can attempt scripted journeys in minutes. Together, they make an attractive promise: test earlier, test faster and discover problems before recruiting a single participant.

But plausible language is not the same as plausible behaviour.

A 2026 ACL benchmark evaluated LLM agents against 31,865 real online shopping sessions. Prompt-based agents predicted the next human action with only 11.86% accuracy. Shopping is not playtesting, and the result cannot be transferred directly to games. It does expose the central risk: an agent can explain its decision convincingly without behaving like the person it claims to represent.

Synthetic players are a linter, not a playtest.

Used carefully, synthetic players may help expose structural problems early: contradictory instructions, dead ends in navigation, missing states or a disclosure flow that never actually discloses.

What they cannot currently validate is the experience itself. An agent does not feel the dread of Metro's oxygen meter or the anticipation of choosing a Hades boon. It also has no personal history with a genre, no tired hands and no reason to abandon the game because something else looked more fun.

Real players do not arrive to complete our test successfully. They arrive with their own expectations, habits and motivations. That inconvenience is exactly what makes them valuable.

Use synthetic players to clean the flow before real people arrive. Never use them as evidence that the experience works.

Measuring the experience, not just the product

At scale, teams naturally measured what was easiest to count. Page views, uptime, latency, seven-day active users and earnings were grouped under the conveniently anatomical acronym PULSE.

These metrics could describe the health of a product. They were much weaker at explaining the quality of the experience. A player can open the inventory every session because it is useful, because it is confusing or because the quest marker refuses to disappear until they do.

In 2010, Google researchers introduced HEART: Happiness, Engagement, Adoption, Retention and Task Success. It did not replace operational metrics. It gave teams a way to connect product goals with human outcomes.

Applied to games, HEART might look like this:

HEART translated into possible game metrics. The framework helps connect product goals to human outcomes; the examples are an adaptation, not a studio standard.

HEART does not tell a team which metrics matter. It forces the team to define what success should mean before reaching for the dashboard.

That distinction matters in games. A shorter session might signal frustration. It might also mean the player completed a satisfying chapter and stopped at exactly the right moment.

Product analytics also gave us useful names for small behavioural signals: rage clicks, dead clicks and hesitation time.

Games may not use those exact labels, but they can instrument equivalent behaviour. Repeated attempts to activate an element, a cursor moving between two nearly identical icons or a menu opened and closed without action can all point to friction.

They are evidence to investigate, not proof that the player failed.

Behaviour still does not explain motivation. Two players can perform the same action for completely different reasons, which is where psychometric models become useful.

The Player Experience of Need Satisfaction model, or PENS, grew from Self-Determination Theory. Its motivational core examines how play supports three psychological needs:

These are not three ingredients every interface must maximise. They are three questions that can help explain why an interaction feels motivating or alienating.

A status icon with no readable meaning can undermine competence: the player has information but cannot turn it into understanding.

Flexible loadouts, remappable controls and meaningful dialogue choices can support autonomy by preserving ownership over play.

Relatedness is usually discussed through relationships with other people or characters. My favourite edge case is Disco Elysium. Its 24 skills interrupt, argue and occasionally betray you until fragments of the protagonist's psyche begin to feel like members of the party.

That is my reading rather than a textbook PENS example, but it shows how an interface can carry a relationship, not merely display one.

Diegetic, spatial, meta: where information lives

In 2009, Erik Fagerholt and Magnus Lorentzon proposed a design space for first-person shooter interfaces based on two questions: does the information exist inside the fictional world, and is it represented in the game's three-dimensional space?

The familiar shorthand of non-diegetic, diegetic, spatial and meta UI draws from that work. It is useful, but the original model was more nuanced than the neat 2×2 diagram it later became.

Four examples of the familiar UI taxonomy, highlighting the dominant treatment in each frame. The categories are useful shorthand, but real game interfaces often move between them.

Non-diegetic UI exists outside the fictional world. Health bars, minimaps and Devil May Cry's combo counter speak directly to the player, not the character.

Its strength is clarity. Information can stay stable and readable regardless of camera position, lighting or environmental noise. Its cost is not automatically less immersion. The real cost is that the interface must earn its visual relationship with a world it does not physically belong to.

Diegetic UI exists inside the fiction and can be perceived by the character. Dead Space places health on Isaac's spine and ammunition on the weapon. Metro turns a wristwatch and gas-mask filter into survival instruments. Far Cry 2 puts navigation inside the car.

This can make information feel inseparable from the world, but the world is not a controlled UI canvas. Lighting changes. Objects turn away from the camera. Combat refuses to wait until the player has found a readable angle.

A diegetic health bar still has to survive a dark corridor.

Dead Space puts health on Isaac's spine, stasis on the suit and ammunition on the weapon. The HUD is not hidden; it is relocated into the fiction.

Spatial UI occupies the game world without necessarily belonging to its fiction. Quest markers float over characters. Loot receives an outline. Witcher Senses place readable emphasis directly onto the environment.

Spatial UI reduces the distance between information and the object it describes. But it can also turn a world into a layer of floating labels if every object demands attention at once.

Meta UI represents something the character experiences without giving it a stable physical location in the world. Blood reaches the edge of the screen. Audio becomes muffled after an explosion. Cyberpunk 2077 distorts the interface when V's cyberware or perception destabilises.

The information belongs to the character's condition, but its presentation belongs to the player's screen.

None of these categories is inherently more advanced or immersive.

Dead Space relocates information into the fiction. Persona 5 does the opposite: its menus are unapologetically non-diegetic, then stylised so aggressively that they become part of the game's identity. Competitive shooters preserve abstract HUDs because milliseconds of readability matter more than pretending the ammunition counter physically exists.

Most interfaces move between categories depending on context.

The useful question is not 'which type is best?' It is 'where can this information live without betraying the promise of the game?'

The psychology under the pixels

By this point, the tools had changed repeatedly. The human constraints underneath them had not. Some friction is cognitive, some physical and some emotional.

Hick's Law describes a familiar pressure: decision time tends to increase as the number and complexity of choices grow.

RPG inventories are almost laboratory conditions for this problem. Hundreds of items compete inside a limited visual hierarchy, but simply removing choices would also remove part of the fantasy. The player is supposed to compare, prepare and decide. The interface should not make them decode the organisation before they can begin.

The Witcher 3's patch 1.20 reorganised the inventory into clearer categories, added item comparison and revised its tooltips. The number of meaningful decisions did not disappear. More of the effort could finally go into making them.

Good information architecture does not eliminate choice. It protects the player's attention for the choice that matters.

The Witcher 3 inventory before and after patch 1.20: a real example of hierarchy, comparison and chunking being revised after release.

Fitts's Law deals with acquisition cost: targets generally become faster to reach when they are larger and closer.

Games complicate the model because a mouse pointer, an analogue stick and a thumb do not move in the same way. Still, the design question survives: how much physical effort and precision does this action demand?

A weapon wheel places every option around a shared origin and gives each one a large directional sector. On mobile, critical controls usually belong inside a comfortable reach zone rather than in a beautiful but painful corner.

Neither pattern removes skill from the game. It removes unnecessary motor negotiation with the interface.

Two approaches to acquisition cost. Different input models, shared design question: how much physical effort and precision does the action demand?

Donald Norman's three levels of emotional design offer another useful lens:

Hades often moves through all three in a single reward. A boon arrives with colour, sound and animation. Its effect immediately changes play. Over time, the choice becomes part of the build the player remembers.

The interface is not merely reporting the reward. It helps construct the feeling of receiving one.

Guidance, manipulation and the line between them

As game worlds became denser and more visually detailed, interactive geometry stopped standing out on its own. A ladder could be climbable, decorative or simply part of the rubble. The player needed a signal.

Yellow paint became the loudest version of that signal. Resident Evil 4 Remake and Final Fantasy VII Rebirth mark climbable surfaces with conspicuous colour, and some players read the result as useful guidance while others see hand-holding.

Underneath the argument is a real production problem: traversal affordance has to be communicated somehow.

Paint is cheap, consistent and difficult to miss. Lighting, composition, material contrast and micro-animation can preserve more of the fiction, but they require tighter environmental control and may still fail under changing weather, camera angles or accessibility needs.

The answer is rarely one universal colour or one perfectly subtle beam of light. It is layered guidance, adjustable where possible and tested with the audience the game actually serves.

The goal is not to make guidance invisible. It is to make the player notice the path before they notice the technique.

The same route, two signals. Paint is louder; light asks more from the environment. Neither works independently of context.

Guidance becomes something else when the interface stops helping the player understand a decision and starts making the decision harder to resist.

Dark patterns hide real prices behind layered currencies, manufacture urgency through misleading countdowns or frame refusal as personal failure: "No thanks, I prefer losing."

Harry Brignull began documenting deceptive interface patterns in 2010. Games gave some of them especially fertile ground: variable rewards, time-limited events, social pressure and currencies whose real-world value is deliberately difficult to calculate.

These patterns can increase conversion without improving the player's experience. That is precisely why ordinary engagement metrics are not enough. A system can be commercially effective and still reduce informed choice.

For online platforms within its scope, the European Union's Digital Services Act explicitly prohibits deceptive design practices. But compliance is a low bar. A technically legal store can still spend player trust faster than it earns revenue.

If conversion requires the player to misunderstand the choice, the interface is not persuasive. It is deceptive.

Yellow paint and a fake countdown timer both direct attention, but they do not serve the same relationship.

One tries to make the game legible. The other tries to make refusal harder.

The ethical question is not whether an interface influences behaviour. Every interface does. It is whose intention that influence serves.

The next interface adapts

The clearest direction is not a visual style. It is adaptability.

Accessibility is increasingly treated as part of the game's design, not an exception added afterwards. The Last of Us Part II shipped with more than 60 accessibility settings, including features designed for blind and low-vision players. Frameworks such as AbleGamers' Accessible Player Experiences move the conversation beyond minimum compliance towards multiple ways of perceiving, understanding and operating the same game.

Many of these features travel further than their original use case. Directional subtitles support players in a noisy room. Alternatives to repeated button presses reduce fatigue. Auto-pickup removes repeated motor effort that may have nothing to do with the challenge the game is trying to create.

Celeste's Assist Mode makes the philosophy unusually explicit. Game speed, stamina, dashes and invincibility become adjustable, but the game explains the intent before presenting the controls.

Accessibility gives more players a way to negotiate the challenge on their own terms.

Celeste explains the philosophy of Assist Mode before presenting its controls. The player receives context, then agency.

Adaptability can also become contextual. Interfaces already respond to combat state, input method, screen size and player performance. The more difficult question is how far that response should go.

A system might enlarge frequently missed targets, delay non-critical notifications during high cognitive load or change how instructions are delivered after repeated failure. Done well, adaptation removes accidental friction. Done badly, it becomes invisible interference: the game changes, but the player no longer understands why.

An adaptive interface therefore needs limits. It should preserve agency, avoid pretending certainty about the player and expose meaningful changes when those changes affect how the game behaves.

A more speculative constraint is emerging alongside adaptation: interfaces may need to become legible to machines as well as people.

If players increasingly use external assistants to compare builds, navigate quests or understand complex economies, the underlying game data matters. In a build-heavy ARPG, an assistant that can read structured item attributes is less likely to invent a stat than one forced to interpret a compressed screenshot.

This is not yet a mature game UX discipline. It may never belong entirely to interface designers. But it suggests that future information architecture will have another audience: systems trying to help the player interpret the game.

The synthetic player and the player's assistant are easy to confuse, but they serve opposite sides of the relationship. One claims to stand in for the player. The other stands beside them.

The question that survived the tools

A two-way mirror, a biometric sensor, a telemetry dashboard and a synthetic player do not reveal the same thing.

The mirror shows behaviour in context. Biometrics mark a change without explaining it. Telemetry finds patterns at scale. Synthetic players may expose structural problems before a real participant arrives. Each tool makes part of the experience visible, and each leaves something out.

Thirty years of new methods have not automated the designer's judgement. They have made that judgement better informed.

We still have to decide whether hesitation means meaningful tension or an unreadable choice. Whether a difficult interaction belongs to the fantasy or merely stands between the player and it. Whether adaptation supports agency or quietly takes it away.

A banking app should not make me fight for a transfer. A game can make me fight for a sword, learn its weight and decide whether it was worth carrying home.

Protect the good friction. Kill the bad friction.

The difficult part has always been knowing which is which.

Read the full case study in the portfolio →