Elena Perminova
UX icon

Unmoderated UX tests

I launched unmoderated UX tests at Dostavista to validate solutions before development. I set up the process using Figma, Maze, and Intercom, documented it with the team in an internal guide, and trained designers and product managers to run the tests independently.

Brief
Dostavista
Role
project leadership, research, team training
  • project leadership
  • research
  • team training
Format

Context

Dostavista operated across several international markets, and the product was localized into several languages. This made classic moderated research harder: each market needed a moderator who spoke the users' language fluently and knew the product.

Unmoderated UX tests removed that constraint: we ran them without a moderator and quickly recruited the right audience in each market.

Borzo geography map

How we ran the tests

We started with the goal: what exactly we wanted to understand and which solution we were testing. Then we formulated hypotheses – concrete assumptions about what the user would notice, understand, or be able to do. From those hypotheses we built tasks, and then a prototype for them. If the scenario was long, we split it into several tasks.

For each task, we planned the main path and alternative paths in advance. This let us see after the test not only whether a person reached the end, but also which path they took. Before launch, we defined the audience, sample size, and success criteria. Since there was no moderator nearby, the task had to be understandable without explanations.

Guide fragment: defining the research goal and hypotheses
Guide fragment: preparing a prototype for the tasks

Excerpts from the internal guide. Test setup

Recruiting

We recruited participants through Intercom: selected the right segment and showed the invitation at the right moment. The message briefly explained why we were inviting them, how long the test would take, and what they needed to do.

Guide fragment: sending test invitations through Intercom

Excerpt from the internal guide. Participant recruitment

How we evaluated results

We reviewed results on two levels. First, we looked at the scenario as a whole: how many participants completed the task and by which path. Then we reviewed screens with a high number of misclicks. If the share of misclicks exceeded 20%, we checked what share of participants completed that step on the first attempt.

We set team guidelines: if 80% or more complete the task, the scenario works; 60⁠–⁠80% means it is worth checking what can be improved; less than 60% means the scenario should be reconsidered.

Guide fragment: analyzing results and task success criteria

Excerpt from the internal guide. Results analysis

Test example: country selection

The app selected the country automatically, but sometimes got it wrong – for example because of a VPN or device settings. We tested three hypotheses: whether the user understands which country is selected and how to change it; whether they know where to look for country and region settings; whether they can change the country through the profile.

Wefast users in India · sample of 100 people

1. Restore the location to Mumbai, India

The app opened with the wrong country and in another language. The task was to restore India and Mumbai. 93 out of 100 participants completed the task; 86 of them – 92.5% – followed the main path.

First country selection task: key screens of the scenario

First task. Key screens

2. Change the region to São Paulo, Brazil

In the second task, participants had to find the region setting inside the app. Already on the Orders screen, misclicks reached 75.8%, and only 46.5% of participants found the right path on the first attempt. The hypothesis that users understand where to look for country and region settings was not confirmed.

After that, the scenario worked noticeably better: Profile – 74.5% completed it on the first attempt, city selection – 81%, country selection – 60.5%. The country change interface was not especially clear for users.

The test helped pinpoint the problem: it was critical to improve the entry point to settings on the Orders screen.

Second country selection task: key screens of the scenario

Second task. Key screens

Result

Two of the three hypotheses were confirmed; only the hypothesis about finding the settings was not. At the same time, 73% of participants rated the task difficulty between 1 and 4 out of 10. That meant the problem was not the whole scenario but the entry point: the Orders screen was the part that needed redesign. The study helped separate a local navigation problem from the overall complexity of the scenario and show exactly where refinement was needed.

Country selection test results: task difficulty scale

Task difficulty scale

Next project: Neurofeedback training scenarios