Unmoderated UX tests
I launched unmoderated UX tests at Dostavista to validate solutions before development. I set up the process using Figma, Maze, and Intercom, documented it with the team in an internal guide, and trained designers and product managers to run the tests independently.
- Brief
- Dostavista
- Role
- project leadership, research, team training
- project leadership
- research
- team training
- Format
Context
Dostavista operated across several international markets, and the product was localized into several languages. This made classic moderated research harder: each market needed a moderator who spoke the users' language fluently and knew the product.
Unmoderated UX tests removed that constraint: we ran them without a moderator and quickly recruited the right audience in each market.
How we ran the tests
We started with the goal: what exactly we wanted to understand and which solution we were testing. Then we formulated hypotheses – concrete assumptions about what the user would notice, understand, or be able to do. From those hypotheses we built tasks, and then a prototype for them. If the scenario was long, we split it into several tasks.
For each task, we planned the main path and alternative paths in advance. This let us see after the test not only whether a person reached the end, but also which path they took. Before launch, we defined the audience, sample size, and success criteria. Since there was no moderator nearby, the task had to be understandable without explanations.
Excerpts from the internal guide. Test setup
Recruiting
We recruited participants through Intercom: selected the right segment and showed the invitation at the right moment. The message briefly explained why we were inviting them, how long the test would take, and what they needed to do.
Excerpt from the internal guide. Participant recruitment
How we evaluated results
We reviewed results on two levels. First, we looked at the scenario as a whole: how many participants completed the task and by which path. Then we reviewed screens with a high number of misclicks. If the share of misclicks exceeded 20%, we checked what share of participants completed that step on the first attempt.
We set team guidelines: if 80% or more complete the task, the scenario works; 60–80% means it is worth checking what can be improved; less than 60% means the scenario should be reconsidered.
Excerpt from the internal guide. Results analysis
Test example: country selection
The app selected the country automatically, but sometimes got it wrong – for example because of a VPN or device settings. We tested three hypotheses: whether the user understands which country is selected and how to change it; whether they know where to look for country and region settings; whether they can change the country through the profile.
Wefast users in India · sample of 100 people
1. Restore the location to Mumbai, India
The app opened with the wrong country and in another language. The task was to restore India and Mumbai. 93 out of 100 participants completed the task; 86 of them – 92.5% – followed the main path.
First task. Key screens
2. Change the region to São Paulo, Brazil
In the second task, participants had to find the region setting inside the app. Already on the Orders screen, misclicks reached 75.8%, and only 46.5% of participants found the right path on the first attempt. The hypothesis that users understand where to look for country and region settings was not confirmed.
After that, the scenario worked noticeably better: Profile – 74.5% completed it on the first attempt, city selection – 81%, country selection – 60.5%. The country change interface was not especially clear for users.
The test helped pinpoint the problem: it was critical to improve the entry point to settings on the Orders screen.
Second task. Key screens
Result
Two of the three hypotheses were confirmed; only the hypothesis about finding the settings was not. At the same time, 73% of participants rated the task difficulty between 1 and 4 out of 10. That meant the problem was not the whole scenario but the entry point: the Orders screen was the part that needed redesign. The study helped separate a local navigation problem from the overall complexity of the scenario and show exactly where refinement was needed.
Task difficulty scale