Case 05 · 2024 · 8 participants

Everybody reordered. Nobody found favourites

Seven of eight people could not find the favourites list. All eight said they were satisfied with the app.

Role
UX research, usability testing
Timeline
2024, three tasks, eight participants
Team
Course research team, TMU
Context
Academic
Outcome
7 of 8 never found favourites
7 / 8 Never found the favourites list
4 / 4 New users failed at least one task without help
8 / 8 Reported being satisfied with the app anyway

The spark

The McDonald’s app is not a product people evaluate. It is a product people use, in a car park, one-handed, deciding in under a minute. That makes it a good subject, because habit hides a lot of friction, and nobody complains about an app they have already learned to work around.

We recruited eight participants across a deliberate split: regular users who order most weeks, and people who had never opened the app. The same three tasks went to both.

The dig

Three tasks, chosen because they cover the app’s three reasons to exist. Place an order for pickup. Find and reorder a previous order. Redeem a reward.

Each session was moderated, think-aloud, on the participant’s own phone where possible, because installing the app on a test device removes the account history that half of these tasks depend on. We logged task success, route taken, and the point at which a participant stopped reading and started hunting.

Reordering worked. It is the flow McDonald’s has clearly invested in, and regular users completed it without hesitating. Everything adjacent to it did not.

The shift

We expected the new-user group to struggle and the regulars to sail through, and to be reporting on an onboarding problem.

What separated the two groups was narrower than that. Regulars had memorised exactly one path each, and outside it they were as lost as the people who had never opened the app. Four of four new users failed at least one task without help. Seven of eight, across both groups, never located the favourites list, which is a feature built specifically for the habit the regulars have.

So the finding is not “new users need better onboarding”. It is that the app rewards a single learned route and punishes any departure from it, and that discovery stops the day you find one thing that works.

The pattern that held across both participant groups.

The build

The deliverable was an audit with prioritised recommendations, ordered by how many participants each finding accounted for.

  • Surface favourites where ordering happens. It exists, it is buried, and it is the exact feature the weekly-order behaviour calls for.
  • Make search return something. Two participants typed an item name, got a dead end, and went back to browsing the full menu rather than trying again.
  • Name the item the way people name it. The Big Mac appeared under a label that matched no participant’s vocabulary, and all eight commented on it unprompted, which is the only unanimous reaction in the study.
  • Give rewards a visible balance. Redemption assumed you already knew what you had.

Constraints

Eight participants and three tasks. That is enough to say a problem is real and not enough to say how often it happens, so every number here is a count of participants, not a rate.

We tested on the live production app, which means we tested one build in one region on one day. Menu labels, promotions, and the rewards balance all change server-side, and at least one of our findings is a content decision that could be different by the time this is read.

There was no route to implementation and no follow-up round. The recommendations were never put back in front of the people who failed, so the audit shows what broke and not whether the fixes work.

And the satisfaction question was asked at the end of the session, immediately after the tasks, which is the worst possible moment for an accurate answer.

The proof

The last two numbers belong in the same sentence. Every new user failed a task, and every participant reported satisfaction. Two of eight hit a search dead end. All eight were frustrated by the Big Mac label and all eight still said the app was fine.

Reported satisfaction measures how well people have adapted, not how well the app works.

Line the team kept returning to in the readout

The lesson

I stopped asking participants whether they were satisfied. The answer is consistently yes, from the same person I have just watched fail a task twice, and it is not a lie: they have adapted, and the adapting is invisible to them. Now I ask what they did the last time it did not work. That question gets an answer that matches the screen recording.