Case 04 · December 2023
Traffic held. Sales fell 17%
Sales down 17% in three months. Traffic flat. The new checkout had tested well. Nothing looked wrong.
- Role
- UX research, strategy and planning
- Timeline
- 8 weeks, four phases, December 2023
- Team
- Shrey Patel, Daniel Hartmann
- Context
- Client engagement
- Outcome
- A 10% lift per change, named before anything ran
The spark
A 17% decline in online sales across three months. Web traffic stable. The order flow had recently been redesigned and had received positive qualitative feedback.
That combination is the interesting part. Every individual signal looked fine, which is exactly the situation where a team starts guessing, and the gap between stable traffic and falling conversion is where the actual story lives.
The dig
Four methods, layered on purpose, in one eight-week cycle. Usability testing to see behaviour: where people hesitate, misclick, or give up. Online surveys for reasoning at a scale one-on-one sessions cannot reach. A/B testing to move recommendations from “we think this would help” to measured impact. Google Analytics to verify the qualitative findings against real flow and decide what to fix first.
Surveys tell you what people say. Usability testing shows what they do. A/B testing proves whether a fix works. Analytics says how often it matters. None of the four covers its own blind spot, which is the reason for running all of them.
Eight weeks, four phases: screener, moderator script, survey approval and recruitment in weeks one and two; sessions and surveys run and validated against analytics in weeks three and four; A/B tests built from the strongest insights and run in weeks five through seven; findings presented in week eight.
The shift
The obvious suspect was the redesign. It was the most recent change, it touched the exact flow that was underperforming, and blaming it would have been defensible in any meeting.
The plan was written specifically not to assume that. Sales declines are usually multi-causal, and a research plan that goes looking for one culprit will find one whether or not it is guilty. So the scope was the whole purchase experience, not the checkout.
All three problems that surfaced were upstream of the order flow. The information architecture had grown over time and no longer matched how customers thought about products. Product pages carried no styling context, so on a lifestyle brand customers could not picture the item in an outfit, which is most of the decision. And gender and category filters reset on every visit, so returning customers had to re-tell the site who they were.
A research plan that goes looking for one culprit will find one, whether or not it is guilty.
Had we audited the checkout, we would have fixed nothing and reported a success.
The build
The deliverable was a research plan and report, not an interface, so the build here is the recommendation set. Each one is tied to a specific finding and each one is testable.
- Simplify the IA. Merge overlapping categories and sub-categories so browsing follows how customers group products, not how the catalogue grew.
- Style the product. A “styled by the model” section and styling suggestions on every product page, so a garment arrives with the context a lifestyle brand sells on.
- Stop asking twice. Automatic product categorisation by gender preference, so a returning customer is not re-filtering on every visit.
- Existing · Category navigation
- Overlapping categories and sub-categories, hard to move between.
- Proposed · Category navigation
- Merged so browsing follows how customers group products.
- Existing · Product page
- No styling context on a lifestyle brand.
- Proposed · Product page
- Styled by the model, plus styling suggestions.
- Existing · Filters
- Gender and category reset on every visit.
- Proposed · Filters
- The preference is remembered.
Each recommendation traces to one observed finding. None of them touch the redesigned checkout.
Constraints
This was a planning and strategy engagement. The plan was written, scoped, and handed over. We did not run the A/B tests ourselves and there is no post-implementation number, so nothing on this page is a measured result.
The 17% is the client’s figure. We had read access to analytics for verification and no route into the order data behind it, which means the decline could carry causes the four methods were never going to see: stock, pricing, paid spend, a channel drying up. The plan says so.
And the recommendations were never tested against each other. Three fixes tied to three findings is a queue, not a ranked queue, and the ranking would have come out of the A/B phase.
The proof
There is no post-implementation number to report. What the plan does have is a bar, set before anything ran.
Rising pages per visit was the second signal, both tracked through Google Analytics so the team could watch results against baseline in real time rather than waiting for a readout. Research without a measurable outcome is opinion with footnotes. Naming the threshold is what turns a recommendation into something that can be wrong.
The lesson
Naming the failure threshold before running anything was the most useful line in the document. “Below 10%, do not ship it” ended more arguments than any individual finding did, because it moved the decision from whose judgement to trust to what the number was. I write that line into every plan now.