Not a Bigger Discount — A Smarter One

Winning Back Users with AI-Powered Personalized Promotions

Role

UX Researcher (Team of 3)

Skills

Survey Design, A/B Testing, Statistical Analysis (Chi-Square, Welch's t-test)

Timeline

4 Weeks, Survey — A/B Test Analysis

Scope

Simulated client engagement (bootcamp project)

Overview

Cravee is a food delivery platform with strong acquisition but persistent retention problems. We ran a 350-person survey to diagnose loyalty/churn drivers, then designed an A/B test comparing AI-personalized vs random promotions. Result: personalized promotions increased monthly retention by 21–35% over three months, boosted order frequency by 45%, and the effect grew stronger over time — suggesting durable behavioral change rather than a temporary bump.

Users Love Deals. They Don't Love Us.

Despite high ordering frequency, users were switching between apps for better deals, leaving after promotional periods ended, and expressing dissatisfaction with delivery times and food quality.

83%

Serial comparison shoppers

Actively compare multiple delivery apps before placing an order.

70%

Deals are existential

Say deals are very or extremely important. Reducing promotions directly triggers churn.

37%

Top loyalty driver: pricing

Deals and pricing emerged as the top retention lever.

36%

Second driver: reliability

Speed and reliability are equally powerful for retention.

Phase 1: Customer Survey

We deployed a survey to 350 respondents across various age groups, with 67% identifying as Gen Z or Millennial. 83% ordered monthly, and 50% lived alone or with roommates — indicating high-frequency, convenience-driven ordering behavior. Key insight: Non-loyal users are rational shoppers, not disloyal by nature. Personalization emerged as the multiplier: 31% preferred deal-based recommendations, 29% preferred past-order-based, and 76% said they would use a "tap to reorder" feature.

Critical Product Question

Can we design a retention intervention that creates lasting behavioral change — not just a temporary bump from discounts?

Phase 2: A/B Test Design

Target audience: Users who had ordered at least once per month for 3+ consecutive months, had not ordered in the past 2 months, but still opened the app at least once every 14 days. Control group received random restaurant promotion — "$10 off $20 order" at any restaurant. Treatment group received AI-personalized promotion — "$10 off $20" at a restaurant tailored to their order history with contextual cues like "You order here most Fridays." Design: 500 users per group, 5-day activation window, 90 days of subsequent behavior tracking.

Hypotheses & Results

H1Supported

Monthly Retention Rate

Personalized promotions will produce a higher percentage of users placing at least one order per month.

Month 1: +21% (p=0.049) · Month 2: +31% (p=0.011) · Month 3: +35% (p=0.011). Effect grew stronger over time.

H2Direction Positive

Repeat Purchase Rate

Among retained users, those receiving personalized promotions will place 2+ orders per month at a higher rate.

50–70% higher rates across all months, but p-values (0.083, 0.056, 0.138) did not reach significance. Subsample of 100–175 too small. Recommendation: scale to 2,000+ users.

H3Supported

Order Frequency

Treatment group will have higher total orders over the 90-day observation period.

Control: 0.87 avg orders → Treatment: 1.27 avg orders. +45% lift (t=3.60, p<0.001, Cohen's d=0.23). Every individual month also significant (all p<0.003).

Outcome

Before

29%

Monthly retention

0.87

Orders per 90 days

Generic

Random promotions for all

After

35%

Monthly retention (+21%)

1.27

Orders per 90 days (+45%)

AI-personalized

Tailored to behavior patterns

Reflection

The most instructive moment was H2. The repeat purchase rate showed a 50–70% lift — a potentially transformative finding — but didn't reach statistical significance. This taught me the difference between "no effect" and "insufficient power to detect an effect." The recommendation to scale the test was just as important as the confirmed findings, because it prevented the team from prematurely dismissing a promising signal. I also learned that translating statistical results into business narratives matters as much as the analysis itself. Saying "p < 0.001" means nothing to most stakeholders. Saying "for every 100 lapsed users we reach, personalization brings back 7 more than random promotions, and they keep ordering" — that drives decisions.