Not a Bigger Discount — A Smarter One
Winning Back Users with AI-Powered Personalized Promotions
Role
UX Researcher (Team of 3)
Skills
Survey Design, A/B Testing, Statistical Analysis (Chi-Square, Welch's t-test)
Timeline
4 Weeks, Survey — A/B Test Analysis
Scope
Simulated client engagement (bootcamp project)
Overview
Cravee is a food delivery platform with strong acquisition but persistent retention problems. We ran a 350-person survey to diagnose loyalty/churn drivers, then designed an A/B test comparing AI-personalized vs random promotions. Result: personalized promotions increased monthly retention by 21–35% over three months, boosted order frequency by 45%, and the effect grew stronger over time — suggesting durable behavioral change rather than a temporary bump.
Users Love Deals. They Don't Love Us.
Despite high ordering frequency, users were switching between apps for better deals, leaving after promotional periods ended, and expressing dissatisfaction with delivery times and food quality.
83%
Serial comparison shoppers
Actively compare multiple delivery apps before placing an order.
70%
Deals are existential
Say deals are very or extremely important. Reducing promotions directly triggers churn.
37%
Top loyalty driver: pricing
Deals and pricing emerged as the top retention lever.
36%
Second driver: reliability
Speed and reliability are equally powerful for retention.
Phase 1: Customer Survey
We deployed a survey to 350 respondents across various age groups, with 67% identifying as Gen Z or Millennial. 83% ordered monthly, and 50% lived alone or with roommates — indicating high-frequency, convenience-driven ordering behavior. Key insight: Non-loyal users are rational shoppers, not disloyal by nature. Personalization emerged as the multiplier: 31% preferred deal-based recommendations, 29% preferred past-order-based, and 76% said they would use a "tap to reorder" feature.
Critical Product Question
Can we design a retention intervention that creates lasting behavioral change — not just a temporary bump from discounts?
Phase 2: A/B Test Design
Target audience: Users who had ordered at least once per month for 3+ consecutive months, had not ordered in the past 2 months, but still opened the app at least once every 14 days. Control group received random restaurant promotion — "$10 off $20 order" at any restaurant. Treatment group received AI-personalized promotion — "$10 off $20" at a restaurant tailored to their order history with contextual cues like "You order here most Fridays." Design: 500 users per group, 5-day activation window, 90 days of subsequent behavior tracking.
Hypotheses & Results
Monthly Retention Rate
Personalized promotions will produce a higher percentage of users placing at least one order per month.
Month 1: +21% (p=0.049) · Month 2: +31% (p=0.011) · Month 3: +35% (p=0.011). Effect grew stronger over time.
Repeat Purchase Rate
Among retained users, those receiving personalized promotions will place 2+ orders per month at a higher rate.
50–70% higher rates across all months, but p-values (0.083, 0.056, 0.138) did not reach significance. Subsample of 100–175 too small. Recommendation: scale to 2,000+ users.
Order Frequency
Treatment group will have higher total orders over the 90-day observation period.
Control: 0.87 avg orders → Treatment: 1.27 avg orders. +45% lift (t=3.60, p<0.001, Cohen's d=0.23). Every individual month also significant (all p<0.003).
Outcome
Before
29%
Monthly retention
0.87
Orders per 90 days
Generic
Random promotions for all
After
35%
Monthly retention (+21%)
1.27
Orders per 90 days (+45%)
AI-personalized
Tailored to behavior patterns
Reflection
The most instructive moment was H2. The repeat purchase rate showed a 50–70% lift — a potentially transformative finding — but didn't reach statistical significance. This taught me the difference between "no effect" and "insufficient power to detect an effect." The recommendation to scale the test was just as important as the confirmed findings, because it prevented the team from prematurely dismissing a promising signal. I also learned that translating statistical results into business narratives matters as much as the analysis itself. Saying "p < 0.001" means nothing to most stakeholders. Saying "for every 100 lapsed users we reach, personalization brings back 7 more than random promotions, and they keep ordering" — that drives decisions.
Thanks for reading!