Does AI-powered product recommendation (upsell/cross-sell apps) actually lift AOV meaningfully for a mid-market catalog, or is that overstated?
AI product recommendation apps carry no independently verified AOV lift range for Shopify catalogs: every widely quoted figure traces to a vendor case study rather than a holdout test on the merchant's own traffic (verified September 2026). Shopify generates related product recommendations automatically at no cost, so a paid engine has to beat free. Rebuy starts at $25 per month.
Why the published numbers do not answer the question
Search for recommendation lift figures and you will find percentages in the teens, the twenties and occasionally the hundreds. Follow any of them back and the chain ends at a vendor's own case study, a vendor-published statistics roundup, or a third party recycling one of those two. We checked in September 2026 and could not find a public, independently run holdout test on a mid-market Shopify catalog. That absence is the finding.
Our own build-vs-buy analysis records the same caution on the buy side: lift claims come from vendor case studies, and the right response is to insist on a holdout test on your own traffic before renewal (Deploi, August 2026).
The measurement problems that inflate the number
- Widget attribution counts orders that were already happening. If a shopper clicks a recommendation for a product they intended to buy, the widget books the revenue. Attributed revenue is not incremental revenue.
- AOV is the wrong denominator. A recommendation engine that adds a cheap accessory to some orders can raise attach rate and lower AOV at the same time. Revenue per session is the number that reflects the business.
- Mix shifts get read as lift. Launch a recommendation engine in September and compare to August, and you have measured the season.
- The comparison baseline is usually nothing, not free. Shopify already auto-generates related product recommendations at no charge. The honest baseline for a paid engine is the native one, configured, not an empty product page.
What has to be true for lift to be available at all
| Condition | Why it matters | Rough threshold |
|---|---|---|
| Catalog breadth | A model needs somewhere to send the shopper | Thin below a few hundred active SKUs |
| Co-purchase density | Patterns need repeat order evidence | Sparse catalogs produce near-random pairs |
| Price ladder | Cross-sell needs an accessory or a step up that exists | A single price point leaves nothing to attach |
| Surface availability | Cart, checkout and post-purchase carry most of the measured gain | Checkout surfaces are plan-gated on Shopify Plus |
| Data freshness | Stale inventory feeds recommend what you cannot ship | Sync frequency in hours, not days |
Deploi assessment, September 2026, not an official Shopify or vendor position. The plan gate on checkout surfaces is per Shopify's documentation, September 2026.
What the tools cost, verified
Rebuy Personalization Engine builds a plan from packages starting "as low as $25/month," with Platform One shown at $534 per month billed monthly and pricing that scales by orders per month (per rebuyengine.com, verified September 2026). The listing is live, rated 4.7 stars across roughly 791 reviews (verified on two URLs of the same listing, September 2026). Nosto is free to install, requires a Nosto account, and publishes no tier pricing at all, 4.7 stars across 61 reviews (verified September 2026).
How to get an answer for your catalog in one quarter
Run the holdout before you renew, not after. Split traffic, leave the recommendation surfaces off for a meaningful share of sessions, and compare revenue per session rather than AOV. Run it long enough to cross a full purchase cycle. If the vendor will not support a holdout, that is information about the vendor.
The result is often positive and usually smaller than the deck. Both of those things can be true, and a small real number beats a large attributed one because you can budget against it.
The Deploi point of view
Our own position, from building on Shopify. Separate from the facts above.
- Our take: Recommendation lift is real and routinely overstated, and the gap is a measurement artifact rather than a vendor conspiracy. Buy the engine if you want the cart, checkout and post-purchase surfaces and someone to own them. Do not buy it on a case-study percentage.
- What we’ve seen: The recommendation work that pays is unglamorous. Correct product data, sensible exclusions, a rail that does not recommend the item already in the cart, and a slot in the right position on the page. Most disappointing launches we have looked at were not model failures; they were a rail placed below the fold, or a feed with stale availability.
- What it takes: roughly 74 to 160 hours of scoped work for a product recommendations and cross-sell build (directional Deploi estimate from a small sample of engagements, not a measured average).
- Where we disagree: The category publishes lift numbers as though they were properties of the product. Lift is a property of your catalog, your traffic and your measurement design. We would rather hand a client a holdout plan than a benchmark.
- What this page adds: that no independent lift range exists in public, why the published ones inflate, and the catalog conditions that make lift possible before any vendor is chosen.
Reviewed by Martin Dejnicki, Director of SEO & AI Search. Facts verified 2026-09-13.
Where we worked this out
Our decision records