What's a reasonable QA budget or time allocation, as a percentage of total dev time, for checkout testing on a mid-market Shopify Plus store?
No published figure exists. Here is the honest answer.
Checkout QA is budgeted by test matrix, not as a percentage of development hours. No published benchmark sets a defensible checkout QA share. Capgemini's World Quality Report 2025-26 does not publish one (verified September 2026). Count the combinations you must prove instead: payment methods × discount stacking × tax and duty × customer type × market.
Why the percentage question has no good answer
Percentage-of-dev is an appealing heuristic because it is easy to put in a spreadsheet. It fails here for a structural reason: checkout QA effort is driven by the number of states a customer can reach, and that number has almost no relationship to how many hours were spent building the storefront. A store with a simple theme and five payment methods across four markets has far more checkout QA to do than a store with an elaborate theme selling one product domestically.
We looked for a citable industry figure to anchor a percentage. The most-cited source for quality spend, Capgemini's World Quality Report 2025-26 (published November 2025), does not publish a QA share of budget in its released highlights (verified September 2026). Figures circulating as "industry standard QA is 20–25% of development" do not trace to a checkout-specific source. We are not going to quote you a number we cannot stand behind.
The unit that does work: the test case count
Build the matrix, count the cells, and cost it. For a mid-market Plus store the axes are:
| Axis | Typical values to enumerate |
|---|---|
| Payment method | Shopify Payments, Shop Pay, PayPal, Apple/Google Pay, gift card, manual/offline |
| Discount state | None, automatic, code, stacked automatic + code, free shipping, at-threshold |
| Tax and duty | Domestic tax, tax-exempt customer, duties-inclusive international |
| Customer type | Guest, logged-in, B2B company contact, subscriber |
| Market / currency | One row per active market |
| Fulfillment | Standard shipping, local pickup, split shipment, pre-order |
Multiply, then prune combinations that cannot co-occur. What you are left with is a countable number of scenarios, each of which is a real test with a real cost. That number is defensible in a budget conversation in a way that a percentage never is.
The platform-specific cases that must be in the matrix
- Checkout extensibility. Shopify Scripts stopped running on 30 June 2026, when Shopify deactivated every Script still published (per Shopify Help Center, October 2026). For Plus stores, checkout.liquid, additional scripts and script tags on the Thank you and Order status pages were sunset on 28 August 2025; other stores followed on 26 August 2026 (per shopify.dev, October 2026). Any discount, shipping or payment logic that used to live in Scripts is now a Function or a checkout UI extension, and it needs its own tests.
- B2B. "B2B doesn't support purchase options, such as subscriptions, pre-orders, and try before you buy" (per shopify.dev, September 2026). If you sell B2B, test that the subscription selector cannot reach a B2B cart. This is a real defect class, not a hypothetical.
- Subscriptions. "The order edits API doesn't support subscriptions" and "Subscriptions can't be used with draft orders" (per Shopify's Help Center, September 2026). Both constrain what support can fix post-purchase, so both belong in the ops runbook the QA phase produces.
- Duties. If Managed Markets is on, duties-inclusive pricing changes what the buyer sees at checkout. Test the display, not just the calculation.
What buying more QA actually gets you
Coverage, and evidence. The deliverable from a checkout QA phase should be a matrix with a pass state per cell, dated, re-runnable before each release. If your agency's QA line produces a Slack message saying "checkout tested", you bought a feeling.
When to spend less. A single-market, single-currency, card-only store with no subscriptions and no B2B has a small matrix, and paying a large QA allocation against it is buying test cases that do not exist. Count first; the count tells you the budget.
The Deploi point of view
Our own position, from building on Shopify. Separate from the facts above.
- Our take: Refuse the percentage. Ask for the test matrix and the cell count, and pay for that. A percentage of dev hours is a proxy for effort that has no causal link to checkout risk.
- What we’ve seen: Discount stacking and tax-exempt customers are where defects concentrate, and both are combinations rather than features, which is exactly why they survive feature-by-feature testing and reach production.
- Where we disagree: Agencies quote QA as a share of development because buyers ask for it that way. We think that norm is the reason checkout defects ship: it budgets testing against how much was built rather than against how many states a customer can reach.
- What this page adds: that the parent's QA line item should not be quoted as a percentage at all, the specific axes that define a Plus checkout matrix, and the platform constraints (checkout extensibility dates, B2B purchase options, subscription order edits) that add mandatory cells to it.
Reviewed by Martin Dejnicki, Director of SEO & AI Search. Facts verified 2026-09-13.