agenciestestingregressionshopifywcagdeveloper

Shopify Accessibility Testing for Agencies: Handoff Checklist and Regression Checks

Shopify accessibility testing for agencies: a handoff checklist, regression checks after theme updates and app installs, a test script, and how to report findings to a client.

By Radoslaw Fedorczuk10 min read

Shopify accessibility testing at an agency usually happens once, before launch. Then the client installs a reviews app, Shopify offers a theme update, a marketer adds a countdown banner, and nobody tests again. Three months later the keyboard path to checkout is broken and the email lands on your desk.

This guide is for the developer who owns that problem: what to check before handoff, how to catch regressions after theme updates and app installs, a small script you can run against a store, and how to write findings up for a client without promising more than a test can show.

Who this is for

  • Agencies and freelancers who build or maintain Shopify themes, including Dawn-based builds and fully custom ones.
  • In-house teams who inherit a store from an agency and need to know what state it is in.
  • Anyone who has to answer a client's question "is our store accessible?" with something better than a score.

Why a Shopify store regresses

  • Theme updates. When a merchant updates a theme, Shopify adds the new version as "Updated copy of" to the draft themes. Settings, sections and app embed settings are copied. Code edits are carried over only when they do not conflict with the update; otherwise the theme card says the code edits could not be included, and your fixes stay behind in the old copy.
  • Apps. An app adds markup and scripts through app embeds and app blocks, and older apps inject code directly. Review widgets, upsell carousels, cart drawers, chat bubbles and cookie banners all bring their own focus handling, or none.
  • Content. New sections, banners with text baked into images, videos without captions and alt texts left empty arrive through the theme editor, not through your repository.
  • Campaigns. Pop-ups, countdowns and auto-rotating sliders appear for a sale and stay.

None of these go through your code review, so none of them get your tests unless you schedule them.

A 90-second smoke test after every change

Run it on the home page, a product with variants, and the cart, on the draft theme before it is published.

  1. Put the mouse away and press Tab from the top. The skip link appears first, and every stop after it is visible.
  2. Open the menu, choose a variant and add to cart with the keyboard alone. Focus moves into the cart drawer and returns when you press Escape.
  3. Check the new thing, whatever changed: the app widget, the new section, the pop-up. Can you reach it, use it and leave it with the keyboard?
  4. Zoom to 200% and resize to a phone width. Nothing is cut off, and the page does not scroll sideways.
  5. Turn on VoiceOver or NVDA and listen to the new element once.

If a step fails, the most recent change is the first suspect.

The handoff checklist

Hand this over with the theme, filled in, together with the report described further down. The criteria numbers link to the W3C explanations.

Structure

  • One h1 per template and no skipped heading levels (1.3.1).
  • header, nav, main and footer landmarks, and a skip link whose target exists and takes focus (2.4.1).
  • A distinct title on every template (2.4.2).
  • lang on the html element that follows the market's language. Dawn writes lang="{{ request.locale.iso_code }}" (3.1.1).

Keyboard and focus

  • Every control reachable and operable with Tab, Shift+Tab, Enter, Space and the arrow keys (2.1.1).
  • A visible focus indicator everywhere, with no outline: none left without a replacement (2.4.7).
  • Menus, drawers and dialogs move focus in, close on Escape and return focus (2.4.3, 2.1.2). Our article on cart drawer accessibility has code for this.
  • No positive tabindex.
  • Sliders and anything that moves for more than five seconds can be paused (2.2.2).

Forms and dynamic content

  • A programmatic label on every field: search, newsletter, contact, filters, quantity (1.3.1, 3.3.2).
  • Error messages in text, next to the field, saying how to fix the input (3.3.1).
  • Cart updates, variant price changes and filter result counts announced through a status element that exists on page load (4.1.3). The variant picker article shows the pattern.

Visual

  • Text contrast of at least 4.5:1, or 3:1 for large text (1.4.3), including text on image banners and in sale badges.
  • Controls, input borders and focus indicators at 3:1 against their background (1.4.11).
  • Content usable at 320 CSS pixels wide without scrolling sideways (1.4.10), and with increased text spacing (1.4.12).
  • Tap targets of at least 24 by 24 CSS pixels (2.5.8).

Media and content

  • Alt text on product images that names what the shopper needs to know, alt="" on decorative images (1.1.1).
  • Captions on videos with speech (1.2.2); no sound that starts by itself.
  • A short content guide for the client: how to write alt text, why text belongs in a text block rather than inside an image, which colour settings keep contrast.

Third parties

  • A list of every app embed, app block and injected script, with the vendor and the flows it touches.
  • Each of them tested with the keyboard in the smoke test above. A barrier inside an app is still a barrier in the client's store.

Catching regressions automatically

Manual tests catch the most, and they depend on someone remembering to run them. A script run on a schedule, or on the draft theme before publishing, catches the rule-based part every time. This one uses Playwright and the open source axe-core rule engine, and adds one keyboard check that breaks often in Shopify themes:

// a11y-regression.mjs
// Usage: node a11y-regression.mjs https://your-store.myshopify.com
import { chromium } from 'playwright';
import axe from 'axe-core';

const base = process.argv[2];
const pages = ['/', '/collections/all', '/products/example', '/cart', '/search?q=shirt'];
const browser = await chromium.launch();
const page = await browser.newPage();
let problems = 0;

// 1. Rule-based checks on each template.
for (const path of pages) {
  await page.goto(new URL(path, base).href, { waitUntil: 'load' });
  await page.addScriptTag({ content: axe.source });
  const { violations } = await page.evaluate(() =>
    window.axe.run({ runOnly: ['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa', 'wcag22aa'] })
  );
  for (const v of violations) {
    problems += 1;
    console.log(`${path}  ${v.id}  ${v.impact}  ${v.nodes.length} element(s)`);
  }
}

// 2. One keyboard check that breaks often: Add to cart moves focus into the drawer.
await page.goto(new URL('/products/example', base).href, { waitUntil: 'load' });
await page.locator('[name="add"]').focus();
await page.keyboard.press('Enter');
await page.locator('[role="dialog"]').first().waitFor({ state: 'visible' });
const focusInDrawer = await page.evaluate(
  () => !!document.activeElement?.closest('[role="dialog"]')
);
if (!focusInDrawer) {
  problems += 1;
  console.log('/products/example  cart drawer opened without moving focus into it');
}

await browser.close();
console.log(problems
  ? `${problems} problem(s) found`
  : 'No problems found by these checks');
process.exit(problems ? 1 : 0);

Install it with npm install playwright axe-core and npx playwright install chromium, replace /products/example with a real product handle, and point it at a preview URL of the draft theme before the merchant publishes it. For a password-protected development store, submit the storefront password on /password in the same browser session first. The script exits with code 1 when it finds something, so it can fail a CI job or a scheduled task.

I ran it against two local test stores built from the same five templates. On the one with a correct drawer and correctly sized links it printed "No problems found by these checks". On the other it reported a button-name violation on a nameless icon button, target-size violations on small inline links, and "cart drawer opened without moving focus into it" for a drawer opened with nothing more than this:

// The drawer appears, focus stays on the Add to cart button behind it.
addButton.addEventListener('click', (event) => {
  event.preventDefault();
  document.getElementById('CartDrawer').hidden = false;
});

Keep the baseline output from launch in the project. When a later run differs, the diff tells you which template and which rule changed, and the date tells you which update or app install to look at.

How to report findings to a client

A report is useful when the client can act on it and can reuse it later, for an accessibility statement or when a customer complains. For each finding, write down:

ID:            CART-03
Page:          /products/linen-shirt (templates/product.json)
Finding:       The quantity buttons have no accessible name. NVDA reads
               "button" twice, with no product or direction.
Criterion:     WCAG 2.2, 4.1.2 Name, Role, Value (Level A)
Affects:       screen reader users, voice control users
Found with:    keyboard and NVDA with Firefox; confirmed in the accessibility tree
Fix:           visually hidden text "Increase quantity for {{ item.title }}"
               in snippets/cart-drawer.liquid
Owner:         agency (theme code)
Status:        fixed in the draft theme, retested on 2026-10-06

Around the findings, the report needs:

  • Scope: the date, the pages and templates tested, the theme name and version, the apps installed at the time, browsers and assistive technology used.
  • Method: what was automated, with which tool and rule set, and what was tested by hand.
  • Owner per finding: agency (theme code), client (content and settings) or app vendor. A client cannot fix a vendor's widget, but they can send the vendor your finding.
  • What was not tested: checkout, account pages, pages behind a login, languages not reviewed.
  • Wording: describe what you measured on which date. Write "the quantity buttons have no accessible name", not "the store is non-compliant". The W3C's WCAG-EM methodology is a good model for how an evaluation defines its scope and sample.

Scores go in an appendix, if anywhere. A client who sees "92/100" on page one reads it as nearly done, whatever the findings below it say.

What AccessifyAI detects

The scanner reads the HTML each page sends, without running JavaScript, and checks it against rules mapped to 15 of the 55 WCAG 2.2 A and AA criteria. How we test lists every rule, its severity and its limits. For agency work:

  • A scan reads up to 10 pages on the free plan, 50 on Basic and 500 on Pro, following links inside the store.
  • Each finding comes with its WCAG criterion, the element and the theme file or section it most likely comes from. The app reads the theme and cannot change it, so the fix goes through your normal workflow.
  • Scheduled scans compare each result with the previous one and flag a regression when the score drops by more than 5 points or a new critical finding appears. That covers the rule-based part of the regressions described above, for example a new app that adds unnamed buttons or a theme update that brings back a missing label.
  • It does not test keyboard flows, focus movement or announcements. The smoke test and the script above cover what it does not.

What no automated tool can judge

  • Whether a shopper can finish a purchase with a keyboard, a screen reader or at 400% zoom.
  • Whether alt texts, link texts and headings say the right thing.
  • Whether an app's widget makes sense when read aloud.
  • Whether error messages help someone correct a form.
  • Whether a design change made the store harder to use for people with cognitive disabilities.

W3C's page on involving users in evaluation explains what testing with disabled users adds on top of expert review. For a larger client, budget for it.

Try it on a client store

Run the free one-page scan on any page of a store to see the rule-based findings. To scan whole stores on a schedule, install AccessifyAI from the Shopify App Store on the client's store. It helps find and fix issues and is not a legal guarantee, so the manual checks stay in your process.

Frequently asked questions

How often should an agency retest a Shopify store?

After every theme update, app install or app update, and redesign, on the draft theme before publishing. Between those, a scheduled automated scan catches rule-based regressions, and a manual smoke test of the purchase path once a quarter catches the rest.

Can an agency promise a client WCAG compliance?

Promise the work and describe the result: which criteria you tested, on which pages, on which date, and what you fixed. A store changes every time the client edits content or installs an app, so a statement about its conformance holds only for the version you tested. Avoid legal wording such as "compliant" in reports and proposals.

Who is responsible when an app causes the barrier?

The store is the client's, so the barrier is in the client's store whatever its source. In the report, name the app and the element, mark the vendor as owner, and suggest either a setting that avoids the problem, a replacement app, or a message to the vendor with the finding.

Is axe-core enough for regression testing?

It is a solid rule engine and a good CI check. Like every automated tool, it covers part of WCAG. Pair it with at least one scripted keyboard check for the flows that matter, and with a manual smoke test after each change.

What should be in a Shopify accessibility handoff?

The filled-in checklist, the dated report with owners per finding, the list of apps and injected scripts, a short content guide for the client's team, and the baseline output of your automated checks.

Sources

Share:

Get accessibility tips by email

One short email per month with new guides and Shopify accessibility updates. No spam, unsubscribe anytime.