Purpose: Turn A/B testing from a series of disconnected experiments into a programme that compounds — where test forty is smarter than test four — by using Claude and Cowork to do the connective thinking that teams normally skip.

Who this is for: Ecommerce operators, CRO leads, and agency teams who are already testing (or about to start) and keep ending up with a spreadsheet of wins and losses that never adds up to a strategy.

What you'll build: A repeatable testing loop plus one living document — a purpose ledger — that records what every test taught you about your customers. Optionally, you'll also wrap the repeatable parts into Claude Skills so the loop runs the same way every time.

Time estimate: Half a day to set the programme up. Two to four hours of human time per test after that, spread across the test's life. The test itself runs for as long as it needs to reach a reliable result — that part you don't rush.


How A/B testing with AI actually works (the 30-second version)

A/B test (also called a split test): you show half your visitors the current page (the control) and half a changed version (the variation), then decide which is better from how people actually behave — not from what someone predicted in a meeting.

The mechanics have been solved for years. The part that fails is the thinking around them. Most teams test each thing on the site in isolation, get a result, and move on, so nothing accumulates — month six looks exactly like month one. That pattern is common enough to have a name: tunnel-vision testing. It's why a year of technically perfect tests can leave you unable to answer the only question that matters — what do our customers actually care about?

Here's the reframe this SOP is built on. The tempting use of AI is to generate variations faster. That's the cheap half of the work, and speeding it up mostly gets you more disconnected tests. The valuable half is the connective thinking: classifying tests so patterns emerge, framing them without bias, and reading results in a way that feeds the next test. That half gets skipped because it's slow and dull — which is exactly the half an AI can help you sustain.

Two ideas hold the whole thing together:

That's the mental model. The phases below are how you run it.


Phase 1: Set up the programme

Goal: Establish whether testing is worth it, take stock of what you already know, and lock the metrics you'll measure on every test from here on.

Step 1: Confirm it's worth running

Check your revenue and traffic against the floor in Before you start. If you're under it, stop and put the effort into traffic instead — there's usually bigger fruit hanging at that stage. A reasonable planning assumption once you're over the floor: a lift in the region of 10% in orders within about six months. Below the floor, that maths doesn't clear.

Deliverable: A documented go/no-go decision.