Why launch QA is different from a catalog audit
An audit and a launch gate answer different questions. An audit of a live catalog asks which existing images are costing you the most money, and ranks fixes by revenue. A launch gate asks a binary question: is this batch fit to publish, yes or no? There is no sales data yet to prioritise with, so every SKU in the batch carries equal weight until it proves it is fine.
Launch QA also has a hard deadline attached. That changes the economics. A defect found in QA costs one re-edit. The same defect found after launch costs a re-edit, a re-upload, a cache purge, a marketplace feed resubmission and, for anything that went into ads or email, new creative. The cheapest moment to fix an image is before anyone outside the team has seen it.
Catalog audit
- Runs on a live catalog
- Ranks fixes by revenue at risk
- Fixes trickle out over weeks
- Goal: find the expensive gaps
Launch QA gate
- Runs on a batch before it is published
- Every SKU weighted equally
- Pass/fail decision by a fixed date
- Goal: nothing embarrassing goes live
The gate is only as good as the standard it checks against. If you do not have a written product photography style guide with framing, background and colour rules, write the short version before QA starts. Reviewers who are checking against their own taste will disagree, and disagreement on launch week is where deadlines slip.
Per-image checks: the ten things every file must pass
These are the checks a reviewer runs on each file individually. Most of them are mechanical, which means most of them can be scripted. Do that wherever you can and save human attention for the checks that need eyes.
| Check | Pass condition | Automate? |
|---|---|---|
| Resolution | Long edge at or above your zoom target (2,000 px is a common floor) | Yes |
| Aspect ratio | Matches the standard exactly (e.g. 1:1 or 4:5), no off-by-one pixels | Yes |
| File format and size | Agreed format, under your upload ceiling, sRGB embedded | Yes |
| Filename | Follows the SKU naming pattern, no "final_v3" or camera numbers | Yes |
| Background | Plate colour within tolerance at the corners and edges | Yes |
| Product present and whole | Nothing cropped off unintentionally (toes, straps, handles) | Partly |
| Cutout edge | No halo, no fringe, no missing fine detail (hair, fringe, mesh) | Eyes |
| Retouching | No dust, lint, stray threads, sensor spots, or visible clone repeats | Eyes |
| Colour accuracy | Matches the physical sample under daylight, side by side | Eyes |
| Correct product | Image matches the SKU and variant it is attached to | Eyes |
The last row causes more customer complaints than the other nine combined. A navy sweater photographed and filed under the black variant passes every technical check and still generates returns. Check SKU-to-image mapping against the product sheet, not against the filename, because the filename is usually where the mistake was made.
Colour accuracy needs the physical sample in the room. Reviewing colour on a monitor against your memory of the product is how a burgundy becomes a brick red across an entire collection. If the samples have already shipped back, flag colour as unverified rather than passing it.
Set-level checks: what only shows up in the grid
This is the part most launch checklists skip, and it is the part shoppers notice first. Lay the whole batch out as a contact sheet at roughly the size it will appear on a collection page, sorted the way the storefront will sort it. Then look for things that are wrong only in relation to the neighbours.
- Frame fill. Products in the same category should occupy a similar share of the frame. A pair of trainers that fills 85% of its square next to one that fills 60% reads as two different-sized shoes. Set a tolerance per category, for example product height within plus or minus 5% of the category target, and measure it rather than eyeballing it.
- Placement. Same baseline for products that stand on the floor, same centre line for products that float. Inconsistent vertical placement makes a grid look like it is bouncing.
- Background match. Two whites that each look white alone can look grey and cream side by side. Sample the plate colour from every image and compare the values, not your impression.
- Shadow treatment. One style per category: none, a soft contact shadow, or a reflection. Mixing them within a collection is one of the fastest ways to make a catalog look assembled from different shoots.
- Angle and orientation. If the standard is three-quarter left, every hero should face the same way. A single right-facing bag in a row of left-facing ones is the outlier everyone notices.
- Colour family consistency. Products sold as the same colour across categories (a "sand" chino and a "sand" jacket) should actually look the same colour in the grid.
Frame fill and placement are where batches drift most, because they depend on how each image was cropped, by whom, and on which day. This is also where a deterministic step pays off. In Retouchable, product-only outputs finish with a framing step that measures the product, scales it by a per-category rule and places it in the same safe area on the exact plate colour, so the grid is consistent by construction instead of by review. Whatever tool you use, the principle holds: framing that is calculated is easier to QA than framing that is judged. For more on the background side, see keeping product backgrounds consistent across a catalog.
Extra checks for AI-generated and AI-edited images
AI tools remove most of the per-image labour, which is exactly why a launch batch is now often hundreds of images produced in hours. They also introduce failure modes that a traditional retouching QA pass was never designed to catch, because a human retoucher does not invent details. Add these checks to any batch where generative tools touched the pixels.
| Failure mode | What it looks like | How to catch it |
|---|---|---|
| Text and logo drift | Label lettering subtly rewritten, logo proportions changed | Zoom to 100% on every label, tag and print; compare to source |
| Construction changes | Extra button, missing pocket, different stitch pattern | Side-by-side with the source photo, not just the sample |
| Colour shift | Generated output warmer or more saturated than the original | Measure product colour against the source, not against the background |
| Texture smoothing | Knit, denim or leather grain flattened into a plastic look | Check at 100% zoom; texture loss hides at thumbnail size |
| Anatomy on models | Hands, fingers, ears or jewellery that do not quite work | Dedicated pass on every on-model image, hands first |
| Identity drift | The same model looks like a different person across SKUs | Review on-model images per model, in sequence |
The common thread is that AI defects are plausible. They do not look like errors at thumbnail size, which is why they survive casual review. The fix is to always compare output to source, never just to judge output on its own. Our guide to evaluating AI product photography quality goes deeper on scoring generated images.
Text and logo drift is the AI defect most likely to become a legal or brand problem rather than a cosmetic one. A misspelled brand name on a woven label is a defect even if nobody notices it for a month. Treat any lettering change as critical.
Grade defects by severity so the launch is not held hostage
Without a severity scale, QA turns into a negotiation. One reviewer holds the launch over a slightly soft shadow; another waves through a wrong-variant image because it looked fine. Agree three levels before review starts and attach a rule to each.
Critical means the image misrepresents the product or the brand: wrong product or variant, wrong colour, altered logo or label text, visible construction changes, a missing hero image, or anything that breaks a marketplace's listing rules. A critical defect takes the SKU out of the launch until fixed. It does not take the whole launch down.
Major means the image is accurate but visibly off-standard: frame fill outside tolerance, wrong background tone, a cutout halo visible at product-page size, a shadow style that does not match the category. These are fixed before launch if time allows, and on a committed date if not.
Minor means something only a trained eye catches at 100% zoom: a faint stray thread on a secondary angle, a slightly soft edge on a detail shot. Log it and batch it into the next maintenance pass.
The rule that keeps this honest: severity is decided by what the defect would cost a shopper, not by how much it annoys the reviewer. A wrong variant image is critical because it causes returns. A slightly off-centre detail shot is minor because nobody buys or returns on it.
How much to check: sampling versus 100% review
Reviewing every image of a 2,000-image launch at 100% zoom is not realistic. Reviewing ten and hoping is not QA. The workable middle is to run automated checks on everything, a set-level grid review on everything, and a human close inspection that scales with risk.
These are working starting points, not industry benchmarks; adjust them to your defect history. Two rules make sampling trustworthy. First, every hero image (position one, the image that appears in collection grids, search results and ads) gets a full close review regardless of batch size, because it carries most of the traffic. Second, sampling escalates: if the sample turns up a critical defect, the rate for that batch goes to 100% until you know whether the defect is a one-off or a pattern. A label rewrite on one AI flat lay is a reason to check every flat lay from that run.
Stratify the sample, too. Pull from every category, every editor or pipeline, and every day of production, rather than the first 50 files in the folder. Defects cluster by source, and a random-from-the-top sample can miss an entire bad batch.
Sign-off, the final pre-publish pass, and the checklist
Name one person who owns the go/no-go decision. That person does not review every image; they confirm that every check ran, every critical defect is resolved or its SKU is pulled, and every major defect has an owner and a date. Shared ownership of launch sign-off usually means nobody signs off until the morning of launch.
Then do one last pass after upload and before publish. Images that passed QA as files can still fail in place: the platform recompresses them, the theme crops them to a different ratio, the wrong image lands in position one, or alt text did not carry over. View the products as drafts or on a staging storefront at desktop and phone widths. On Shopify specifically, image position decides the featured image, so a correct file in the wrong slot is a launch defect; our Shopify launch imagery checklist covers that platform's steps in detail.
The condensed gate, ready to paste into a launch template:
- Written standard exists: aspect ratio, frame fill per category, plate colour, shadow style, angle.
- Automated checks run on 100% of files: resolution, ratio, format, colour profile, filename, plate colour.
- SKU-to-image mapping verified against the product sheet, including variants.
- Contact-sheet grid review of the whole batch, sorted as the storefront sorts it.
- Frame fill and placement measured against category tolerance.
- Colour checked against physical samples, or flagged unverified.
- Every AI-touched image compared to its source for text, construction, colour and texture.
- Every hero image close-reviewed at 100% zoom.
- Sample rate applied per batch risk; escalated to 100% on any critical find.
- Defects graded critical, major or minor; critical SKUs pulled or fixed.
- Post-upload check on staging at desktop and phone widths, including featured image and alt text.
- One named owner signs the go/no-go.
Keep the defect log after launch. The defects you found this time are your best predictor of the ones you will find next time, and they tell you which checks to automate before the next batch arrives.