A jewellery listing photographed on Amazon’s mandatory white background answers exactly one question: what does this look like. It does not answer the question that actually stops a purchase, which is what does this look like on me. That second answer comes from a lifestyle image — the same ring on a hand, the same pendant on a neck — and its absence is one of the more common, least visible gaps in a large catalogue.
The problem is finding out which listings have that gap. Not a handful you happen to notice while browsing the catalogue, all of them, across every ASIN and every image in every gallery.
Why this doesn’t get checked by eye
At a small catalogue, someone opens each listing and looks. At 1,911 ASINs, each with its own set of gallery images, that stops being a task and becomes a project nobody has time to run — so it either doesn’t happen, or it happens as a spot check that catches a handful of egregious cases and misses the rest.
A spot check is worse than it sounds, because the gap is not evenly distributed. Some product lines were shot with a full lifestyle set from day one; others were added later, by a different photographer, on a tighter budget, with white-background-only coverage. Sampling ten listings tells you almost nothing about the other 1,901.
The only way to actually know the shape of the problem is to check every image on every listing. That means fetching the full gallery for every ASIN and looking at each photo individually — which is a scale of image inspection nobody does by hand, so it has to be automated or it doesn’t happen.
Building the check: fetch, then classify
The audit runs in two stages, because each one fails differently and needs to be debugged on its own.
Stage one pulls the gallery. For each ASIN, fetch its product page and extract the image set from the page’s own embedded data, not from scraping thumbnails off the rendered page. This matters more than it sounds: Amazon’s product pages embed the image sets for every sibling variant of a listing on the same page, and grabbing everything that looks like a product image pulls in photos that belong to a different ASIN entirely — one URL turned up shared across 43 siblings in this catalogue. The fix is reading the one JSON block scoped to the ASIN actually requested, not every image block the page happens to contain.
Stage two classifies each image. A local YOLO11 person detector checks every photo in every gallery for a human body part — hand, wrist, neck, ear, torso — because the working rule for “lifestyle” here is simple: if any image in the gallery shows a person wearing or holding the product, the listing has lifestyle coverage. If none do, it’s flagged.
Where the naive version breaks
The detector alone is not enough, and this is the part worth explaining because it’s the part that would quietly wreck the audit if skipped.
Run the detector with a normal confidence threshold and check the results against a hand-labelled sample, and the threshold does nothing useful. On 116 manually checked images — 79 real model shots, 37 false positives — the median detection confidence was 0.248 for the real ones and 0.165 for the false ones. Those distributions overlap almost completely. A threshold strict enough to exclude the false positives excluded every real model shot too. A threshold loose enough to catch real model shots let in all 37 false positives.
The false positives had a pattern, though, once we looked at what they actually were: the detector latching onto a curved bracelet edge or the outline of a charm sitting on Amazon’s plain white product background, and reporting it as a small, low-confidence person. Real model shots looked geometrically different in two specific ways — a wrist or neck genuinely fills most of the frame, and there’s no white border around it because it’s a photograph of a person, not a product on a seamless background.
Those two measurements — how much of the frame the detected region fills, and how white the border of the image is — turned out to separate the categories cleanly where confidence alone could not. Scored against the same hand-labelled set, the combined rule got 96.6% right. Cases that fell near either boundary were routed to manual review instead of forced into a yes/no bucket, so the ambiguous minority gets a person’s judgment and the clear majority doesn’t need one.
What it found
Across 1,911 ASINs, 766 had at least one lifestyle image somewhere in their gallery. 1,145 — roughly 60% — did not. That’s not a rounding error in a catalogue; it’s the majority of the listings having no answer to the “what does this look like on me” question anywhere in their gallery.
Before any of that number went into a spreadsheet or a client conversation, it was checked against reality — stratified sample sheets pulled from the flagged listings, reviewed by eye, weighted toward the failure mode that costs the most: a listing wrongly told it’s fine when it isn’t. A detector that scores well on paper and hasn’t been checked against a sample of its actual output is a detector nobody should trust with a catalogue-wide claim.
The same pipeline now has a second job
A few weeks after this audit ran, the question it answers got a second, sharper edge. Since June 2026, Amazon requires sellers to tag any listing image containing a fully AI-generated, photorealistic person with the keyword contains-synthetic-performer in the file’s XMP metadata — a response to New York’s synthetic-performer disclosure law. Miss the tag on a main image and Amazon can flag it; a listing without a compliant main image can get suppressed from search entirely, not just penalised.
That rule is not about lifestyle images specifically — an AI-generated model counts whether or not the shot is a genuine lifestyle photo. But it lands on the same gallery this audit already fetches, which is the point worth noticing: stage one of this pipeline, pulling the correct scoped image set for every ASIN, is identical work whether stage two is asking “is there a person here” or “is the person here real.” A catalogue that has already been fetched and classified once is most of the way to being checked for the second thing too, because the expensive part — getting the right images, per ASIN, without pulling in sibling variants — is already done.
For a jewellery catalogue in particular, this is worth checking rather than assuming away. AI-generated lifestyle shots are a cheap way to backfill the exact gap this audit found — a wrist or ear wearing the product, without booking a real model — and that shortcut is precisely the kind of image the new rule targets.
The point of doing it this way
None of this is really about YOLO or bounding boxes. It’s about the difference between “we think some listings are missing lifestyle images” and “these specific 1,145 ASINs are missing lifestyle images, and here’s the exact image gallery each one has, so a photographer or a listing team can go fix it directly.” The first is a guess. The second is a work list, and turning a catalogue-wide gap into a work list is what makes it actionable rather than a vague ongoing worry — the same image and infographic direction that a listing needs to actually convert once someone finds it.