Looking for evidence of claim and refund fraud, we read 2,400 posts from the two largest open image-editing communities. We found no fraud at all. What we found was the same techniques, discussed as ordinary craft, by people with no idea what they are being used for. The capability and the abuse sit in separate rooms, and nothing runs between them.
What the corpus showed
The pull was straightforward. Two communities where open image-editing models are built and discussed, 2,400 posts, filtered for the vocabulary of damage editing: adding, removing, extending, and repairing damage in photographs of objects.
Not one post discussed faking a claim or a refund. The words "claim", "insurance" and "refund" appear only in their ordinary English senses. What the corpus is full of instead is legitimate craft: object removal, inpainting, restoration, upscaling, and removing furniture from real-estate photographs so a room shows better.
A null result is easy to dismiss and this one should not be. We went looking for the abuse where the capability is built, and the abuse was not there. That is the finding.
Two rooms, one technique
Run the two literatures side by side and the overlap is exact. In the builder room, the task is removing a sofa from a listing photo so the floor behind it is continuous. In the fraud room, the task is adding a dent to a door panel so the metal behind it is continuous. Same model, same mask, same interpolation problem, opposite intent.
That symmetry is why the split matters. A detection team reading the builder room would learn exactly which operations are cheap this quarter and where they still break. They do not read it, because nothing in a fraud team's day points there. A builder publishing a mask-coverage tip would learn that the technique is landing in claim files. They never hear it, because fraud losses are not published.
What arrives at an insurer or a marketplace is the same shape we documented in the claim photos insurers are already catching and in the refund photos that do not fray right: a real photograph with one region altered, not a wholly generated picture. That is an inpainting job. It is the builder room's bread and butter.
The same failure, found twice
The clearest evidence that these rooms are solving one problem is that they keep finding the same limits independently.
In our own generation work we found a threshold in mask coverage. Below roughly a third of the image, the fill interpolates correctly and structure carries across the masked region: plank joints line up, edges continue, the geometry holds. Past about three quarters, the model stops interpolating and starts inventing, producing structure that is internally plausible and inconsistent with the rest of the frame.
A builder reported the identical threshold while removing furniture from real-estate photographs. Same model family, same coverage behaviour, same point where interpolation turns into invention. Neither party knew the other was looking.
That threshold is a detection opportunity and nobody is holding both halves of it. The builder knows where the model stops interpolating. The fraud team knows which claims cleared. Put the two together and you get a test. Leave them apart and you get a folk tip on one side and an unexplained loss on the other.
Why no loop forms on its own
It would be convenient if this gap closed under market pressure. It does not, for three reasons that all point the same way.
Builders have no reason to look. The work is legitimate, the licences are permissive, and a technique that removes a sofa is not improved by knowing it also removes a scratch from a claim photo. Publishing is the norm and there is no feedback channel pointing back from misuse.
Fraud teams cannot publish what they learn. Detected fraud is commercially sensitive, and the detail that would be most useful to a defender, the exact recipe that cleared a check, is the detail most dangerous to release. Insurers publish totals, not methods.
And no market force pays for the connection. Detection vendors sell against a threat model they assemble themselves. Buyers have no independent way to check whether that model matches what is arriving, which is the same procurement problem that makes a headline accuracy number worth so little. Nobody is paid to stand between the room where the capability is made and the room where it lands.
Where that leaves the measurement
A threat model assembled from caught fraud describes the fraud the existing checks were already capable of catching. That is the structural consequence of the split: the techniques nobody has seen in a claim file are not rare, they are published openly in a room that produces no incident data.
The recurring rule applies here in an unfamiliar direction. The question is never whether an edited photograph looks real, it is whether the check reads a property the editor cannot supply, and which properties those are changes with whatever the capability room shipped this quarter.
That is the gap our own work sits in. We read the room where the capability is published, reproduce what is landing there, and measure it against detection before it arrives as a loss. The corpus behind the finding in this article came from exactly that process.
Margen does not sell a detector or a claims product. We read the room where the capability is published, build the attacks it makes cheap, and measure whether your checks hold against them.
Related reading
- Methodology notesA detector score is meaningless until you name the condition.The same two detectors score 1.000, about 0.70, and 0.243 on the same images, depending only on how the files were prepared. A score without its condition is not a weak claim, it is an incomplete one, and the condition that broke these models was not the one anybody expected.
- Fraud storiesWhen a 15-dollar AI ID passed a live exchange KYC check.The OnlyFake case is the cleanest documented example of an automated control being beaten by generative AI: a synthetic passport image cleared a crypto exchange's document verification.
- Fraud storiesTwo deepfake CEOs, two executives who did not fall for it.A cloned voice of Ferrari's CEO and a deepfake of WPP's CEO both failed, and neither was stopped by a detector. Each was stopped by a person verifying something the impostor could not supply.