NewThe detectors that scored perfect collapsed the hardest under attack.

Questions buyers ask first.

Still unsure whether an independent evaluation is worth it? Start here, then talk to us.

Anything we missed: info@margensoftware.com

Vendor numbers are usually measured on clean, familiar data. In our own benchmark, two detectors that scored a perfect AUC of 1.000 fell to 0.243 and 0.340, below random, once real and synthetic images were re-encoded through one pipeline. They had been reading the file format rather than the image. The other six lost 0.047 on average, so the collapse is not universal. The question is not whether the number is high, it is what it was measured on.

Internal teams test what they know how to build and inherit the same blind spots as the system under test. We bring a pre-registered taxonomy, balanced demographic coverage, and attacker-realistic budgets. Findings are exhibits, not anecdotes.

It starts with a conversation about what you are trying to assess. We learn the problem, work out which metrics actually matter for it, red-team your detector against them, and hand back a report on where it holds and where it fails. On request we then work with your detector vendor, or with your own team, to close the gaps.

No. Your model and any data you share stay inside the engagement. Consent rules restrict processing biometric data without permission, so a neutral third party that already holds the attacks is often the most practical way to get an honest result.

Yes, on an annual retainer. We retest quarterly against new attacker releases and track where each detector loses ground, by attack class and demographic group.

Structurally. A vendor cannot pay to raise its score or its position, the methodology is fixed and published before an engagement starts, and the buyer owns the report. We sell no detector of our own, so a best-fit recommendation is driven by the measured results. If we ever earn a referral or share revenue with a vendor, we disclose it, and it never moves a number.

Detection vendors who need independent validation, identity-verification and KYC providers, hiring and interview platforms, insurers and marketplaces checking claim and return imagery, red-team firms, and enterprise security leaders. In short, anyone whose fraud defense leans on a detector.

It is test material for detectors. The API delivers labeled real and AI-generated imagery so a team can score their own detection system against it. Real faces are licensed from commercial stock-media providers, synthetic faces are generated in house, and each delivered item is licensed for your evaluation use under the Margen data license.

Still weighing it up?

Tell us what you are protecting and which detection system is protecting it. We will tell you whether an evaluation is worth your time.