Passport and ID
Face-Injection
Does your detector still catch an AI-generated face after it has travelled a real submission path, phone camera, print, scan, photocopy, rather than arriving as a clean studio render?

Clean-render accuracy is not field accuracy.
Detectors that score well on clean synthetic faces often degrade sharply once an image has been through a phone camera or a document reproduction step. This benchmark isolates that gap.
Every item is labeled on a two-part axis. The full benchmark targets balanced counts across all twelve cells; this sample spans a representative subset.
Twelve cells, read as a grid.
Skin tone uses a 6-band ordinal scale reconciled against Individual Typology Angle, so cells stay comparable across items. Swatches are indicative of the band, not a measured value.
Each condition emulates a real submission path.
Applied on top of the generated face. The set is extensible; additional capture and document chains are added as the build progresses.
How the corpus is built and labeled.
Every item names its attack with a TTP tag.
A pass or fail count does not tell you what to fix. Each item carries a Tactic, Technique, and Procedure tag: the tactic is impersonation, the technique is an identity-preserving edit, and the procedure is the specific editor. Group misses by any of the three to target a fix.
What it is for today, and what it is not yet.
- Detector evaluation and robustness testing
- Locating where a detector breaks, by cell and by condition
- Training and fine-tuning as the corpus grows, the same set serves the red team measuring a gap and the detection team closing it
- This is a sample release from an in-progress build
- Per-cell counts and condition coverage expand in the full benchmark
- Demographic labels are perceptual consensus review, not biometric ground truth
Synthetic attack faces, as they arrive.
Query the corpus the way you would query a database.
Every dimension is a filter. Compose them, read a live .total, then pull only what you need.
GET /v1/items ?benchmark=passport-pad-v1 &kind=fake &skin_tone=brown &gender=male &generator=qwen_image_edit
Pull it and score your detector.
Access is via the Margen platform. Generate an API key, then follow the API docs to pull the catalog and measure per cell and per condition.
Currently held out. This benchmark is being run through an active evaluation, so only a test set of 10% of its identities is available to pull. The remaining held-out set is measured through a managed evaluation, not delivered as raw data, so nobody scored against it has trained on it.


















