NewThe detectors that scored perfect collapsed the hardest under attack.

Off-the-shelf attack data, pulled through an API.

Labeled real and AI-generated imagery, broken out by group and by platform condition.

A map of where a detector fails, not a single score.

Per-demographic breakdown

Results per cell, not one pooled average that hides the group you fail on.

Platform conditions

Clean imagery plus re-encode variants for Facebook, Instagram, TikTok, and X.

Real and fake, paired

Every synthetic item traces back to a real source, for like-for-like scoring.

Commercially licensed

Cleared for building and evaluating detection under the Margen data license.

Three places to start.

Data cards

What is in each benchmark: composition, cells, and conditions.

See the data cards

Docs

Auth, endpoints, filters, pagination, and the Python SDK.

Read the docs

Portal

A free sample on signup. Credits unlock the full catalog.

Get an API key

Teams that build or evaluate deepfake detection.

Find where your detector fails before a customer does.

Score each release against labeled real and synthetic media, and catch the regression before it ships.

Regression trackingPre-release QAModel cards

Licensed for commercial use, including building detection.

Real faces

Licensed from commercial stock-media providers.

Synthetic faces

Generated in house.

Your use

Each delivered item is licensed for your evaluation use under the Margen data license.

Start pulling.

Sign up, generate a key, and run the free sample through your own pipeline.