Batch export
Download a whole benchmark as a few large parquet files instead of one request per image. passport-pad-v1 and synthetic-face-v1 are available now.
GET /api/v1/data/release?benchmark=<id>
Call it with a benchmark. It returns signed download URLs for every shard set your key can access, and the credits the pull will cost.
1!pip install margen2from margen import Margen3from margen.ergonomics import pull_release45client = Margen(bearer_auth="mgn_...") # your key from the portal67# Download a whole benchmark into one DataFrame.8# Change the id to "synthetic-face-v1" for the other dataset.9df = pull_release(client, benchmark="passport-pad-v1")10print(len(df), "rows")1112# Decode an image. The image column is a {bytes, path} struct.13import io14from PIL import Image15img = Image.open(io.BytesIO(df.iloc[0]["image"]["bytes"]))
Ways to pull
Batch export gives you a whole benchmark. To pull a specific slice instead, use the per-image helpers: they page through your filter and download each image, charging one credit each (net of what you already own).
1from margen.ergonomics import pull_release23df = pull_release(client, benchmark="passport-pad-v1")
Versions and filters
A benchmark can have more than one release. Without release you get its default, which is a deliberate choice on our side and not necessarily the newest; every response lists all of them under available_releases, flagging the default and naming the generators each one contains. Name a release to pull that version, and add generator, kind or condition (also accepted as perturbation, the per-image name) to take only some of its shard sets. Filters only narrow what you are already entitled to, and you pay only for the sets served.
generator= keeps the real (bona fide) sets alongside the generator you named, so the pull can still score a detector. Add kind=fake for the attacks alone. That is how an existing buyer adds a newly released generator without re-downloading anything they already own: only the new sets are transferred, and only they are charged.
1from margen.ergonomics import pull_release23# A specific version of the benchmark.4df = pull_release(client, benchmark="passport-pad-v1", release="passport-pad-v1.1")56# Just one generator's attacks from it. If you already own the rest of the7# benchmark, this transfers and charges only the new sets.8df = pull_release(9client, benchmark="passport-pad-v1", release="passport-pad-v1.1",10generator="flux2_klein_4b", kind="fake",11)
A release that does not exist is a 404 unknown_release naming the valid ids. A filter that matches nothing is a 200 with no shards and available_filter_values showing what would match.
Shard sets
A benchmark is split into shard sets, one per condition, and on releases with more than one generator one per condition and generator (a set key like chain_fax__flux2_klein_4b). Each set holds every image under that condition, and the sets share no images, so you can buy any subset of them without overlap. A large set is split into numbered parts under 1 GB; load all of a set's parts, or none.
Set sizes are not uniform. On synthetic-face-v1, clean is the full 27,973-image corpus while each robustness set covers a 2,399-image subset, and clean_baseline re-ships that subset so you can take it alongside the perturbations. Read each set's rows before you buy.
Credits
One credit per image, the same as pulling images one at a time. credits_required is the price for this pull and is shown before anything is charged.
You are never charged twice for the same image. Buying a set means you own its images, so re-pulling it, or downloading any of those images individually, is free. If a release includes images you already own, only the new ones are charged.
Coverage
The coverage field tells you what you got:
full— the whole benchmark.scoped— the sets your key is entitled to.partial— your access does not line up with whole sets (for example a per-cell restriction). No shards are returned, but you can still pull what you are entitled to per image. This is a200, not an error.empty— nothing here matches your access.
On a partial or scoped response, pull your slice per image with the same filter your access allows:
1from margen.ergonomics import iter_items, download_selection23# your entitled slice, e.g. one demographic cell4items = iter_items(client, benchmark="synthetic-face-v1", skin_tone="dark", gender="female")5download_selection(client, items, out_dir="out/")
You are never charged twice, so a per-image pull of images already covered by a set you own is free.
Signed URLs
Part URLs expire after six hours (expires_in, in seconds) and support HTTP Range, so an interrupted download can resume. If a URL expires, call /release again; re-calling a release you own is free and returns fresh URLs.
Loading and integrity
Each part is a parquet file that loads with pandas.read_parquet or datasets.load_dataset. The image bytes are in an image column as a { bytes, path } struct; media_id joins the same base image across conditions. Every part carries a sha256; check it after download to catch a truncated transfer.