Conversation
The 18 v23 holdouts that had no usable upstream repo and were mirrored to
34data (bmcore PR #359). Uploaded and verified: 18/18 repos hold data,
56.92 GB across 326k source files.
real_images +7 medical imaging (kidney-stone-ct, malaria, pulmonary
chest x-ray, dermnet, coronahack, lits-256, tcga-coad)
synthetic_images +3 synthetic-human, ai-generated-ecommerce x2
synthetic_videos +7 VideoGen-RewardBench x5 (easyanimate/opensora1.2/
qingying/tongyi/vidu) + edtalk, float-talking-head
synthetic_audio +1 indic-tts-urdu
Archive bounds are per-run download limits, not sample counts — the per-dataset
sample cap comes from calculate_weighted_dataset_sampling. In full mode the YAML
values are respected (DOWNLOAD_SIZE_OVERRIDES is debug/small only), so -1/-1
would fetch each dataset in full every run to use a few hundred samples, and
would also lose variety: _select_files_to_download draws a seeded random subset
when the count is bounded, but returns everything when it is -1. Bounded values
give per-run economy and rotate archives across rounds as the seed changes.
image 1000/1 -> 1000 available vs 355 cap
video 500/1 -> 500 available vs 149 cap
audio 500/2 -> 1000 available vs 336 cap
For these 18 that is 56.9 GB -> 20.6 GB per full run (64% less), concentrated in
synthetic-human (16.6 -> 2.4), float-talking-head (13.1 -> 4.4), tcga (5.4 ->
1.8) and edtalk (5.3 -> 1.8). Video entries declare source_format: mp4 but hold
zips; the loader's fallback recalculates n_files on the archive branch, so
archives_per_dataset governs.
Raise full-mode targets so 18 new datasets do not dilute the existing ones:
image 69000->72500, video 33000->34000, audio 47000->47500. Per-dataset caps are
held at 355 / 149 / 336 (audio +1), unchanged from before this wave.
edtalk and float-talking-head keep media_type: semisynthetic, joining the 29
existing semisynthetic video entries.
The other 33 v23 entries are not here: 22 defer to existing ungated upstream
repos and will land separately with hf_revision pinned, 7 have gated upstreams
and still need mirroring, and 4 wild-* stay withheld.
Registry loads clean at 565 datasets, no duplicate names or paths, and all 18
paths resolve on HuggingFace with data.
data: add released v23 holdout datasets to public registry (18 mirrors)
Second half of the v23 release. These 22 holdouts already exist in public upstream datasets, so instead of copying Wasabi -> 34data they point at the upstream repo directly. Same approach already used for cml-tts-german/spanish, ai4bharat-speecharenabench-kn/mr and indic-tts-kannada/malayalam. ylacombe/cml-tts 4 dutch, italian, polish, portuguese Sumsub/Swappir 5 GPEN x2, roop x2, simswap scaledf/ScaleDF 4 FFpp Deepfakes/Face2Face, E4E, DiffusionCLIP RekaAI/RekaDaily-10k-raw 3 phone shards 00000-00002 obscure-entropy/conceptual_captions_hu_filtered 1 eltorio/ROCOv2-radiology 1 rshaojimmy/DGM4 1 manipulation/StyleCLIP.zip MCG-NJU/TimeLens2-93K 1 videos-00003-of-00018.tar mvp-lab/LLaVA-OneVision-2-Data 1 mid_training_video/30s WenhaoWang/TIP-I2V 1 i2vgenxl_videos_subset_1.tar Upstream filenames match the Wasabi path segments exactly, so these resolve to the same material that was held out. Every repo is ungated and every entry pins hf_revision, so an upstream force-push cannot silently change what the benchmark reads. include_paths is a substring match, which is a trap here: 'CelebA_HQ_roop.zip' also matches 'GPEN_CelebA_HQ_roop.zip', and 'fairface_roop.zip' also matches 'GPEN_fairface_roop.zip'. Both carry exclude_paths: ['GPEN_'] so they resolve to one file each. Every include_paths was checked against the pinned revision's actual file listing: 20 resolve to exactly 1 archive, rocov2 to its 27 train shards, cml-tts to 373/61/12/42 per language. ScaleDF is pinned to scaledf/ScaleDF (the org repo) rather than WenhaoWang/ScaleDF; both carry all four tars. Archive bounds follow the same reasoning as the mirrored half — image 1000/1, video 500/1, audio 500/2, all clearing the per-dataset caps, which land at 335 / 145 / 327 with these 22 added. Registry loads clean at 587 datasets, no validation failures, no duplicate names and no duplicate (path, include_paths) pairs introduced.
…owup data: register 22 v23 holdouts against upstream repos (no 34data mirror)
Adds ai4bharat-speecharenabench-gu/ml/or/ur, pointing at ai4bharat/SpeechArenaBench with the same hf_revision, hf_subfolders and data_columns shape as the existing kn/mr entries. No upload required. Completes the v23 audio release: 9 of 9 holdouts registered, 4 real / 5 synthetic. Registry loads at 591 datasets, no validation failures, no duplicate names.
data: register the 4 v23 SpeechArenaBench holdouts
…-column Set Conceptual Captions image column
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Promotes
devtomain: the v23 holdout dataset registrations (#146, #147, #149).34data/*hf_revisionpinnedReleased holdout mapping
The v23 bmcore holdout datasets released in this version, mapped to the deterministic obfuscated holdout directory names they used on the benchmark volumes (computed from the v23 holdout configs via the same fingerprint logic gasbench uses for holdout loading).
conceptual-captions-0-1mreal-image-holdout-ed700208imagerealcoronahack-chest-xrayreal-image-holdout-e66cad4dimagerealdermnetreal-image-holdout-b92141ceimagerealkidney-stone-ctreal-image-holdout-3933fc95imagereallits-256real-image-holdout-0db5eb2dimagerealmalaria-bounding-boxesreal-image-holdout-67cde42eimagerealpulmonary-chest-xray-abnormalitiesreal-image-holdout-bc18aa99imagerealrocov2-radiologyreal-image-holdout-f2fc06ceimagerealtcga-coad-msi-mssreal-image-holdout-946aa138imagerealdgm4-styleclipsemisynthetic-image-holdout-1bf9b5c9imagesemisyntheticscaledf-diffusionclipsemisynthetic-image-holdout-9ba5d572imagesemisyntheticscaledf-e4esemisynthetic-image-holdout-a7ed6b28imagesemisyntheticscaledf-ffpp-deepfakessemisynthetic-image-holdout-a8436881imagesemisyntheticscaledf-ffpp-face2facesemisynthetic-image-holdout-a09c6e01imagesemisyntheticswappir-celeba-hq-roopsemisynthetic-image-holdout-9cfc90feimagesemisyntheticswappir-celeba-hq-simswapsemisynthetic-image-holdout-f2148bf3imagesemisyntheticswappir-fairface-roopsemisynthetic-image-holdout-7e6292ceimagesemisyntheticai-generated-ecommerce-damaged-productsynthetic-image-holdout-8ae74aaaimagesyntheticai-generated-ecommerce-fake-logisticssynthetic-image-holdout-31bcbe43imagesyntheticswappir-gpen-celeba-hqsynthetic-image-holdout-dae9b633imagesyntheticswappir-gpen-fairfacesynthetic-image-holdout-962de462imagesyntheticsynthetic-humansynthetic-image-holdout-1fd16e77imagesyntheticllava-onevision-2-mid-training-00000real-video-holdout-5f78ecbdvideorealrekadaily-phone-shard-00000real-video-holdout-845771ffvideorealrekadaily-phone-shard-00001real-video-holdout-bf6137f9videorealrekadaily-phone-shard-00002real-video-holdout-5828d24avideorealtimelens2-93k-00003real-video-holdout-6ba084acvideorealedtalksemisynthetic-video-holdout-e931a355videosemisyntheticfloat-talking-headsemisynthetic-video-holdout-d9dc4ee1videosemisynthetictip-i2v-i2vgenxlsynthetic-video-holdout-eca6dd07videosyntheticvideogen-rewardbench-easyanimatev4synthetic-video-holdout-a984ce1fvideosyntheticvideogen-rewardbench-opensora1-2synthetic-video-holdout-fc6a3601videosyntheticvideogen-rewardbench-qingyingsynthetic-video-holdout-bf19d2d8videosyntheticvideogen-rewardbench-tongyisynthetic-video-holdout-832fd9afvideosyntheticvideogen-rewardbench-vidusynthetic-video-holdout-9bd9f4acvideosyntheticcml-tts-dutchreal-audio-holdout-1d84e222audiorealcml-tts-italianreal-audio-holdout-59302feaaudiorealcml-tts-polishreal-audio-holdout-e453319baudiorealcml-tts-portuguesereal-audio-holdout-f6d2434caudiorealai4bharat-speecharenabench-gusynthetic-audio-holdout-c0a1dad1audiosyntheticai4bharat-speecharenabench-mlsynthetic-audio-holdout-6cf0b951audiosyntheticai4bharat-speecharenabench-orsynthetic-audio-holdout-4ca4e29aaudiosyntheticai4bharat-speecharenabench-ursynthetic-audio-holdout-6b072d50audiosyntheticindic-tts-urdusynthetic-audio-holdout-4ce4357caudiosynthetic