-
Notifications
You must be signed in to change notification settings - Fork 0
MedMCQA standard-battery replication (tracking) #291
Copy link
Copy link
Closed
Labels
dataset:medmcqadifficulty: intermediateTouches one subsystem; some context neededTouches one subsystem; some context neededexperimentExperiment runner / study designExperiment runner / study designpriority: highDo this soon; unblocks the paper or other workDo this soon; unblocks the paper or other work
Description
Activity
Metadata
Metadata
Assignees
Labels
dataset:medmcqadifficulty: intermediateTouches one subsystem; some context neededTouches one subsystem; some context neededexperimentExperiment runner / study designExperiment runner / study designpriority: highDo this soon; unblocks the paper or other workDo this soon; unblocks the paper or other work
Tracking issue for the MedMCQA standard-battery replication. Each item mirrors an experiment from the finished MedQA/NIH template so results are apples-to-apples across datasets (see the workflow: replicate the standard battery on a new dataset).
Dataset adapter is already implemented (
benchmaxxing/datasets/medmcqa.py, #112). These are replication, not new features: run the existing experiments on a MedMCQA manifest through the Gemini API, and write results underexperiments/medmcqa/results/in the same JSON/JSONL format asexperiments/medqa/results/.Foundational (run first)
Seed shape
Social pressure / committee structure
Hierarchy / orchestrator
Controls & robustness
Break-it arms
Related: #129 (MedQA vs MedMCQA cross-dataset consistency, @Agastya191) is adjacent - coordinate to avoid duplicate cascade runs.
Referee (was missing; the named contribution)