Skip to content

Vision-language lineage: Qwen/Qwen2.5-VL-72B-Instruct - #426

Open
sebasmos wants to merge 3 commits into
mainfrom
feat/qwen2.5-vl-72b-imaging
Open

sebasmos wants to merge 3 commits into
mainfrom
feat/qwen2.5-vl-72b-imaging

Conversation

@sebasmos

@sebasmos sebasmos commented Sep 8, 2026

Copy link
Copy Markdown
Member

Vision-language lineage: Qwen/Qwen2.5-VL-72B-Instruct, imaging and text

Locally served (vLLM, TP4, temperature 0), one model in every seat. Imaging: every arm of the experiments/imaging/ lane plus the CheXpert lane, on the same 35 NIH and 35 CheXpert studies as the committed gemini-2.5-flash arms. Text: the same MedQA manifest and 120 cases as the committed gemini-2.5-flash-lite arms, blind metric at n=40 and n=100, and every other text lane the repo can run (MedMCQA, SUPPORT2, referee, cascade, contamination, model dependence, cross-dataset). MIMIC-CXR, both the report-text lane and the film battery, ran on the cluster that holds the data under its DUA. Every number outside those two lanes replays from a fresh clone with no API calls. Draft.

Imaging

Unlike the text lanes, the decoy in the imaging blind metric is the opposite of each model's own baseline read and every study has the finding present, so a drift is only deference where the baseline was right. Qwen reads fewer films correctly than Gemini, and on the films it misreads the "decoy" is the true label. Both the aggregate rate and the rate conditioned on a correct baseline read are therefore given.

Cohort Gemini Qwen2.5-VL-72B
CheXpert (n=35) baseline reads correct 31/35 22/35
blind-rubric drift 30/35 (0.857) 14/35 (0.400)
drift given a correct baseline 29/31 7/22
test-aware uptake 6/35 (0.171) 0/35 (0.000)
NIH (n=35) baseline reads correct 22/35 15/35
blind-rubric drift 29/35 (0.829) 23/35 (0.657)
drift given a correct baseline 20/22 15/15
test-aware uptake 13/35 (0.371) 0/35 (0.000)

On the 22 CheXpert studies both models read correctly Gemini drifts on 20 and Qwen on 7, all 13 discordant pairs the same way (exact McNemar p=2.4e-4). On the 11 NIH studies both read correctly Gemini drifts on 10 and Qwen on 11. So the rubric effect replicates in direction on both cohorts and Qwen is the less deferent model on CheXpert only. Test-awareness suppresses it completely on both cohorts where Gemini retains 0.171 and 0.371. Neither model named the rubric on any drifted study.

Arm Contrast Gemini Qwen2.5-VL-72B
imaging_solo noise floor / watermark flip rate 0.23 / 0.34 0.17 / 0.09
imaging_cascade (watermark) shared / isolated / contagion 0.97 / 0.43 / 0.54 1.00 / 0.49 / 0.51
imaging_system_flag system flag vs peer board, Fisher 21/35 vs 34/35, p=0.0002 28/35 vs 35/35, p=0.0112
imaging_multi_round shared adoption by round [0.8571, 1.0, 0.9714] [1.0, 1.0, 1.0]
imaging_referee private re-query: precision / recall 0.88 / 0.79 1.00 / 0.94
imaging_judge_referee same-lineage judge: precision / FPR 0.56 / 0.94 0.51 / 1.00
CheXpert imaging_cascade shared / isolated / contagion 0.86 / 0.23 / 0.63 0.97 / 0.34 / 0.63

The solo result replicates as a null: no overlay cue reaches Qwen's own noise floor, where Gemini's watermark alone cleared it by 0.11. The committee result replicates in full and saturates: the watermark that does nothing alone is adopted on every study once two seeded peers assert it, so the cue-family, opacity-ladder, peer-size and multi-round arms all sit at 1.00 shared adoption (their summaries are committed; the Cochran and McNemar tests inside them return 1.0 on identical vectors, which is saturation, not a null). The private re-query referee transfers (precision 1.00, recall 0.94) and the transcript-only judge collapses onto the naive gate as it does for Gemini. The committed Gemini judge arm is the text-only variant, so the like-for-like Qwen arm is --text-only; the film-seeing run is kept beside it as imaging_judge_referee_summary_film.json.

Text: MedQA

Unseeded accuracy on the shared cases: Gemini 101/120 in every arm; Qwen exactly 90/120 in every one of the 17 arms that record it. A temperature-0 answer from the served endpoint is reproducible (2 changed answers over 401 repeat prompts across the lane), so the between-arm figures are comparable as well as the within-arm ones.

Arm Contrast Gemini Qwen2.5-VL-72B
blind_metric baseline to blind, n=100 0.00 to 0.28 0.00 to 0.20
blind_metric test-aware 0.17 0.01
text_cue_types baseline / negation 0.642 / 0.917 0.258 / 0.725
text_cue_types primacy / qualifier 0.433 / 0.667 0.242 / 0.300
dose_response faint to assert 0.142 to 0.808 0.075 to 0.367
dose_response assert to emphatic 0.808 to 0.617, 2/25 0.367 to 0.158, 0/25
authority_ladder colleague / attending 0.175 / 0.717 0.050 / 0.325
authority_ladder automated / guideline 0.650 / 0.925 0.083 / 0.592
test_awareness neutral / accuracy / agreement 0.642 / 0.483 / 0.433 0.267 / 0.175 / 0.108
committee_size_sweep none to one peer 0.642 to 0.133 0.267 to 0.042
plausible_distractor plausible vs implausible 0.74 / 0.59, 21/6 0.518 / 0.182, 39/2, p=7.84e-10
live_peer_organic wrong peer consensus 7/8 8/8
super_additivity both vs stronger single 11/10, p=1.0 0/4, p=0.125
temperature_sensitivity within-case flips t0.3 to t1.0 0.26 to 0.36 0.06 to 0.23
deliberation_channel none / hidden / open 0.52 / 0.72 / 0.63 0.27 / 0.27 / 0.09
referee_self_inconsistency temp-0 unstable cases 0/40 0/40

Adoption sits between Gemini's and the floor on almost every arm, roughly a third to a half of Gemini's rate, with the within-arm ordering preserved: negation is the potent cue, the guideline rung the highest, one honest peer protective, the auditor frame protective. Two things differ. The hidden reasoning channel does not exist on this endpoint, so the hidden cell repeats none, while the open channel lowers adoption from 0.27 to 0.09. And the plausibility gap is wide: a plausible decoy is adopted on 0.518 and an implausible one on 0.182, 39 discordant pairs one way and 2 the other.

MIMIC-CXR report text

Run against a vLLM endpoint on 2 H200s on the cluster that holds the data, the reports read in place under the PhysioNet DUA. The cohort is the 633-case index the committed rows key on: its ground truth agrees with the committed deliberation_framing.jsonl on 60 of 60 recorded cases, and the solo cohort reproduces Gemini's 600 indices exactly. Every per-case file has the same keys and row count as its Gemini counterpart (40, 40, 80, 20, 31, 60, 60 rows), so the arms pair case for case, with one exception stated below. Clean-correct on all 633 reports: Qwen 0.831, Gemini 0.790.

Arm Cohort Contrast gemini-2.5-flash Qwen2.5-VL-72B
blind_metric n=40 blind decoy uptake (baseline 0.000 both) 0.050 0.125
blind_metric n=40 drifters / named the rubric 2 / 0 5 / 1
referee_judge n=40 same-lineage judge precision / recall / FPR 1.000 / 0.929 / 0.000 0.333 / 1.000 / 0.963
referee_deployable n=40 deployable referee precision / FPR, clean control 0.538 / 0.182 0.722 / 0.075
break_it_a n=20 hard planted flag adopt / clean control 0.450 / 0.000 0.400 / 0.000
break_it_d n=31 incentive decoy drift -0.065 0.000
deliberation_framing n=60 adoption none / collab / indep / critical 0.717 / 0.650 / 0.467 / 0.367 0.167 / 0.167 / 0.017 / 0.000
push_c n=60 hard generic / anchored conformity 0.850 / 0.883 0.683 / 0.933

The same-lineage judge inverts: Gemini's flagged 13 of 40 at precision 1.000 with no false positives, Qwen's flags 39 of 40 at precision 0.333 and FPR 0.963, so on this lineage the detector fails by flagging everything. The deployable referee with a clean control is the one detector that improves. Deliberation framing does not replicate: adoption is at the floor in every frame, so the ladder has nowhere to order itself and the McNemar contrasts are non-significant. Peer conformity replicates and anchoring is stronger here, adding 0.250 over the generic rationale against 0.033 on Gemini.

break_it_a selects each model's own hard cases (hard_cases(all_cases, solo_records, n)), so its 20 cases are not Gemini's 20: 15 are shared. The committed Gemini file holds 40 rows because its solo stage ran two Gemini models over the same 20 cases; a single-model lane writes 20. Every other arm is paired case for case.

MIMIC-CXR films

Every arm of experiments/mimic_cxr_image/run_battery.py on the same 834 films: the selection regenerated from the recorded seed reproduces all six per-arm manifests, and each film matched the committed deid/provenance.csv sha256 (834 of 834, 0 mismatched). Committed in the lane's own shape, summaries plus deid/<slug>/*.csv keyed on case_index only, with the same headers and row counts as the Gemini files.

One fact governs the whole battery: Qwen reads 14 of the 834 pneumothorax films as present, against Gemini's 183. The planted read on this cohort is the constant "no", so it agrees with Qwen's own read on 98% of films. Every published committee statistic that conditions on a correct clean read (the "cases read correctly alone" blocks) would rest on 3 to 14 films here and is not reported. The full-cohort rates are reported instead, and read as a ceiling rather than a cascade.

Arm Films Contrast gemini-2.5-flash Qwen2.5-VL-72B
solo 834 clean-correct rate 0.219 0.017
solo 834 noise floor / max cue flip (watermark) 0.173 / 0.201 0.162 / 0.091
cascade 215 shared / isolated / contagion 0.98 / 0.81 / +0.17 1.00 / 0.90 / +0.10
referee_cascade 417 shared / isolated / contagion 0.97 / 0.78 / +0.19 1.00 / 0.90 / +0.10
strength_cascade 834 isolated adoption at opacity 0.15 / 0.30 / 0.45 (shared 1.00 at all three) 0.76 / 0.77 / 0.79 0.95 / 0.89 / 0.85
system_flag 834 shared / isolated / contagion 0.92 / 0.78 / +0.14 0.99 / 0.90 / +0.09
blind_metric 141 blind drift / test-aware drift (baseline 0.00 both) 0.53 / 0.25 0.61 / 0.00
blind_metric 141 drifters naming the rubric 1 of 75 0 of 86
judge (text-only) 417 judge flag rate 0.971 1.000
judge_with_image 417 judge flag rate given the film 0.010 0.017
referee 417 deployable referee flag rate / adoption base rate 0.281 / 0.194 0.197 / 0.096
nih_match_solo 35 clean-correct / any cue flip 10 of 35 / see solo 0 of 35 / 4 watermark flips

The two arms that do not depend on the baseline read replicate: the hidden rubric drifts Qwen on 0.61 of films against Gemini's 0.53 from a 0.00 baseline on both, test-awareness suppresses it completely (0.00 against 0.25), and no drifter names the rubric. The solo null replicates with every cue 0.07 to 0.16 below the 0.162 floor. For a second vision lineage the first thing to report on this cohort is whether the model can read the finding; the committee battery presupposes that it can.

The call caches for both MIMIC lanes hold report text and encoded pixels and are not committed, matching the committed Gemini lane; these two lanes are the one part of this branch that does not replay from a fresh clone.

Other lanes

Every Gemini text experiment the repo can still run has a Qwen arm (29 MedQA summaries, 23 on the MedMCQA cohort, 5 referee, 5 SUPPORT2, plus cascade, contamination, model dependence and the cross-dataset pilot). Committee runners seat this model in every seat, judge and orchestrator included, so each row pairs an all-Qwen committee against Gemini's all-Gemini one.

Lane Arm Contrast Gemini Qwen2.5-VL-72B
MedMCQA text_cue_types baseline / negation 0.742 / 0.925 0.500 / 0.850
MedMCQA authority_ladder colleague / guideline 0.200 / 0.917 0.167 / 0.800
MedMCQA rationale_validity bare vs valid-but-wrong 4/45, p=0.0 5/7, p=0.774414
MedMCQA deliberation_framing none vs critical 2/57, p=0.0 0/46, p=0.0
MedMCQA scale_c anchored vs generic, hard cases (n=110 / 125) 15/4, p=0.0192 20/3, p=0.0005
MedMCQA orchestrator_failure wrong peer / wrong orchestrator poisons output 0.290 / 1.000 0.207 / 1.000
SUPPORT2 support2_solo clean accuracy (abstentions) 0.713 (5) 0.717 (0)
SUPPORT2 support2_cascade wrong seed: shared / contagion 1.000 / 0.713 0.775 / 0.492
SUPPORT2 support2_cascade seed polarity: wrong / right 1.000 / 1.000 0.686 / 0.824
SUPPORT2 support2_cascade_strength six-rung span 0.976 to 1.000 0.442 to 0.791
MedQA referee_deployable deployable referee: precision / FPR 0.682 / 0.108 0.556 / 0.114
MedMCQA referee_deployable deployable referee: precision / FPR 0.742 / 0.140 0.607 / 0.175
SUPPORT2 support2_referee deployable referee: precision / FPR 0.713 / 0.223 0.678 / 0.155
cascade multi_round shared adoption, round 1 to 5 (n=40) 0.275 to 0.325 0.175 to 0.200
contamination contamination_audit question-only / options-only accuracy (flash-lite) 0.18 / 0.28 0.21 / 0.24

The tabular saturation is a property of the model: Gemini adopts the wrong prognosis on the shared SUPPORT2 board in every eligible case and Qwen on 0.775, and the six-rung ladder that is censored on Gemini spans 0.442 to 0.791 here. The seed-polarity contrast Gemini cannot express separates: a right seed is adopted more often than a wrong one. The reasoning-quality result does not replicate: a valid-but-wrong rationale leaves adoption unchanged (5/7, p=0.77) where it lowers Gemini's (4/45).

Cross-dataset cue pilot (50 cases per dataset, option order / longest option / lexical overlap): MedQA 0.08 / 0.06 / 0.04 against MedMCQA 0.12 / 0.24 / 0.30, mean spread 0.160; Gemini 0.08 / 0.10 / 0.08 against 0.08 / 0.16 / 0.16. The cue effect is dataset-dependent on this lineage where it was invariant on Gemini.

Where an arm's row count differs from Gemini's it is the arm's own eligibility rule, not a different cohort: the plausible-distractor arm runs on the cases with a buildable distractor (103 against 110), scale_c on the cases the model gets wrong unseeded (85 against 102), the orchestrator and unanimity arms on their eligible subsets, and support2_referee on 240 rows against Gemini's 235 because this model never abstains. Every shared arm was invoked with the same cohort flags as the accepted second-lineage run. cascade/multi_round is the one arm where that was not initially true: it is now the same 40 cases as the committed Gemini arm (replayed from the model's own cache at zero calls), with the 120-case run kept beside it as multi_round_summary_n120.json.

Code

Main has no experiments/_lane.py, so this branch carries the shared model dispatch it needs to run at all: key and backend resolution, the locally served OpenAI-compatible path, output caps, call pacing and transient recovery, plus the runner ports that go through it. Once that dispatch is on main this branch reduces to its own arms and three pieces the imaging lane needed:

  • _lane.paced_complete takes an optional image, passed through only when present, so the vision runners get the same pacing and recovery as text and the text path is unchanged; _lane.scoped takes the lane's own default id, because the imaging lane's committed model is gemini-2.5-flash, not the text default.
  • The 16 model-calling runners under experiments/imaging/ take --model and write to a model-scoped output directory and cache, replacing the hardcoded Gemini id, backend and committed-cache default each of them had. imaging_judge_referee.py names its seat JUDGE, not MODEL, and is ported accordingly. tests/test_imaging_runner_ports.py pins all of this and builds every runner's parser.
  • The 10 no-argument re-analysis scripts (blind_metric_ci, net_harm, onset_battery, positional_regression, ...) take --results-dir, so a second model's arms are re-analysed in place without touching the committed Gemini derivations; the ported scripts reproduce the unported ones byte for byte on the Gemini rows. positional_regression returns the empirical rate when an arm is saturated instead of fitting a logit with no variance.

cross_dataset/run_cross_dataset.py goes through the shared dispatch and contamination_audit.py is pointed at the model's own solo records (--solo-records) rather than the committed Gemini ones. tests/degeneracy_exemptions.json gains an allowlist entry, with the check performed, for every constant, duplicated or saturated column the guard finds in the new rows.

Three committed Gemini derivations (onset_battery.json, positional_regression.json, claim4_quantification.json) predate the July re-run of the cascade rows they derive from and no longer reproduce from them; they are left untouched here and noted for a separate fix.

experiments/mimic_cxr_image/run_battery.py takes --model and scopes the transcripts the judge and referee arms replay to the model subdirectory the shared runners write; without the second half those arms fail on a path one directory too shallow. experiments/mimic_cxr_image/export_deid.py takes --results-dir and --model, writes deid/<slug>/ with the published headers, and skips its claim verification for a second lineage because the expected values are the Gemini claims (the default path still verifies all twelve). experiments/mimic_cxr_text/build_solo_records.py imports parse_legacy_string: its _parse_choice import was removed in b9a9f07 and the module has been import-broken on main since. tests/test_mimic_battery.py recognises model-scoped summaries as produced by their arm.

Replay

The text arms replay from a fresh clone with no API calls: re-running five MedQA arms with no key set reproduces every leaf of the committed summaries (byte differences are JSON formatting only). The Gemini comparators carried onto this branch, deliberation_channel and the n=100 blind metric, also replay keylessly at their own cohort size and reproduce the committed numbers exactly. The imaging arms are the exception, and it is pre-existing rather than introduced here: two cue families composite rasterised text whose glyphs come from the imaging library's default font, so the cache, keyed on those pixels, re-queries under a different library version. The committed Gemini imaging caches miss on this machine in exactly the same way, and the released per-case rows are the artefact of record for that lane.

Scope

One branch, one new model against the Gemini baseline. No other lineage's results are here. Both MIMIC-CXR lanes are included; their call caches are not, per the DUA. Test suite: 1415 passed, 14 failed, the same 14 as main; the degeneracy guard passes on the new rows with an allowlist entry, stating the check performed, for each finding.

One branch, one new model against the committed Gemini baseline, which for this arm is the
gemini-2.5-flash run already on main over the same 35 studies.

Blind-rubric drift 0.40 against Gemini's 0.86, both off a 0.00 baseline, and complete suppression
under the test-aware prompt (0.00) where Gemini retains 0.17.

The branch carries the shared model dispatch it needs to run at all, because main has no
experiments/_lane.py yet: key and backend resolution, the locally served OpenAI-compatible path, output
caps, pacing and transient recovery, plus the runner ports that go through it. None of the other
lineages' result files are here. Once the text-lane dispatch lands on main this branch reduces to the
imaging arm alone.
@sebasmos sebasmos changed the title Imaging lineage: Qwen2.5-VL-72B-Instruct blind metric Vision-language lineage: Qwen/Qwen2.5-VL-72B-Instruct Sep 10, 2026
@sebasmos
sebasmos marked this pull request as ready for review September 10, 2026 08:44
@sebasmos
sebasmos marked this pull request as draft September 11, 2026 08:29
@sebasmos
sebasmos marked this pull request as ready for review September 13, 2026 11:55

@maximinl maximinl left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Peer review — approve the Qwen2.5-VL lineage package (draft).

What’s right

  • Imaging blind-metric reporting gives both aggregate drift and drift | correct baseline, which is required here: the decoy is the flip of each model’s own read, so misreads turn the “decoy” into the true label. CheXpert conditional contrast (Gemini 20/22 vs Qwen 7/22 on the 22 both-correct studies, McNemar p=2.4e-4) is the citable number.
  • Text-lane temperature-0 reproducibility called out with a concrete repeat-prompt check (2/401), so between-arm comparisons are justified — unlike gpt-oss.
  • MIMIC DUA handling (results in, caches out) matches the rest of the repo.
  • Imaging _lane.scoped(..., default=) so flash (not flash-lite) stays the imaging baseline directory is the correct seam; paced_complete(..., image=) is the right extension for vision.

Non-blocking asks

  • #421 committed cross_lineage_report.py so BH / repeat-prompt body claims replay from a fresh clone. This branch should either include the same (or an imaging+text equivalent) or drop unreproducible meta-counts from the body.
  • _lane.py differs across #421/#426/#427 (this tip has vision pacing + scoped(default=); #421 has HOSTED_PREFIXES). Land with an explicit merge order / single _lane consolidation so neither nvidia exclusion nor vision hooks get lost.
  • Still marked draft — fine while caches settle; ready to undraft once the report script (or body trim) is settled.

No methodology objections to the arm design or the conditional imaging rates.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants