Skip to content

Add SAM2/SAM3 browser segmentation and stereo auto-population - #1672

Merged
BryonLewis merged 41 commits into
mainfrom
sam3-onnx-web
Sep 25, 2026
Merged

BryonLewis merged 41 commits into
mainfrom
sam3-onnx-web

Conversation

@BryonLewis

@BryonLewis BryonLewis commented May 27, 2026 •

Copy link
Copy Markdown
Collaborator

DIVE Web now supports point-prompt segmentation with SAM2.1 Tiny (default) or SAM3, selected in Track Settings. The regular Segment controls replace the embedding-only prototype: positive/negative clicks produce editable polygons, preserve disconnected components and holes, and support confirm/reset/cancel. The selected ONNX model runs in the browser, loads on demand, and caches two camera-frame embeddings. Hardware WebGPU is preferred, with WASM fallback; software GPU adapters use WASM directly.

SAM2 and SAM3 now share a bundled 2 KB ONNX post-processing graph for best-mask selection, bilinear resize, unpadding and binary thresholding. Only the winning candidate is upscaled. DIVE receives a binary mask for polygon editing, while the graph and its weight-free exporter can also be used by desktop ONNX hosts. The exporter includes reference validation; the desktop service has not been switched to this new graph.

New boxes and lines can auto-populate masks and head/tail geometry using the shared annotation application code. In calibrated stereo mode, segmentation prompts and interior mask samples transfer through the existing correspondence model, followed by SAM inference on the other camera. Generated lines and mask bounds finish before the stereo length is refreshed, including subsequent multi-polygon refinements. Manual counterpart edits are preserved, and cancelled or stale work cannot replace newer geometry. Image and video frames are captured without annotation overlays.

Stereo transfer now follows desktop's Fast-FoundationStereo post-processing defaults: cached dense disparity, foreground-biased 90th-percentile sampling over a 7×7 neighbourhood at original-image pixel spacing, and an 11-sample straight-line disparity fit with up to three outliers and a 10-pixel residual. Failed fits retain endpoint matching; curves keep individually mapped vertices. Dense disparity is independent of NCC search limits.

Sparse stereo sampling now runs in a bundled 6 KB ONNX graph with a weight-free exporter (plugins/onnx/export_stereo_sampler.py in VIAME). Interpolation at native-image spacing, finite/positive filtering, clipped-neighbourhood counts and 90th-percentile selection are inside the graph. It gathers only 49 neighbours per point from cached dense disparity. Straight-line transfer batches all eleven samples into one graph call, including its endpoint fallback. The old JavaScript sampler is retained only as a test oracle; native desktop adoption of the graph is a separate change.

Mask transfer uses desktop's deep-interior/farthest-point sampling (five seeds overall, at least two per component), per-component median-offset filtering, and original labelled-click fallback if no interior matches survive. The other-camera mask must be within a 2.5× area ratio, accounting for holes. Browser rectification, model exports, and polygon rasterization can still differ from native results.

This branch incorporates the existing multi-polygon/auto-populate dependencies from #1952 and #1941 and is updated to current main. It has not been merged into viame/main. Native desktop inference continues to use VIAME; its auto-populate geometry application is now shared with the web provider.

Validation:

  • Full client tests, with focused coverage for component/hole conversion, head/tail extraction, prompt caching, model switching, stale results, cancellation, manual geometry protection, box/line auto-population, and repeated stereo mask refinement.
  • Client typecheck, lint (two existing Vue warnings), and web production build.
  • Real Chromium SAM2 smoke test passed with the new post-processing graph, producing the expected mask bounds.
  • The shipped SAM post-processing graph runs in ONNX Runtime/WASM tests; its exporter passes six native ONNX Runtime/OpenCV reference cases.
  • Sparse stereo ONNX validation: 24 native ONNX Runtime reference cases, browser/WASM parity tests, and real Chromium warm sampling at 960×576 model / 1920×1080 source resolution. Eleven samples took 1.9 ms median / 3.7 ms p95 over 60 runs on this machine (excludes model loading and dense inference).
  • Stereo parity regression tests use reference seeds generated by desktop OpenCV and disparity-fit results from compiled VIAME C++. The final full suite passed 1,943 tests with one optional model test skipped.
  • Real Chromium SAM2 ONNX runs: positive/negative points, box prompts, software WebGPU, WASM, and automatic software-adapter fallback.
  • Real Chromium SAM3 ONNX/WASM run: positive and negative prompts produced valid masks (about seven minutes including initialization and both predictions on this CPU). Software-WebGPU SAM3 did not finish within ten minutes; the loader now avoids that adapter type. Hardware-GPU performance still needs validation.

The browser path supports image sequences and loaded video frames; tiled large-image datasets remain unsupported. The SAM3 option is its point/box tracker model, not text-prompt search. Model files download from Hugging Face; image pixels remain local. SAM3 is substantially heavier than SAM2, and hardware-GPU performance still needs validation.

The helper exporter sources have moved to VIAME’s plugins/onnx/ in VIAME/VIAME#303, using snake_case names and the ONNX plugin installation list. This PR removes the DIVE copies and retains both bundled ONNX assets. Regeneration passes all 30 reference cases and changes only model producer/documentation metadata, with computational graphs unchanged.

BryonLewis and others added 22 commits May 27, 2026 12:50
With auto compute on the other camera, the mapped shape gets the same mask
and/or points pass once the transfer succeeds; a derived head/tail follows
the source camera's direction.
Lets the VIAME service keep a line-prompted mask in scale with the line.
A box warped to the other camera whose mask overlaps it by less
than half its union takes the mask's bounds instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The service's polygons list, with holes, is stored as keyed polygons on the
detection; a refinement drops the components it no longer has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… is on

A confirmed mask on a brand-new detection now takes the same keypoint pass
as a drawn box, on the source camera and on the stereo copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With auto-segmentation on, points inside the source mask are warped instead
of the box corners, the other camera is segmented from them, and its box is
that mask's bounds. Corner warping remains the fallback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…moves it

The pass runs after each click's prediction on both cameras, replacing only a
line it derived itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The other camera's mask comes from the service, which seeds from inside the source mask and refuses an out-of-scale result, instead of DIVE warping prompts and predicting itself. Click segmentation sends every source polygon and draws every returned one.
@mattdawkins mattdawkins changed the title SAM3 Onnx Web Support Add SAM2/SAM3 browser segmentation and stereo auto-population Sep 22, 2026
# Conflicts:
#	client/dive-common/apispec.ts
#	client/dive-common/components/EditorMenu.vue
#	client/dive-common/components/TrackSettingsPanel.vue
#	client/dive-common/recipes/segmentationpointclick.spec.ts
#	client/dive-common/store/settings.ts
#	client/dive-common/use/autoPopulate.spec.ts
#	client/dive-common/use/autoPopulate.ts
#	client/dive-common/use/useModeManager.spec.ts
#	client/dive-common/use/useModeManager.ts
#	client/platform/desktop/backend/native/interactive.ts
#	client/platform/desktop/backend/native/segmentation.ts
#	client/platform/desktop/frontend/autoPopulate.spec.ts
#	client/platform/desktop/frontend/autoPopulate.ts
#	client/platform/desktop/frontend/components/ViewerLoader.vue
ImageData's data, width and height are prototype getters, so spreading it
produced an object without pixels; Firefox then failed the first SAM click.
When the higher quality model fails to load or to run on this browser,
the setting drops to template matching so it matches what runs, and the
error says so instead of asking the user to switch.
GeoJS ends editing on its own (right-click, completion) and leaves the
box behind in disabled mode; skipping disable() there kept it as a
ghost on every frame with a hover highlight.
A failed load was cached for the session and surfaced later as a bare
'could not be loaded'; now the load is retried on the next warp and the
caller's error carries the underlying reason.
@BryonLewis
BryonLewis marked this pull request as ready for review September 25, 2026 15:37
@BryonLewis
BryonLewis merged commit 18cb766 into main Sep 25, 2026
3 checks passed
@BryonLewis
BryonLewis deleted the sam3-onnx-web branch September 25, 2026 15:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants