docs/detection_recall_analysis.md conditions its central result — recall by distance, and by
apparent size — on a distance axis computed from flat-ground geometry at a fixed camera height
(depth_extract_da3.py: CAM_H = 2.5). GSV serves its own metric depth, and it says that axis is
stretched by roughly a quarter to a third.
The measurement
GSV's depth payload is a list of planes plus a per-pixel index; the dominant ground plane's distance
is the camera height, exactly, and its normal is the ground tilt. Read that way, camera height is
per-panorama and mostly well below 2.5 m — it tracks capture vintage, with 2025/2026 imagery
lowest (medians 1.89 m and 1.70 m). Details and the survey are in
sidewalk-auto-labeler#40.
Measured at 272 operational detections across 120 panoramas (paterson, gainesville), comparing
flat-ground range against the depth map sampled at the detection pixel:
|
paterson |
gainesville |
| median range, flat-ground @ 2.6 m |
14.98 m |
14.98 m |
| median range, depth |
11.37 m |
10.40 m |
| ratio |
1.286 |
1.354 |
| ratio after correcting only the camera height |
1.029 |
1.019 |
Depth was valid at 100% of those detection points. Correcting the height alone flattens the
residual across every range bucket (1.017 / 1.021 / 1.035 / 1.018 in paterson) — so the cotangent
form is right and only the constant was wrong. At CAM_H = 2.5 rather than 2.6 the stretch is
slightly smaller, on the order of 24–30%.
What moves
1. The distance buckets. "Reliable to ~18 m and effectively blind past 25 m" is stated on the
stretched axis. Deflated by ~25%, those thresholds are closer to ~14 m and ~20 m — a materially
different claim about the model, and the headline of the document.
2. Apparent size, in the same direction and for the same reason. The doc derives subtended pixels
as W/d × (4096/2π), so an overestimated d underestimates apparent size. Every size bucket edge
("20–32 px → recall 0.189") shifts upward, meaning the ramps RampNet misses are larger in pixels
than currently reported.
3. Therefore the retraining-resolution argument needs re-checking. The doc concludes "there is
simply not enough signal left in the pixels" and prices a higher-resolution retrain at +10 pts
saturating ~0.88. If misses are nearer and larger than the axis says, that diagnosis is weaker than
it appears, and the ranking of the three levers may change. This is the part worth redoing before any
retraining decision rests on it.
4. The DA3 cross-check becomes a calibration. §0 reports DA3 and flat-ground agreeing to 6.5–8.5%,
read as mutual confirmation. Against real depth, both are long — so the agreement was two methods
sharing a bias, not two methods being right. The useful move now is to calibrate DA3 against GSV
depth and carry the calibrated model to Mapillary, which serves no depth of its own and is where
flat-ground is worst (the ρ = 0.81 vs 0.95 gap in §0 is exactly that).
Second, unrelated use of the same data: occlusion
The depth payload's per-pixel plane index is a free segmentation of the scene into planar
surfaces. A ramp that is missed because a parked car, a utility pole or a street tree sits in front of
it is a fundamentally different failure from one missed because it is small and far — and the two are
currently indistinguishable in the miss taxonomy, where "small and far away" is the leading
hypothesis for everything.
Concretely: for each missed GT ramp, compare the depth along its ray against the ground-plane
prediction at that pixel. A large negative discontinuity means something is between the camera and the
ramp. That partitions the 150 misses into occluded versus genuinely under-resolved, and only the
second group is addressable by resolution or by denser panorama sampling.
Scope
Depth is fetched via streetlevel with download_depth=True — a flag on the metadata request, not an
extra round trip. The auto-labeler side is harvesting it for all four GSV cities; benchmark panoramas
are a small subset of that and are the highest priority.
🤖 Generated with Claude Code (claude-opus-5[1m])
docs/detection_recall_analysis.mdconditions its central result — recall by distance, and byapparent size — on a distance axis computed from flat-ground geometry at a fixed camera height
(
depth_extract_da3.py:CAM_H = 2.5). GSV serves its own metric depth, and it says that axis isstretched by roughly a quarter to a third.
The measurement
GSV's depth payload is a list of planes plus a per-pixel index; the dominant ground plane's distance
is the camera height, exactly, and its normal is the ground tilt. Read that way, camera height is
per-panorama and mostly well below 2.5 m — it tracks capture vintage, with 2025/2026 imagery
lowest (medians 1.89 m and 1.70 m). Details and the survey are in
sidewalk-auto-labeler#40.
Measured at 272 operational detections across 120 panoramas (paterson, gainesville), comparing
flat-ground range against the depth map sampled at the detection pixel:
Depth was valid at 100% of those detection points. Correcting the height alone flattens the
residual across every range bucket (1.017 / 1.021 / 1.035 / 1.018 in paterson) — so the cotangent
form is right and only the constant was wrong. At
CAM_H = 2.5rather than 2.6 the stretch isslightly smaller, on the order of 24–30%.
What moves
1. The distance buckets. "Reliable to ~18 m and effectively blind past 25 m" is stated on the
stretched axis. Deflated by ~25%, those thresholds are closer to ~14 m and ~20 m — a materially
different claim about the model, and the headline of the document.
2. Apparent size, in the same direction and for the same reason. The doc derives subtended pixels
as
W/d × (4096/2π), so an overestimateddunderestimates apparent size. Every size bucket edge("20–32 px → recall 0.189") shifts upward, meaning the ramps RampNet misses are larger in pixels
than currently reported.
3. Therefore the retraining-resolution argument needs re-checking. The doc concludes "there is
simply not enough signal left in the pixels" and prices a higher-resolution retrain at +10 pts
saturating ~0.88. If misses are nearer and larger than the axis says, that diagnosis is weaker than
it appears, and the ranking of the three levers may change. This is the part worth redoing before any
retraining decision rests on it.
4. The DA3 cross-check becomes a calibration. §0 reports DA3 and flat-ground agreeing to 6.5–8.5%,
read as mutual confirmation. Against real depth, both are long — so the agreement was two methods
sharing a bias, not two methods being right. The useful move now is to calibrate DA3 against GSV
depth and carry the calibrated model to Mapillary, which serves no depth of its own and is where
flat-ground is worst (the ρ = 0.81 vs 0.95 gap in §0 is exactly that).
Second, unrelated use of the same data: occlusion
The depth payload's per-pixel plane index is a free segmentation of the scene into planar
surfaces. A ramp that is missed because a parked car, a utility pole or a street tree sits in front of
it is a fundamentally different failure from one missed because it is small and far — and the two are
currently indistinguishable in the miss taxonomy, where "small and far away" is the leading
hypothesis for everything.
Concretely: for each missed GT ramp, compare the depth along its ray against the ground-plane
prediction at that pixel. A large negative discontinuity means something is between the camera and the
ramp. That partitions the 150 misses into occluded versus genuinely under-resolved, and only the
second group is addressable by resolution or by denser panorama sampling.
Scope
"not enough signal in the pixels" conclusion on the corrected axis
be used on Mapillary with a measured rather than assumed bias
Depth is fetched via
streetlevelwithdownload_depth=True— a flag on the metadata request, not anextra round trip. The auto-labeler side is harvesting it for all four GSV cities; benchmark panoramas
are a small subset of that and are the highest priority.
🤖 Generated with Claude Code (claude-opus-5[1m])