Skip to content

Add drop-test protocol, Edison synthesis, and first-data analysis - #86

Draft
sgbaird with Copilot wants to merge 67 commits into
copilot/add-drop-test-protocolfrom
copilot/add-drop-test-protocol-again
Draft

Add drop-test protocol, Edison synthesis, and first-data analysis#86
sgbaird with Copilot wants to merge 67 commits into
copilot/add-drop-test-protocolfrom
copilot/add-drop-test-protocol-again

Conversation

Copilot AI commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

Issue is operational (coordinating the first crush/drop tests on Jeff Hill's tower) and this repo is the LaTeX MRG proposal with no code or existing test-protocol surface. This PR consolidates the moving parts from the issue thread, adds a literature synthesis, and analyzes the recorded accelerometer data.

Added

  • docs/drop-test-protocol.md — single source of truth for the drop-test setup, covering:

    • Equipment + links to the TP4 Quick Start / User's Guide PDFs attached on the issue and the training video (https://youtu.be/RNjpAmWWmkQ)
    • The tower is bungee-assisted (base accelerates past 1 g): §3.1 reframes the pre-impact specimen lift-off as intrinsic rig physics rather than a setup artifact, and §5 leads with Jeff's first-pass fix (cap how far the specimen's top can rise relative to the base) plus tie-to-base options
    • Quantities of interest tied to the BO objective stack: g_max, SEA, full ~10 s ringdown (not just the 200 ms shock), reusability, slow-mo framing from t=0
    • Failure modes from the first instrumented drop: bungee-driven specimen lift-off pre-impact, ~25° cage tilt from loose rod/hole clearance, slow-mo starting after hoist release
    • @sgbaird's three-test next-iteration plan: bare specimen → plate-only (uninstrumented) → instrumented cage drop
    • Mitigations: constrain specimen to base / cap upward travel, tighter rod/plate tolerance (re-drill or thin metal plates, optionally linear bushings), top-plate retention clips, longer-term vertex-mounted accelerometer inside an acrylic cage, independent lab access for all three students
    • Cross-references to companion modalities out of scope for the first drop (high-speed camera, shaker transfer function, slug-firing gas gun, Polytec LDV)
  • edison-trajectories/drop-test/ — Edison Scientific LITERATURE_HIGH synthesis (task 653d7d39) on drop-tower troubleshooting for small 3D-printed lattice/tensegrity specimens: ~57 KB report, full JSON dump, submission record, and README. Idempotent driver at scripts/edison/submit_drop_test.py. Surfaces a standards stack (ASTM D5276/D7136/D3332, ISO 6603/1683/5347, MIL-STD-810 method 516, SAE J211) and recommendations (linear sleeve bearings, magnetic/elastic top-plate hold-down, ≥10 s ring-buffer DAQ, SAE J211 CFC filtering, n ≥ 5 + CV, ≥5000 fps DIC; closest analogues Pajunen 2019, Dwyer 2023).

  • First drop-test data analysis of the five TP4 accelerometer exports posted by @me-madsen (Signal 10–14, 4-channel, 125 kHz, 0.2 s window):

    • data/drop-tests/raw/ — committed raw export files
    • scripts/analysis/drop_test_analysis.py — loader, SAE J211 CFC-1000 / CFC-180 filtering, peak/pulse/PSD metrics, and figure generation
    • data/drop-tests/figures/ — full-window CH1 overlay, per-run impact zoom (raw vs filtered), peak-g bar chart, PSD, and CH4 trigger-artifact plot
    • docs/drop-test-analysis.md + data/drop-tests/README.md — findings: the "audrey" tensegrity specimen reduces CFC-180 peak acceleration ~74–79 % vs the no-specimen control (~370–463 G vs ~1,792 G), while the PETG run's raw peak is within ~1 % of the control (≈ direct plate-on-plate hit), strong evidence of the bungee-driven lift-off; CH4 carries a fixed ~1.4 kG trigger/release artifact at t≈4.2 ms in every run. Caveats noted: unconfirmed channel map, 200 ms window only, n = 1 for control/PETG, no Δv/SEA quoted yet.
  • Vertex vs. acrylic-plate T3-prism drop-test data and analysis posted by @ctrhjk (PR Add drop-test protocol, Edison synthesis, and first-data analysis #67), under data/drop-tests/vertex-acrylic/:

    • raw/ — eight TP4 exports {n0jdwk, m6cyoq, T3_0103, T3_0000}_Signal{1,2}.csv (Signal1 = vertex-mounted, Signal2 = acrylic-plate), single drop per configuration at 13 ft, 200 ms / 125 kHz
    • README.md — channel map (CH1 removed; CH2–CH4 tri-axis; CH5 single-axis; CH4 = 1000 G trigger) with full-scale/sensitivity per channel, per-specimen file index, and @ctrhjk's observations (clip-height/no-trigger issue, hot-glue z-axis mount limitation, m6cyoq strut and T3_0103 TPU-tendon damage after the acrylic test, and the invalid T3_0000 acrylic run where the accelerometer fell off)
    • scripts/analysis/drop_test_vertex_acrylic_analysis.py — locates the impact via the triggered CH4 channel (windowed ±1.5 ms peak search in the first 10 ms, not a global max), baseline-corrects, and reports raw / SAE J211 CFC-1000 / CFC-180 peaks for the single-axis CH5 (primary go-forward sensor) and tri-axis CH4, auto-flagging invalid / no-clean-impact runs
    • data/drop-tests/vertex-acrylic/figures/ — vertex-vs-acrylic CH5 impact windows, CFC-180 peak-g bar chart, and vertex CH5 PSD
    • docs/drop-test-vertex-acrylic-analysis.md — findings + an explicit SOP / test-method section: vertex mounting is repeatable (4/4 clean, CFC-180 229–284 G, CV ≈ 9 %) while the acrylic configuration is not (3/4 runs registered no clean impact — clips too low so the plate seats on the specimen, plus the fell-off T3_0000); the vertex peaks do not yet discriminate geometry so fresh intact distinct-geometry samples (vertex-only, n ≥ 5) are needed before peak-g is a trustworthy BO objective; replace the hot-glue mount with a z-axis-aligned seat; and the single-axis sensor's raw peaks reach 70–90 % of its 9,442.9 G full scale (near saturation on the m6cyoq-acrylic run). Caveats: n = 1 per (specimen, mount), 200 ms window only, partial-pulse Δv, unconfirmed CH4/CH5 axis correspondence.
  • Clip-height sweep & base-plate accelerometer-check diagnostic posted by @ctrhjk, under data/drop-tests/clip-height/ — drilling into why the acrylic-plate configuration repeatedly fails to trigger:

    • raw/Accelerometer_check_Signal1.csv + README.md — the one triggered base-plate CSV (tri-axis on the bottom plate, 13 in drop) plus a setup README documenting both experiments: the clip-height sweep (extra bungees cured fly-off; tri-axis on the acrylic plate; clips at 0.5/1/1.5/2 in, two drops each; 0/8 drops triggered, video only) and the base-plate accelerometer check, with the shared channel map
    • scripts/analysis/drop_test_clip_height_analysis.py + data/drop-tests/clip-height/figures/ — windowed CH4 impact location, SAE J211 CFC-1000 / CFC-180 peak/pulse/Δv metrics, and figures (base-plate impact window, full-window CH4, PSD)
    • docs/drop-test-clip-height-analysis.md — findings: the base-plate hit triggers cleanly (CH4 raw 3072 G ≈ 3.1× the 1000 G trigger, CFC-180 280 G, Δv ≈ 3.3 m/s; CH4 dominates the off-axis channels ~23–55×), so the acrylic-plate "no trigger" failure (0/8 across the clip sweep) is a load-path problem — the plate seats on / is damped by the bungee-restrained specimen — not the sensor, DAQ, or trigger level.
  • Input-output (transmissibility) drop-test data and analysis posted by @ctrhjk (PR Add drop-test protocol, Edison synthesis, and first-data analysis #67), under data/drop-tests/input-output/@ctrhjk's input-output instrumentation design: a single-axis accelerometer on the bottom plate = input (now the triggered channel CH5), a tri-axis accelerometer hot-glued to the top vertex = output (CH2–CH4), bungees removed, four distinct-geometry specimens (practice, n0jdwk, yqpmx1, h8Lbev) each dropped five times at 13 in:

    • raw/ — 20 TP4 exports {practice,n0jdwk,yqpmx1,h8Lbev}_Signal{1..5}.csv (Signal index = drop number) + README.md with the channel map (trigger moved to the single-axis input CH5) and @ctrhjk's setup notes
    • scripts/analysis/drop_test_input_output_analysis.py — locates the impact on the triggered CH5 (windowed ±1.5 ms peak), baseline-corrects, and reports raw / SAE J211 CFC-1000 / CFC-180 peaks for the input (CH5) and the tri-axis output resultant, the transmissibility T = output/input, pulse width and Δv, with per-specimen mean ± 1σ / CV aggregates
    • data/drop-tests/input-output/figures/ — input-vs-output impact windows (5 drops overlaid), transmissibility bar chart, input repeatability, output PSD
    • docs/drop-test-input-output-analysis.md — findings: the input-output design works — 20/20 drops triggered cleanly, removing the bungees makes the input nearly constant (235–248 G CFC-180, ≤1.7 % CV), and transmissibility now discriminates geometry (yqpmx1 ≈ 0.96 is the only attenuator, h8Lbev ≈ 1.09, practice/n0jdwk ≈ 1.17–1.19), making T (or output-peak-at-fixed-input) a usable BO objective; a mild within-run drift across the five cyclic drops is flagged as most likely hot-glue-mount-driven. Caveats: n = 1 specimen per geometry (5 repeat drops), 200 ms window, unverified tri-axis orientation, IDs not yet tied back to design parameters
    • edison-trajectories/input-output/ — Edison Scientific ANALYSIS (task fe044079) that independently reproduced the transmissibility values exactly to two decimals, confirmed the within-run drift is statistically real and mount-driven (pooled +0.015/drop, p = 0.0001), endorsed T as a first-pass screening objective (recommending FRF / SRS-band metrics and output-peak-at-fixed-input as it matures), and gave a prioritized SOP (rigid z-aligned keyed sensor seat, keep bungees removed, extend capture past 200 ms, n ≥ 5 distinct prints per geometry with randomized order, anchor in SAE J211 / ISO 5347 / ASTM D3332). Idempotent driver scripts/edison/submit_input_output.py + fetch scripts/edison/fetch_input_output.py; a cross-ch...
  • Drops-per-specimen variance, sample-size, and timing meta-analysis answering @me-madsen's question (PR Add drop-test protocol, Edison synthesis, and first-data analysis #82 comment 5026945744: minimum drops per specimen, variance so far, and set duration at ~42 s/drop at 60 in), under data/drop-tests/sample-size/ (derived — no raw data of its own):

    • scripts/analysis/drop_test_sample_size_analysis.py — aggregates the within-specimen coefficients of variation already emitted by the committed per-dataset analyses (large-campaign stabilized_ols/burn_in_drops, the felt-sheet per-condition CVs, and the n = 5 mount-validation writeups), then computes precision-based sample sizes (n = (t·CV/MoE)²), two-design discrimination power, and timing; emits figures/sample_size_metrics.json + two figures
    • docs/drop-test-sample-size-analysis.md + folder README.md — findings: the go-forward output (top-vertex tri-axis CFC-180 peak / transmissibility T) has pooled within-specimen CV 0.31–3.5 % (median ≈ 1.0 %), best ≈ 0.3–0.6 % on the wax key-seat mount, while the base-plate/input channel is the noisy near-saturation one. Recommendation: baseline 2 warm-up (discarded) + 5 recorded = 7 drops/specimen (5 recorded gives a ±1.2 % 95 % CI on the per-specimen mean and resolves the ≥10 % between-design differences), stepping to 10 recorded for the noisier input/T channel or a hot-glue mount. Timing at ~42 s/drop (60 in, auto): baseline ≈ 4.9 min/specimen; a 20-design BO batch × 7 drops ≈ 1.6 h (lower drop heights log ~12–20 s/drop, ~2.5× faster)
    • Dataset index (data/drop-tests/README.md) and .gitignore (__pycache__/, *.pyc) updated

    Note: I cannot open new PRs (only push commits to this branch), so @sgbaird's request to create a new PR is not actioned here; work continues on this branch.

@sgbaird

sgbaird commented Jul 20, 2026

Copy link
Copy Markdown
Collaborator

@me-madsen claude pings will probably work here now, but you might want to point back to specific comments or areas from the previous 2 PRs: #82 and #67 (now the work is split across three PRs)

Copilot AI changed the title [WIP] Add drop-test protocol, Edison synthesis, and first-data analysis Add drop-test protocol, Edison synthesis, and first-data analysis Jul 20, 2026
Copilot AI requested a review from sgbaird July 20, 2026 23:45
@ctrhjk

ctrhjk commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

@claude

We tested the combination of 60 in of drop height and 5 felt sheets. We used two specimens, 7xadt6 and 9GMQYQ.

The test setup was the same as I used in #82 (comment), and a single-axial accelerometer and a tri-axial accelerometer were used for this test, which means that we still got rid of the bottom tri-axial accelerometer.

Ch 5 measured data from a single accelerometer attached to the bottom acrylic plate, and Ch 2~4 measured data from a tri-axial accelerometer attached to the top vertex of the tensegrity structure. Ch2 is connected to the X axis, Ch3 to the Y axis, and Ch4 to the Z axis.

This test is to prove the validation of the experiment setting using a 60-in drop height and 5 felt sheets.
Conduct the OLS regression and report the results to me. Also, suggest an appropriate drop height and number of felt sheets if you have them after analyzing the data.

Here are the data
7xadt6:
https://github.com/vertical-cloud-lab/tensegrity-optimization/tree/copilot/add-drop-test-protocol/data/drop-tests/7xadt6%20_60in_5felts%20folder

9GMQYQ:
https://github.com/vertical-cloud-lab/tensegrity-optimization/tree/copilot/add-drop-test-protocol/data/drop-tests/9GMQYQ_60in_5felts

@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @ctrhjk's task in 9m 42s —— View job


60 in / 5 felts validation analysis (7xadt6 & 9GMQYQ) ✅


TL;DR

The 60 in / 5 felt setting is validated — keep it — but the felt is a consumable. 201/201 captures triggered cleanly, the output is the strongest and most repeatable of any condition tested (461–475 G, CV ≤ 1.7 %), and the two geometries separate decisively. The one campaign-scale effect in the data is felt compaction: the raw CH5 spike tripled (2.1 → 6.5 kG, 22 % → 69 % of full scale) across the two back-to-back sessions, which is what drives every drift term below. Transmissibility T = TOP/CH5 cancels almost all of it.

Capture health

7xadt6 9GMQYQ
real drops / captures 100/100 101/101
impact lands at 4.06 ± 0.07 ms 4.00 ± 0.04 ms
median cadence 41 s 41 s (one 12.9 min pause before Signal 101)
worst CH5 raw 35.1 % FS 68.6 % FS
worst top-vertex axis (CH4/Z) 20.1 % FS 22.1 % FS

OLS regression (stabilized phase, burn-in = 5 drops)

The burn-in scan finds no k ≤ 20 with a non-significant trend on either specimen — the drift is campaign-scale felt wear, not a seating transient — so the SOP 5-drop burn-in is used.

7xadt6 (drops 6–100, n = 95):

metric mean CV slope (%/drop) 95 % CI (per drop) p DW
TOP CFC-180 (G) 461.4 1.74 % +0.062 % [+0.275, +0.298] G 1.4e-68 0.96 1.10
CH5 input (G) 446.2 1.66 % +0.059 % [+0.252, +0.274] G 3.2e-66 0.96 1.05
T = TOP/CH5 1.034 0.12 % +0.003 % 2.7e-17 0.54 1.76

9GMQYQ (drops 6–101, n = 96):

metric mean CV slope (%/drop) 95 % CI (per drop) p DW
TOP CFC-180 (G) 475.2 1.00 % +0.028 % [+0.110, +0.155] G 4.2e-20 0.59 0.77
CH5 input (G) 462.7 0.81 % +0.013 % [+0.036, +0.085] G 4.7e-06 0.20 0.59
T = TOP/CH5 1.027 0.45 % +0.015 % 4.3e-37 0.82 0.48

Reading it: the TOP output climbs ~0.03–0.06 %/drop, but in lockstep with the input (per-drop corr(TOP, CH5) = 0.999 for 7xadt6) — so T absorbs it, drifting only +0.3 % / +1.4 % over ~95 drops. Pulse width and ringdown frequency are flat → this is a rig/felt effect, not specimen degradation. (Durbin–Watson 0.5–1.1 means residuals are autocorrelated — trust the slopes and R², not the literal p-values.)

Specimen discrimination (stabilized drops, Welch t): TOP 461.4 vs 475.2 G (−2.9 %, p = 3.7e-30, |d| = 2.1); T 1.034 vs 1.027 (+0.7 %, p = 6.4e-27, d = 2.1). These two geometries happen to be close responders, yet they still separate cleanly — sub-1 % differences in T are resolvable at n ≈ 5 drops.

The felt-wear finding

CH5's raw spike grows monotonically all evening: 7xadt6 starts at the fresh-felt level the sweep measured (~2.1 kG ≈ 22 % FS) and crosses the FS/3 head-room target around drop 67; 9GMQYQ (run second) is above FS/3 from drop 1 and ends at ~6.5 kG (69 % FS). The CFC-180 input barely moves (446 → 463 G) — compacted felt mostly adds high-frequency spike content, which is exactly what eats raw-peak head-room. Signal 101, recorded after the 12.9 min pause, drops back to 56 % FS — consistent with viscoelastic felt recovery. (See figures/03_saturation.png.)

Recommendation on height / felt count

  1. Keep 60 in / 5 felt sheets as the operating point — no change needed to the condition itself.
  2. Replace/refresh the felt when CH5's raw |peak| exceeds FS/3 ≈ 3.1 kG — roughly every ~100 drops at 60 in (ideally fresh felt per specimen), and log the felt state per session.
  3. If mid-campaign felt swaps are impractical: 60 in / 6 felt sheets. One extra sheet attenuates the base hit ×0.36 (felt-sheet model) — would have held even this evening's worst case at ~25 % FS — at near-zero cost in output signal. Fallback if hoisting is the bottleneck: 50 in / 4–5 felt.
  4. Use T = TOP/CH5 (not raw TOP peak) as the BO objective — it cancels the felt-wear drift (CV 0.12–0.45 % vs 1.0–1.7 %).

One question back, @ctrhjk: was the felt stack changed between the two specimens? The analysis assumes the same stack all evening (the monotonic CH5 trend suggests so). If it was changed, the 9GMQYQ starting level (~42 % FS) means felt sheets vary unit-to-unit — which makes logging felt state and normalizing by the input even more important.
· Branch

Stabilized-phase OLS drift, felt-wear saturation audit, and specimen
discrimination for the two 100-drop validation campaigns at the
felt-sheet sweep's recommended operating point.

Co-authored-by: Jinkwan Han <282969777+ctrhjk@users.noreply.github.com>

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ctrhjk

ctrhjk commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

@claude

The stack of felt sheets wasn't changed

@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @ctrhjk's task in 1m 5s —— View job


Felt-stack confirmation ✅


Thanks @ctrhjk — that confirms the analysis. With a single felt stack all evening:

  • The monotonic CH5 raw-spike growth (2.1 → 6.5 kG over 201 drops) is cumulative felt compaction, exactly as read. The 9GMQYQ starting level (~42 % FS) isn't unit-to-unit felt variation — it's simply where 7xadt6's 100 drops left the stack.
  • No conclusions change; the confirmation removes the one alternative explanation. The recommendations stand as written: keep 60 in / 5 felt sheets, refresh the felt when CH5's raw |peak| crosses FS/3 ≈ 3.1 kG (~every 100 drops at 60 in, ideally fresh felt per specimen), log drops-on-stack per session, and use T = TOP/CH5 as the BO objective since it cancels the felt-wear drift (CV 0.12–0.45 %).
  • Practical corollary now that the wear curve is known to be one stack's trajectory: a third same-evening campaign on that stack would have pushed CH5 toward saturation — so either swap felt between specimens or run 60 in / 6 sheets if mid-campaign swaps are impractical.

The doc now records the confirmation in §3 and closes the corresponding caveat in §6.

--- · Branch

@me-madsen

Copy link
Copy Markdown
Collaborator

@claude recommend a minimum number of drop tests per specimen to get accurate data. How much variance have we had in our data per specimen being tested so far? Also provide how long each set of tests would take given that it currently takes ~42 seconds/drop at 60 inches with automatic dropping. For evaluating this, reference the data from #86 (comment) and #82 (comment) and #67

@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @me-madsen's task in 4m 45s —— View job


Minimum drops per specimen, variance so far, and set timing ✅


TL;DR

Minimum: 5 recorded drops per specimen, after 2 discarded warm-up drops → 7 total ≈ 4.9 min/specimen at 60 in. This was the recommendation from the earlier meta-analysis, and folding in the two datasets you referenced (the 60 in/5 felts validation and the 5-vs-10-in comparison) strengthens it — the newest, largest datasets are among the tightest we've recorded.

1. How much variance per specimen so far?

Every dataset where one specimen was dropped repeatedly, with the within-specimen coefficient of variation (CV) of the go-forward metrics (top-vertex CFC-180 output peak, and transmissibility T = output/input):

dataset drops output CV T CV
60in-5felts 7xadt6 (#86) 95 1.7 %* 0.12 %
60in-5felts 9GMQYQ (#86) 96 1.0 %* 0.45 %
5vs10 5 in / 10 in (#82) 30 + 30 0.7 % / 0.3 % 0.4 % / 0.8 %
drift-calibration 1+2 (wax mount) 19 + 45 0.5–0.6 % 2.0–2.9 %
100drops / 200drops / 5in-100drops 91–190 1.0–2.5 % 0.4–2.4 %
felt-sheet sweep (9 cond × 5) 45 0.2–2.1 %
input-output, hot-glue mount (#67) 4 × 5 1.3–3.5 % 0.8–4.6 %
key-mounted / wax / burn-in-wax (n = 5 each) 15 0.3–1.3 % 1.1–2.4 %

* the 60 in TOP/CH5 CVs are inflated by the campaign-scale felt-wear drift (both channels climb in lockstep); T cancels it, which is why its 0.12–0.45 % over ~95 drops is the tightest of any dataset.

Pooled: output CV 0.31–3.5 %, median ≈ 1.0 % (90th pct ≈ 2.3 %); T median CV ≈ 1.2 %. The worst numbers all trace to the retired hot-glue mount; the key-seat (± wax) mount sits at 0.3–1.0 %. For scale, between-design differences we've measured are ~10–26 % (and even the two close responders 7xadt6 vs 9GMQYQ separate by 2.9 % with |d| ≈ 2), so signal-to-noise for ranking designs is comfortably high. (PR #67's vertex-acrylic set is n = 1 per configuration, so it contributes a between-specimen spread — CV ≈ 9 % across 4 specimens — not a within-specimen CV.)

2. Minimum drops per specimen

Two requirements, both computed from the observed CVs (n = (t·CV/MoE)²; two-design power at 80 %):

  • Precision: at the median output CV (~1.0 %), 5 recorded drops give a ±1.2 % 95 % CI on the per-specimen mean; even a conservative 2.5 % CV reaches ±3 % by n = 6.
  • Discrimination: resolving a 10 % between-design difference needs only 1–2 drops at these CVs — precision, not discrimination, drives the count.
  • Warm-up: fresh-mount burn-in settles by drop ~2 on the key-seat/wax mount (5–10 on hot-glue), so discard 2 before recording.
plan warm-up recorded total when
minimal 0 5 5 wax key-seat mount, T objective
baseline 2 5 7 default SOP
conservative 2 10 12 noisy input channel, hot-glue mount, or a design on a decision boundary

3. Time per set at ~42 s/drop (60 in, auto)

Your 42 s figure is confirmed at campaign scale — the 201-drop 60 in/5 felts session logged a median cadence of 41 s.

plan drops @ 60 in @ 5–10 in (~11–16 s/drop)
minimal (5) 5 3.5 min ~1.3 min
baseline (2+5) 7 4.9 min ~1.9 min
conservative (2+10) 12 8.4 min ~3.2 min

Batch scale at 60 in, baseline plan: 10 designs ≈ 0.8 h, 20 designs ≈ 1.6 h, 48 ≈ 3.9 h, 96 ≈ 7.8 h (conservative plan ≈ 1.7× those).

One caveat carried over from the writeup: these CVs are repeat drops of one physical article — print-to-print reproducibility of the same geometry isn't characterized yet, so plan on ≥3 replicate prints per geometry before treating peak-g/T as a settled BO objective. Full details, math, and figures: docs/drop-test-sample-size-analysis.md.

--- · Branch

github-actions Bot and others added 2 commits July 21, 2026 16:37
Co-authored-by: Jinkwan Han <282969777+ctrhjk@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Refreshes the drops-per-specimen variance/sample-size/timing meta-analysis
with the two datasets referenced in the PR #86 re-ask: the 60 in / 5 felts
validation (7xadt6/9GMQYQ, ~95 stabilized drops each) and the 5-vs-10-in
comparison (30 drops/height). Pooled output CV 0.31-3.5% (median ~1.0%);
T median CV tightens to ~1.2% (0.12-0.45% at the 60 in / 5 felts operating
point). Recommendation unchanged: 2 warm-up + 5 recorded = 7 drops/specimen.

Co-authored-by: Marcus Madsen <265197858+me-madsen@users.noreply.github.com>

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ctrhjk

ctrhjk commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

@claude

These are the slo-mo videos of the two specimens' drop tests.

7xadt6:
https://youtube.com/shorts/Nab3hfuF4Dw?si=f_s9pycnNpNioYzx

9GMQYQ:
https://youtube.com/shorts/zkum2JlHpYk?si=iam6iLkiIMK11_Mc

Can you do a video analysis?

@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @ctrhjk's task in 8m 56s —— View job


Slow-mo video analysis (7xadt6 & 9GMQYQ) — done, with one limitation ⚠️

  • Attempt full video download on the runner — blocked: YouTube gates both the player API and the watch page behind a sign-in/bot check for datacenter IPs (tried yt-dlp with 8 player clients, Invidious mirrors, and the embedded player)
  • Fall back to the real video frames YouTube serves publicly — 8 frames total: ~25/50/75 % of each runtime (1280×720) + 1080×1920 poster per video
  • Visual + quantitative frame analysis (crops, orange-strut pixel masking, edge detection, inter-frame diffs)
  • Committed frames + writeup (data/drop-tests/60in-5felts-validation/video/, §7 added to docs/drop-test-60in-5felts-analysis.md) — pushed (c2dd07f)

TL;DR

I couldn't pull the full videos from CI (YouTube bot-gate — details below), but the 8 real frames YouTube serves publicly were enough for a meaningful check. Everything visible is consistent with the accelerometer record: correct bungee-free setup, specimens intact in every frame, and the felt shows the expected impact-zone wear. For true frame-by-frame kinematics I need the original files — one small ask at the bottom.

What the frames show

Setup verification (both videos) — matches the campaign SOP exactly:

  • T3-prism specimen upright on the acrylic plate atop the grey drop carriage, carriage riding two vertical guide rails, brown-topped felt stack below (lighter sheets visible beneath the top sheet, consistent with the 5-sheet stack).
  • Tri-axial accelerometer seated at the top vertex in the key-seat mount (grey wax visible), white cable routed with generous slack — no visible cable load on the specimen. No bungees on the carriage — the bungee-removed SOP is confirmed on camera.
  • 7xadt6 close-up additionally shows the specimen tie-down (dark tether from a bottom node to a plate anchor) and a silver puck on the plate consistent with the stud-mounted single-axis input accelerometer (CH5).

Quantitative frame comparison (9GMQYQ, 50 % vs 75 % frames):

  • Carriage bottom edge at the identical pixel row (y = 444) in both frames, and the specimen's strut bounding box matches to ≤1 px — these frames catch a parked hold above the felt with the specimen sitting perfectly still. The only inter-frame change localizes to the top-vertex accelerometer/cable region (sub-strut-scale wobble).
  • Specimen integrity: struts straight, tendons taut, geometry upright in every frame of both videos — visual confirmation of what the CVs already implied (both specimens survived ~100 drops each with ≤1.7 % output CV).
  • The top felt sheet shows a mottled darker zone at the impact center — consistent with §3's cumulative-compaction finding (CH5 raw spike 2.1 → 6.5 kG over the evening), though lighting can't be fully excluded from a still.

Nothing in the frames contradicts the accelerometer analysis; the setup-level failure modes from earlier iterations (bungee lift-off, hot-glue mount, missing tie-down) are all visibly absent.

Why not full video analysis, and how to enable it

YouTube refuses stream access from GitHub's datacenter IPs ("Sign in to confirm you're not a bot") — every yt-dlp player client, Invidious mirrors, and the watch/embed pages are all gated. The preview frames above are the only real frames served publicly. To unlock frame-by-frame work (impact deformation, rebound kinematics, contact time vs the DAQ's ~1 ms pulse, drop-height verification):

  1. Attach the original video files directly to a PR comment (GitHub accepts .mp4 up to 10 MB — a trimmed impact clip is fine), or commit them / add to a release.
  2. Tell me the recording frame rate (e.g. 240 fps) and playback speed so video time maps to real time.
  3. Optional: something of known size in the focal plane (ruler, plate width) for pixel-to-mm calibration.

Post those and ping me — I'll do displacement/velocity tracking through the impact.

--- · Branch

@me-madsen

Copy link
Copy Markdown
Collaborator

Minimum: 5 recorded drops per specimen, after 2 discarded warm-up drops → 7 total ≈ 4.9 min/specimen at 60 in

Good to note that there seems to be relatively low variance in our testing sessions per specimen. We will currently plan to do 50 drops per specimen as a baseline for our testing, and may increase this up to 100 in the future; simply to have a relatively large dataset for our research.

github-actions Bot and others added 2 commits July 21, 2026 17:24
YouTube blocks full downloads from CI (sign-in/bot gate), so this commits
the eight real preview frames YouTube serves publicly (25/50/75% + poster
per video) with frame-level observations: bungee-free rig verified,
key-seat mount + tie-down visible, specimen intact in every frame, felt
impact-zone mottling consistent with the compaction finding, and carriage/
specimen parked to <=1 px between the two mid-run 9GMQYQ frames.

Co-authored-by: Jinkwan Han <282969777+ctrhjk@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sgbaird

sgbaird commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

@claude threw those videos into the branch, top-level dir. Rename and organize as you see fit. Not sure what the fps, but I'm certain that this is available somewhere else in this repo, perhaps in one of the earlier PRs. You may need to do an extensive sweep through all PR comments or do a repository-level search or equivalent to find it.

@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @sgbaird's task in 30m 8s —— View job


Videos organized + full frame-by-frame kinematics done ✅


TL;DR

The committed files unlock exactly what the preview-frame pass couldn't do, and the headline result is that the video independently corroborates the accelerometer record: the impact pulse is ≤2 capture frames ≈ 1–2 ms (DAQ says ~1.6 ms), and the mystery "parked carriage" from the preview frames turns out to be an anti-rebound brake catch — no secondary impacts ever reach the specimen.

Time base — solved exactly, no fps guesswork left

The uploads are YouTube's 30 fps re-encodes of the camera's 24p/960 fps HFR clips. Instead of assuming a mapping (the burn-in-wax pass had to carry a ±25 % caveat here), the script detects and removes the 24→30 pulldown duplicate frames: 19.8 % duplicates in 7xadt6 (textbook 1-in-5) and 14.1 % in 9GMQYQ. The surviving frames are the camera's capture frames, so real time = unique frame / 960 exactly. Sanity check: 22.57 s of playback ÷ 40× = 0.564 s real ≈ the 0.557 s free-fall time from 60 in.

Kinematics (one drop per specimen)

7xadt6 9GMQYQ
deceleration bracket ≤2 frames ≈ 1–2 ms ≤2 frames ≈ 1–2 ms
top-vertex snap-back / impact speed 0.70 0.68
sustained rebound / impact speed 0.35 0.43
brake deceleration 2.1 g 2.4 g
brake catch after impact +89 ms, 130 mm rise +86 ms, 150 mm rise
  • Rig behavior finding: after impact the carriage rebounds at ~0.4× impact speed and is braked at ~2.1–2.4 g to a dead stop 130–150 mm above the felt, holding there — each capture is a single clean shock. (This is why the preview frames showed it "parked above the felt".)
  • Specimen response: the top vertex snaps back at ~0.7× impact speed on both specimens, and the 7xadt6 montage visibly shows the struts bowed through contact/turnaround/+15 ms and straight again at the hold — elastic flexure, fully recovered. Intact in every inspected frame, matching the CV ≤ 1.7 % accelerometer story.
  • Calibration honesty: there's no scale bar and the visible descent is too short to self-calibrate from g (the pixel velocity is flat to ±2 % — perspective gradient cancels the free-fall gain), so the pixel scale is anchored on free-fall arrival from 60 in (5.47 m/s; the DAQ plate Δv of 5.53/5.69 m/s corroborates). Nice cross-check: the two independently framed videos then imply the same physical specimen size (82 vs 78 mm orange-strut extent) — that consistency would break if the 960 fps time base or the 60 in height were wrong.
  • Limit: peak specimen compression happens between capture frames (inside the 1–2 ms pulse), so 960 fps can't resolve it — the Edison synthesis's ≥5000 fps DIC recommendation stands for deformation measurement.

Figures: velocity/trajectory · impact zoom · montages: 7xadt6, 9GMQYQ

The legacy preview JPGs stay in video/ for the record; their observations are folded into §7.

--- · Branch

github-actions Bot and others added 4 commits July 21, 2026 20:32
Move @sgbaird's root-level video uploads to
data/drop-tests/60in-5felts-validation/video/{7xadt6,9GMQYQ}_slomo.mp4 and
run the frame-by-frame pass that was previously blocked by YouTube's bot
gate. fps recovered from the repo record + PR #67 camera spec: Sony RX100
IV HFR at 960 fps; the script deduplicates the 24->30 pulldown so real time
is exactly unique-frame/960 (19.8%/14.1% duplicates found).

Findings: impact pulse <= 2 capture frames (~1-2 ms) corroborating the
DAQ's ~1.6 ms CFC width; anti-rebound brake catch at ~2.1-2.4 g holding the
carriage 130-150 mm above the felt from ~86-89 ms after impact (no
secondary hits); top-vertex elastic snap-back at ~0.7x impact speed with
visible strut flexure fully recovered; cross-framing scale consistency
(82 vs 78 mm specimen extent) under the free-fall anchor (DAQ dv 5.53/5.69
m/s corroborates). Compression itself falls between frames - >=5000 fps DIC
still needed.

Co-authored-by: Sterling G. Baird <45469701+sgbaird@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@me-madsen

Copy link
Copy Markdown
Collaborator

Given these results the tower should be fine to resume regular testing. Depending on the results of this inquiry we may be changing how we run this campaign to collect data soon #97 (comment)

@me-madsen

Copy link
Copy Markdown
Collaborator

Note the following conversation regarding beginning testing.

image

@me-madsen

Copy link
Copy Markdown
Collaborator

@claude search the repository to see if anything is recorded on the accelerometer calibration settings that we set on the drop tower computer. It's likely that if anyone recorded this it would be @ctrhjk

@claude

claude Bot commented Aug 18, 2026

Copy link
Copy Markdown

Claude finished @me-madsen's task in 1m 51s —— View job


Accelerometer calibration settings — what's recorded in the repo ✅

  • Sweep the repo (READMEs, docs, protocol, raw TP4 export headers, series tables) for calibration/sensitivity settings
  • Check what @ctrhjk recorded in the dataset READMEs and original PR comments
  • Report what exists, what's missing, and where it lives

Yes — the TP4 per-channel calibration settings are recorded, in many places, and your instinct is right: the source is @ctrhjk. The fullest record traces back to @ctrhjk's channel-map post on PR #67, which was committed into data/drop-tests/vertex-acrylic/README.md and then propagated (in abbreviated form) into ~20 dataset READMEs.

The most complete record — vertex-acrylic/README.md (from @ctrhjk, tests of 06/22)

This is the only place that captures the full TP4 channel configuration, including coupling and ICP:

Channel Sensor Full scale (G) Sensitivity (mV/G) Coupling ICP Waveform analysis
CH2 tri-axis (X) 14,492.8 0.69 AC yes Half Sine
CH3 tri-axis (Y) 14,992.5 0.667 AC yes Half Sine
CH4 tri-axis (Z) 13,624.0 0.734 AC yes Half Sine
CH5 single-axis 9,442.9 1.059 AC yes Half Sine

The axis assignment (CH2 = X, CH3 = Y, CH4 = Z) was stated by @ctrhjk in the 60 in/5 felts comment and is recorded in the READMEs of every recent dataset (abc123-blind, 60in-5felts-validation, print-defects, pu-configs, etc.).

The retired low-range tri-axis (bottom-vertex "BOT", CH6–CH8) is recorded too, e.g. data/drop-tests/100drops/README.md: 1,002.0 / 991.1 / 989.1 G at 9.98 / 10.09 / 10.11 mV/G.

These sensitivities have been identical in every dataset from June through the current campaigns — nobody has changed the per-channel calibration entries on the TP4 since @ctrhjk recorded them.

Capture-setting history (also recorded, per dataset README)

era record / rate pre-trigger trigger
June–early July (125 kHz era) 200 ms / 125 kHz / 25,000 samples 2 % = 4 ms 1000 G (CH4, later CH5)
late July (PU sweeps) 20 ms / 1.25 MHz ~0.39 ms 300 G, then 150 G
current SOP (since abc123 blind test) 100 ms / 1.25 MHz / 125,000 samples 2 % = 2 ms 150 G

The current SOP is pinned in docs/drop-test-ab-blind-protocol.md and data/drop-tests/abc123-blind/README.md.

What is not recorded anywhere — the actual gaps

  1. Sensor make/model/serial numbers and factory calibration certificates. The protocol doc says only "Accelerometer Get started on this repo #1 (PCB / similar)" (docs/drop-test-protocol.md:23). So we know the numbers entered into the TP4, but not whether the sensitivities came from each sensor's factory cal sheet or were typed from something else — and there's no way to check drift against a certificate we don't have.
  2. Where the full-scale values come from. They're consistent with the TP4's ±10 V input divided by each sensitivity (e.g. 10 V ÷ 1.059 mV/G = 9,443 G), i.e. derived ceilings, not independent settings — but that derivation isn't confirmed in any committed record.
  3. Mechanical shock-survival ratings of the sensors (separate from measurement range) — flagged as missing back in the accelerometer-ranges answer, still missing.
  4. No screenshot of the TP4 Channel Settings page is committed — everything we have is transcribed.

Suggested 5-minute close-out (for @ctrhjk, next time at the tower)

Photograph (a) the TP4 Channel Settings screen and (b) the model/serial labels on the three sensors — plus the calibration sheets if they're in the case — and post them here. I'll fold them into the protocol doc as the canonical calibration record. That would close gaps 1–4 in one pass and make the sensitivity entries auditable against the factory values.

No code or docs were changed — this was a search/report task.
· Branch

@me-madsen

Copy link
Copy Markdown
Collaborator

@claude note that these are what I believe are the right settings for future reference. The settings were cleared and so I want to ensure the settings are giving the right sensitivities and so on. Compare the following data of the same specimen before the reset to see if the sensitivities are right:

Before clearing of setting (101 drops): https://byu.box.com/s/it5499hkyw24twg7179smsn0fv0bodal
After (30 drops): https://byu.box.com/s/4tttcvt6lx008tr05meruslkwasqqskh

20260818_172017

@claude

claude Bot commented Aug 19, 2026

Copy link
Copy Markdown

Claude finished @me-madsen's task in 13m 55s —— View job


Post-reset sensitivity verification — the re-entered settings check out ✅

  • Read the settings screenshot + gather context (canonical channel map, prior pipeline)
  • Download both Box datasets — before: 8-17-2026 bpx68c - 60 in - 0.5 in mat - 101 drops, after: 8-18-2026 bpx68c - calibration testing (30 drops, captures time-stamped 08-19) — 131 captures, ~1.2 GB
  • Check settings vs the historical record (screenshot transcription + in-data capture format)
  • Data-level before/after comparison — scripts/analysis/drop_test_calibration_check_analysis.py
  • Write up — docs/drop-test-calibration-check-analysis.md + figures/metrics + screenshot committed as the canonical record — pushed (bf274dd)

TL;DR

The re-entered settings are right — resume testing on them as-is. The screenshot matches the June-era channel map digit-for-digit, and the data confirms the entries took effect: every observable that a wrong sensitivity must move is continuous across the reset to ≲1 %, an order of magnitude below the smallest plausible typo (3.4 %). The one thing that did change (+9–11 % on the peak levels) has the signature of the rested/re-seated mat, not of a channel gain — a gain error can't narrow a pulse, and the pulse narrowed −8 % while impulse, Δv and hop timing stayed put.

1. Screenshot vs the canonical record — exact match

CH2/CH3/CH4 = 0.690 / 0.667 / 0.734 mV/G (14,492.8 / 14,992.5 / 13,624.0 G), CH5 = 1.059 mV/G (9,442.9 G) with the 150 G trigger, all AC + ICP + Half Sine; 100 ms / 1.25 MHz / 125,000 samples / 2 % = 2 ms pre-trigger. Identical to @ctrhjk's channel map recorded since June (vertex-acrylic/README.md) and the current-SOP capture settings. CH1/6/7/8 correctly inactive. The screenshot is now committed at data/drop-tests/calibration-check/tp4-settings-2026-08-18.jpg and linked from the protocol doc §1 as the canonical settings record.

2. The data-level check

A mis-entered sensitivity is a per-channel multiplicative factor: it must scale that channel's peak, Δv and noise floor by the same amount while leaving all timing/shape metrics untouched. So the test is scale-coherence, not raw level matching:

observable before (08-17) after (08-19) ratio reads as
trigger crossing (raw CH5 > 150 G) 2.06 ms 2.05 ms trigger + 2 ms pre-trigger identical
input Δv 3.497 m/s 3.501 m/s 1.001 (p = 0.84) CH5 gain unchanged
e_rebound (scales as 1/gain) 0.0311 0.0308 0.990 CH5 gain unchanged
T = out/in 1.021 1.012 0.991 tri-axis vs CH5 gain unchanged
hop delay t_second 22.0 ms 22.0 ms 1.01 same arrival velocity — sessions comparable
CH5 peak / pulse width 199 G / 2.48 ms 220 G / 2.27 ms 1.106 / 0.917 pulse reshaped; peak × width ∝ impulse moves only +1.4 %
noise floors CH2–CH5 0.46–0.70 down while peaks went up — impossible under a gain increase

Candidate typos, each excluded: a channel left at the 10 mV/G default (×0.07–0.10 — nowhere close), CH5 entered as 1.000 (×1.059 on Δv and e_rebound — both flat), a Z-axis value swap (would drag T by 6–10 % — T moved 0.3 %). The one candidate below the check's noise floor is a CH2↔CH3 value swap (±3.4 %, opposite signs on the two small lateral axes): the observed lateral shifts are same-sign (a physics pattern, not a swap pattern), and even if it existed it would move the campaign metric (the resultant, dominated 95 % by CH4/Z) by ≲0.3 %. Figures: continuity across the reset · per-channel ratios vs candidates.

Bonus continuity: the 08-17 session's input (197.5 G, 2.48 ms) lands right on the 08-11/12 speed-decay record at this arrangement (195–201 G, 2.5–2.6 ms), so the chain runs unbroken from the pre-reset era through the reset.

3. What the visible +10 % shift actually is

The after session hits harder and shorter at the same arrival energy — rested mat (2 days idle after 101 drops) plus whatever re-seating happened during the settings work. The mirror image is inside the before session: its Δv slid 3.96 → 3.11 m/s over 101 drops as the mat's rebound faded, the exact within-session mat drift the speed-decay campaign documented. Practical corollary: after any mat rest/re-seat, budget ~5–10 warm-up drops (the after session takes that long to reach its plateau, a bit more than the usual 2).

Two small things back

  1. The status bar in your screenshot shows the DB sitting on a session named ebdna8 - 60 in - 1/2" mat - 101 drops — a specimen ID that appears nowhere in the repo. If that's a completed 101-drop campaign, it hasn't been uploaded yet.
  2. The definitive close-out for calibration remains the same 5-minute task as before: photos of the sensor model/serial labels + factory calibration sheets, so the 0.690/0.667/0.734/1.059 values can be audited against their source. This check verifies the entries are the same as they've always been and internally consistent — the factory sheets would verify they were right in the first place.
    · Branch

…ere it matters; commit the settings screenshot as the canonical record

Co-authored-by: Marcus Madsen <265197858+me-madsen@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@me-madsen

me-madsen commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

With the complete SOBOL +S0 batch, the campaign has officially begun. 4/9 of our SOBOL + S0 batch have been tested at the following conditions:

  • 60 inch drop
  • 1/2 inch polyurethane mat (exact material and specs here)
  • 101 drops per specimen
  • 3 slow motion videos taken per specimen

The data will be analyzed for average transmissibility per specimen, and any other performance measurements that @claude can extract from our data as has already been set as a precedent in this PR, #94, and #97, and anywhere else as discussed in this repository.

The data is being uploaded to box. It will then be posted here and analyzed. The BO will most likely be handled in a separate thread. The initial transmissibility analysis may also be handled in that separate PR. @sgbaird will handle implementation of the BO at this time. It takes roughly 90 minutes to run each specimen through the testing. We can expect to have this data at some point tomorrow.

@me-madsen me-madsen mentioned this pull request Aug 19, 2026
@claude

claude Bot commented Aug 19, 2026

Copy link
Copy Markdown

Claude finished @me-madsen's task in 6m 45s —— View job


Campaign kickoff — batch pipeline ready for the data ✅

  • Gather context (per-capture pipeline, campaign SOP state, BO hand-off needs)
  • Prepare a reusable campaign analysis script, smoke-tested on real Box captures
  • Commit + push (8703400)
  • Readiness checklist + metadata asks below

TL;DR

Nothing to analyze yet (data lands tomorrow), so I used the window to make the analysis a one-command turnaround: scripts/analysis/drop_test_campaign_analysis.py is a generalized batch pipeline — point it at a folder with one subfolder per specimen (loose CSVs or the TP4 zips as uploaded, both read directly) and it produces the full per-specimen workup plus a BO-ready campaign_summary.csv for @sgbaird's thread. Smoke-tested end-to-end on real captures pulled from Box via the committed speed-decay manifests. Staging folder + usage: data/drop-tests/sobol-campaign/.

What gets extracted per specimen (the full precedent stack)

  • Transmissibility T = TOP/CH5 under both CFC-180 and CFC-1000 (the scripts have emitted both since the Edison review), stabilized mean ± CV after the 2-drop warm-up discard, plus OLS drift slopes
  • Output/input CFC peaks, raw peaks + worst %-of-full-scale per channel, pulse FWHM
  • Δv with an automatic rig-health verdict per session — healthy ≥ 5.28 m/s / settled 4.55–4.66 / alarm < 4.40 — so a rail-friction relapse mid-campaign gets flagged before it contaminates a specimen's numbers
  • t_second and e_rebound (the specimen-hop constant — in the blind test these discriminated specimens 3–10× more strongly than T)
  • Ringdown f_n / ζ, r²-gated per the abc123 convention (low-r² drops honestly excluded rather than averaged in)
  • Invalid-capture screening (raw input below the 150 G trigger, like abc123's Signal 12), pause/cadence log
  • Campaign level: cross-specimen ANOVA + pairwise Welch/effect sizes, ranking figure, and the discrimination check against the known noise floors (print-to-print ~2 % spread in T; cross-session mount re-seat ~1.5–2 %)

The one thing the BO hand-off still needs: the specimen-ID → design-parameter map

This has been an open gap since the first input-output run ("IDs not yet tied back to design parameters"). The script takes --params params.json ({"<id>": {"<param>": value, ...}}, IDs case-insensitive) and passes the parameters straight through into campaign_summary.csv, so the BO ingests (parameters, objectives) rows directly. If the SOBOL batch's design table exists anywhere (CSV/JSON/spreadsheet), posting it alongside the Box link makes the hand-off automatic.

Asks for the upload (all cheap, all bite later if missing)

  1. One Box folder per specimen, folder/session name starting with the specimen ID (e.g. bpx68c 60in 101drops) — the pipeline parses the ID from the first word, lowercased.
  2. Session dates/order + any mount re-seats. With 9 specimens across multiple days, the ~1.5–2 % session-to-session T shift from re-seating is the same order as some between-design gaps — the sequencing log is what keeps that attributable. Worth considering: re-drop one reference specimen (e.g. bpx68c, already characterized twice) for 5–10 drops every couple of sessions as a free session-normalizer.
  3. Which specimen is S0, so the baseline is labeled in the outputs.
  4. Videos: camera-original landscape files + XML sidecars on Box (your current practice) — the sidecars carry the exact 959.04 fps time base.

One connection worth confirming: the 08-18 settings screenshot showed the TP4 database sitting on a session named ebdna8 - 60 in - 1/2" mat - 101 drops — a specimen ID not seen in the repo before. If that's one of the 4/9 already tested, it's presumably in the upload; if it predates the settings reset, flag it, since the calibration check only verified continuity for post-reset data (the 08-17 bpx68c session showed pre-reset data is fine too, but it's worth knowing which side of the reset each specimen sits on).

When the Box links are posted, ping me and the turnaround is: fetch → run → per-specimen writeup + campaign ranking + campaign_summary.csv for the BO thread.


· Branch

…pload

One command per Box upload: per-specimen stabilized metrics (T in both
CFC bands, dv rig-health verdict, t_second, e_rebound, r2-gated ringdown
fn/zeta), invalid-capture and warm-up handling, cross-specimen
ANOVA/pairwise stats, figures, and a BO-ready campaign_summary.csv with
optional specimen-id -> design-parameter passthrough. Smoke-tested on
Box captures from the committed speed-decay manifests (loose CSVs and
zip archives both exercised).

Co-authored-by: Marcus Madsen <265197858+me-madsen@users.noreply.github.com>

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@me-madsen

Copy link
Copy Markdown
Collaborator

@claude, @sgbaird, I've uploaded drop data for 8/9 specimens. Slo-mo video data will take me a little more time to upload. The last specimen will be tested tomorrow morning. Do your thing to analyze it @claude as you mentioned in the last message.

bpx68c - https://byu.box.com/s/it5499hkyw24twg7179smsn0fv0bodal
amdjwm - 87 drops - https://byu.box.com/s/ig3vgyld9t4kbuf2tstvg8v3m2jrwns0 / 101 drops - https://byu.box.com/s/xd2ftvi3y57jopvxyg1pae7qx61jdsvu
bag26v - https://byu.box.com/s/n5fkbur86gzronh04rf3f00diw0yz2es
autv5r - https://byu.box.com/s/2uztrfbblzzwhefmdelt523v4xi1exz0
6lhxfy - 35 drops - https://byu.box.com/s/hkyw0nv9s2l27r893sapb6kxumk0l0d9 / 101 drops - https://byu.box.com/s/q5tyg1as1h0pgqrppa8nsuhnllbhjsnu
9hhbkp - https://byu.box.com/s/bbwutg2r6a4l3q71vv2q84nvrue7h1h6
nvxsrv - https://byu.box.com/s/cym881rlx1je5rec4eowj8khglopax2i
6nheas - https://byu.box.com/s/y7v10k08mqb3f093dxfl4pn8wafq6n9i

The appropriate information, including digital file and pictures of each, and mass, can be found herehttps://github.com//issues/98#issuecomment-5363923933 and in the associated issue.

@claude

claude Bot commented Aug 21, 2026

Copy link
Copy Markdown

Claude finished @me-madsen's task in 23m 3s —— View job


SOBOL + S0 campaign analysis — 8/9 specimens, 942 drops ✅

  • Gather context (design-parameter/mass key from the issue-T-3_01 Prints #98 print-key job; campaign pipeline)
  • Download all 10 Box datasets (8 specimens + 2 partial sessions, 942 captures ≈ 8.4 GB) — all 942 captures clean, zero spurious triggers
  • Catch + fix a data-quality issue that would have corrupted the BO objective (pre-trigger contact foot, §below), validated against the TP4's independent series-table numbers
  • Run the campaign pipeline: per-specimen metrics, ranking, session-repeatability checks
  • BO-ready campaign_summary.csv with design parameters joined
  • Findings doc — docs/drop-test-sobol-campaign-analysis.md — + figures, READMEs, Box manifests; committed and pushed (80d42b1)

TL;DR

The campaign discriminates decisively, and there's a headline: 6lhxfy (Sobol spec 01) is the first strong attenuator this program has ever measured — T = 0.893, an 11 % peak reduction, reproduced across two sessions on different days to 0.13 %. T spans 0.893–1.062 (16.8 %) across the 8 specimens with within-specimen CV ≤ 0.5 %. Also: the tower quietly finished recovering (Δv 5.3–5.5 m/s ≈ healthy bar all campaign week), and one pipeline fix was needed before any of this was trustworthy.

The ranking (T = TOP/CH5, CFC-180; stabilized drops)

rank specimen spec T180 (CV) e_rebound Δv health
1 6lhxfy 01 0.893 (0.47 %) 0.050 healthy
2 amdjwm ? 0.980 (0.26 %) —† settled
3 6nheas 05 0.997 (0.32 %) 0.040 settled
4 bpx68c S0 ref 1.011 (0.23 %) 0.020 healthy
5 9hhbkp 00 1.018 (0.17 %) 0.022 healthy
6 nvxsrv 04 1.027 (0.43 %) 0.027 healthy
7 autv5r 02 1.040 (0.34 %) 0.027 healthy
8 bag26v 08 1.062 (0.48 %) 0.024 settled

amdjwm's hop detector was bimodal this session; its e_rebound isn't reliable.

Every adjacent pair separates at |d| ≥ 2.8, but against the known ~2 % print-to-print floor the honest read is four design tiers: 6lhxfy alone (9.8 % clear of the field) → amdjwm/6nheas → the near-unity middle (bpx68c/9hhbkp/nvxsrv, mutually unresolved at print level) → the amplifiers (autv5r, bag26v, the tallest article). 6lhxfy's corner: thickest struts (9.24 mm), thinnest cables (2.55 mm), highest twist (77.4°), lightest mass — the compliant-cable/stiff-strut corner is the live region for the BO. Nice sanity check: spec 00 differs from S0 mainly in cable Ø and lands 0.7 % from it — within print noise. Figures: per-drop series · ranking boxplots

⚠️ The fix that had to happen first: the pre-trigger is no longer a valid baseline

The carriage now arrives fast enough (~5.4 m/s) that mat contact begins > 2 ms before the 150 G trigger — the entire 2 ms pre-trigger window rides a +10–20 G contact foot. Baselining on it (the abc123-era convention) made Δv read 2.0–3.5 m/s (nonsense) and — much worse — biased T by up to ~5 % in a specimen-dependent way (nvxsrv 1.074 → 1.024 after the fix). The pipeline now re-baselines on the record tail; validated against the TP4 series tables' independently computed raw peaks (1–2 % agreement) and Δv (2–3 %) on all 10 sessions. The old mode is untouched, so prior committed analyses reproduce unchanged.

Corollary that corrects the standing record: the same artifact (smaller) was in the 08-10/08-12 rail sessions. On the TP4's own numbers the tower was at 4.9–5.0 m/s then (not ~4.6), kept improving with traffic, and by campaign week sat at 5.3–5.5 m/s ≈ the healthy-tower bar — the pin-break deficit is effectively gone without further intervention. The within-session "velocity decay" was largely the foot growing as the mat warms. Correction notes added to the speed-decay and greasing docs; the qualitative conclusions there (real greasing step, no cumulative decay) survive.

Session repeatability — the interrupted sessions were free crossover checks

6lhxfy 0.8943 (35 drops, 08-19) vs 0.8931 (101 drops, 08-20): 0.13 % apart across a day + mount re-seat. amdjwm 0.9862 → 0.9805 same-day (0.58 %). T180 transfers across sessions far better than the historical worst cases. T1000 does not transfer (amdjwm 1.16 → 1.01) — it's mount-coupling-dominated; keep CFC-180 T as the objective and T1000 as a diagnostic.

BO hand-off + questions back

campaign_summary.csv has one row per specimen: objectives ± sd (t180, t1000, out/in peaks, Δv, t_second, e_rebound, ringdown fn/ζ) + design parameters (R, H, twist, strut/cable Ø, mass, spec) from params.json. Recommended objective: t180_meant180_sd), e_rebound as a secondary if multi-output. Raw data stays on Box; committed per session: TP4 series table + Box-ID manifest (re-fetchable via the new scripts/fetch_box_shared_folder.py).

  1. @me-madsen: which design is amdjwm? It's the 2nd-best specimen (T = 0.980) but appears in no print key, Sobol table, issue, or PR anywhere — its parameter row is blank, so the BO can't use it until it's mapped (re-print of spec 03/06/07? re-label?). Its mass would be good to have too.
  2. ebdna8 (spec 03): the 08-18 settings screenshot showed the TP4 database sitting on a completed ebdna8 - 60 in - 1/2" mat - 101 drops session that has never been uploaded. If real, uploading it completes a 9th specimen for free.
  3. Spec 08 official-article ambiguity from the print key (dea4ls per .3mf vs bag26v per issue comment): bag26v is what was tested here.
  4. Capture settings for the remaining specimen(s): if the TP4 allows, a 5 ms (5 %) pre-trigger would put the whole contact foot in-record and retire the baseline issue at the source. Everything else can stay exactly as is.

When the 9th specimen + videos land, the turnaround is one command — ping me and I'll fold them in and refresh the ranking + summary CSV.
· Branch

… fix the pre-trigger contact-foot baseline, confirm the tower healthy via the TP4 series tables, rank the designs (6lhxfy T=0.893 is the first strong attenuator), and emit the BO summary

Co-authored-by: Marcus Madsen <265197858+me-madsen@users.noreply.github.com>

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sgbaird

sgbaird commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

@claude the mass seems to be varying more than I would have expected. #29 #30 #35 (especially latter) has the details around how that would have been generated I think. Unless the constant mass constraint wasn't implemented there

@claude

claude Bot commented Aug 21, 2026

Copy link
Copy Markdown

Claude finished @sgbaird's task in 3m 51s —— View job


Mass variation explained — the constraint was implemented, but it holds solid volume constant, not printed grams ✅


TL;DR

The constant-mass constraint was implemented and converged — the batch CSV predicts 30.90–30.97 g for all 9 designs (±0.1 %). The measured spread (18.50–22.04 g, CV 5.9 %) comes from a definitional gap: the constraint equalizes solid CAD volume, but the printer gives PLA ~57 % effective density (walls + sparse infill) while the thin TPU cables print ~solid. Since the PLA/TPU split varies strongly by design (TPU fraction 12–36 %), a constant-solid-mass batch deterministically becomes an ~±9 % printed-mass batch. Regression on the generator's own per-material predictions explains R² = 0.78 of the variation; the residual ±0.3–0.8 g (~3 %) is ordinary print-to-print variance.

1. The constraint exists and did its job

You remembered right: bo/t3_prism_sobol_batch.py (PR #35, adapting PR #30's campaign script per PR #33's hybrid-campaign Routes) implements Route A — constant cell mass: every Sobol design is uniformly re-scaled (twist and shape ratios preserved) until its estimated mass equals m* = 30.95 g, iterating on rendered STL volumes to |m − m*| ≤ 0.15 g. m* is the solid-volume mass of the S0 reference STLs at ρ_PLA = 1.24 / ρ_TPU = 1.21 g/cm³. The committed bo/t3-prism-bo-batch.csv shows all 9 rows at mass_g 30.90–30.97, mass_ok=True. So the projection worked — in solid-volume space.

2. Where the grams went

The CSV also predicts the per-material split (pla_g / tpu_g), and that's the tell. Regressing the 7 measured masses (issue #98) on the predicted solid masses:

measured ≈ 0.565 · m_PLA,solid + 0.986 · m_TPU,solid (R² = 0.78, residual sd 0.64 g)

specimen spec PLA / TPU solid (g) TPU frac measured fit resid
6lhxfy 01 27.33 / 3.64 0.12 18.50 19.02 −0.52
bpx68c S0 23.42 / 7.53 0.24 20.23 20.65 −0.42
nvxsrv 04 24.04 / 6.88 0.22 20.66 20.36 +0.30
bag26v 08 23.45 / 7.50 0.24 21.42 20.63 +0.79
9hhbkp 00 20.04 / 10.89 0.35 21.62 22.05 −0.43
6nheas 05 19.84 / 11.11 0.36 21.73 22.16 −0.43
autv5r 02 21.74 / 9.19 0.30 22.04 21.34 +0.70

Reading it: the slicer prints PLA at ~57 % of solid and TPU at ~99 % (thin cables are all walls — no room for infill). The mass projection preserves each design's shape, so the thick-strut/thin-cable corner (6lhxfy) is PLA-heavy and loses the most grams to infill, while TPU-rich designs (9hhbkp, 6nheas) keep theirs. The measured/solid ratio correlates with TPU fraction at r = 0.88. The leftover scatter (~3 %) matches the print-to-print variance the defect study measured — nothing anomalous there.

3. What to do about it (if anything)

For this batch, nothing is lost: the actual masses are recorded in params.json/campaign_summary.csv, so the BO can carry mass as a feature or normalize by it. But note the confound: printed mass is now correlated with cable_d/TPU share by construction, so "mass-controlled" comparisons across this batch aren't.

For future batches, two one-file fixes in bo/t3_prism_sobol_batch.py (PR #35 branch):

  1. Cheap: weight the scale solve's per-material volumes by effective print densities (the fitted 0.565 / 0.986). One-line change to the mass model — but the coefficients are print-profile-dependent (wall count, infill %), so they'd need a refit if the profile changes.
  2. Exact: iterate the Route-A solve on the BambuStudio CLI's sliced per-filament grams instead of STL volumes — the --assemble path is already wired in, and sliced grams are the ground truth for the active profile. Slower per iteration, exact by construction.

Which is "right" depends on the constraint's purpose: if the point is attenuation-per-gram-of-printed-structure at fixed mass budget, printed mass is the physical quantity to hold constant, and option 2 is the honest version.

4. A useful side effect: the fit can help identify amdjwm

The regression predicts printed masses for the untested specs: 03 (ebdna8) ≈ 20.4 g, 06 ≈ 19.2 g, 07 ≈ 21.5 g. So weighing amdjwm (the mystery 2nd-place specimen with no parameter row) would help disambiguate it — a ~19 g reading points at spec 06, ~21.5 g at spec 07. Cheap tiebreaker alongside a caliper check of R/H.

Recorded as §7 of docs/drop-test-sobol-campaign-analysis.md (642b8c0).
· Branch

…s constraint held solid volume constant (30.90-30.97 g predicted), but PLA prints at ~57% effective density vs ~99% for TPU, so the design-dependent PLA/TPU split maps it onto the observed +/-9% printed-mass spread (fit R2=0.78)

Co-authored-by: Sterling G. Baird <45469701+sgbaird@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sgbaird

sgbaird commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

Aside: seems like 100% infill of the PLA struts would have been better. Was there not a conversation around this in the repository somewhere about setting the infill to 100%?

claude Bot pushed a commit that referenced this pull request Aug 21, 2026
…suggest round-2 batch

Multi-objective (t180, e_rebound), fully Bayesian, batch, existing-data
Honegumi template adapted for the physical campaign: Sobol init step
dropped (PR #35 batch was the init), results ingested from the PR #86
campaign_summary.csv snapshot, specs 03/06/07 attached as pending
trials, amdjwm skipped until its spec mapping is resolved. Records 9
suggested base-space designs with posterior predictions, the AxClient
state, and a Pareto + parameter-space figure.

Co-authored-by: Sterling G. Baird <45469701+sgbaird@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
claude Bot pushed a commit that referenced this pull request Aug 21, 2026
Task 3e398131: Edison derives its own objective set from the per-drop
campaign data before reading ours, then attacks the three legs of the
hand-off claim (t180+e_rebound objective pair, the n=8 trade-off/Pareto
framing, and the per-drop-SEM noise model that ignores the ~2% print floor).
Bundle: 23 files across three branches (per-drop metrics, series tables,
analysis doc+script, print-defect floor study, #97 energy review, BO script).
Follows the d9092c5a precedent from PR #86.

Co-authored-by: Sterling G. Baird <45469701+sgbaird@users.noreply.github.com>
@me-madsen

Copy link
Copy Markdown
Collaborator

Note this about databases when using the droptower. Databases can fill up quickly due to the high fidelity data we are gathering. This shows how to create new databases and change between them: https://youtu.be/BSK_UcERTVw

claude Bot pushed a commit that referenced this pull request Aug 22, 2026
…33 #94 #97 #85 #86 #98 #101)

Pivot from planned-methods SEA/eta_c framing to the executed campaign:
drop-tower objectives t180 (filtered peak-acceleration ratio) and rebound
energy per drop, SAASBO round 1 on the printed Sobol seed batch (real
results table, real Pareto/feature-importance/LOOCV figures with labels
regenerated for naming consistency), round-2 batch in fabrication,
constant-solid-mass projection + printability screens, as-printed
fabrication record (manual painted supports, TPU dry box, high-flow
nozzle), corrected J211 filter provenance, simulation screening ladder
from PR #33 with honest scorecard, metal-analog metric switched to t180,
Edison adversarial objective review reflected in Discussion/round-3 plan.
SI rewritten: print key, drop-tower protocol/rig characterization,
printed-mass model, simulation ladder. Em-dash sweep per style guide.
Rebuilt all four PDFs.

Co-authored-by: Sterling G. Baird <45469701+sgbaird@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants