loner kickvox: draw her real waveforms into the timeline blocks, and the energy trim they exposed master
@jeffrey, reading the study: "up comes too soon" · "curled is too short" · "check the length of the actual waveforms in the utterances, not just ur trim etc" · "please map / render those waveforms directly into the clips in the mp4 so i can see" · "and add live timecode". bin/timeline.py now draws each word's ACTUAL audio inside its block — per-column peak of the vox4/ lead render (the thing the study plays), normalized to the whole take, so dead air inside a slot is visible — plus a live timecode and bar·beat address in the corner, and the active word outlines in pink instead of filling so the waveform stays readable while it plays. Which immediately showed the trim was lying. Whisper's word boundaries are handoffs, not note ends: it gave "led" 0.99 s when she stops singing after ~0.55 and the rest is decay, and build_warp was time-stretching that silence across the slot along with the note — "led" sang to 78% of its block and then sat there. trim_units() pulls each unit's source span back to its real audio end (5 ms RMS against a −36 dB gate of the take's peak, +50 ms release margin) and the silence is DROPPED rather than warped, so only sung frames stretch. Guards: the last unit is untouched (the tail/release machinery owns it), trims under 80 ms aren't worth the surgery, no unit loses more than 65% of its span. On the whole line exactly two fire — led −390 ms · think −445 ms — and every word now sings to ≥92% of its slot (led 78% → 95%). Receipt: `trims` in vox4/.manifest.json. Bar 1 hand-pinned against that picture: CURLED alone fills it (cur 2 + led 2 — her own 0.74/0.99 s split says led ≥ cur), and UP, her 0.25 s tonic release, lands ON the bar-2 downbeat as an up-in pickup pair into "myself", which stays anchored at beat 9 so think/of/stone hold their tuned spots. timeline.py wants system python3 (PIL), not pop/.venv — noted in the README along with the whole study-reading recipe.