macneopolitan: can the words be heard? — an offline singer, a Whisper judge, and the knobs it found master
Jeffrey, hearing the dialogs: "the words are hard to hear / listen to after output." This makes that measurable instead of arguable, without playing a note. `singrender` is a new product in slab/menuband: Menu Band's own MenuBandSinger.swift (symlinked into the target, one source of truth) driven from the shell into WAV files — no audio device, no window, no app. What it writes is what the room would hear. MenuBandSingerVoice moves to its own file so the renderer takes the singer without the engine. `bin/hear.mjs` renders every sung line of a score, sends each WAV back through Whisper, and scores the transcript against the lyric as word error rate. It scores the spoken TTS source too — the ceiling the singing can only fall from. `--compare A B` tables two runs and lists the lines that moved; `--rescore` recomputes every stored run after a fairness fix to the scorer (contractions expanded, numbers spelled, "and" dropped from hundreds); `--env K=V` reaches the core's knobs. Judge is whisper small.en: base.en was too unsteady on sung words to trust. The measurement, 8 dialogs, 421 words: the spoken source scores 3.8% and the singing pass 18.5%. So the words are lost in the singing, not in the speech. live/singer.c grows knobs for that, every one defaulting OFF so today's sound is unchanged: SINGER_SUSTAIN_DB stretches only the loud core of a vowel, so an n or a b murmur is not held into a syllable of its own; SUSTAIN_BAND judges that core by 400 Hz to 4 kHz formant energy; SUSTAIN_MINFRAC and MINMS keep it from freezing a sliver, since a 25 ms zone held 24 times over is one frozen spectrum for a second; GAP_MS stops a held vowel before the next onset consonant so a stop's closure is silence; CMIX blends the spoken original over voiced consonants, not only unvoiced ones; and PRESENCE_DB, XF, LOOP and SINGER_TRACE=1 round it out, the last printing every unit's placement. Results, the failures included, are recorded in hear/ with a README: loop sustain garbles at 127.6% and the formant-band zone is too tight at 24.0%, while the gap and the consonant gain each win about two points. bin/offload.sh moves the whole loop to another Mac, because 8 GB neo wedged under it at load 227: a generated self-contained Swift package for the renderer, whisper.cpp built static in $HOME — brew's binary cannot be relocated without sudo, its ggml compiles in one backend search path and looks nowhere else — and the spoken-source cache, so a host can render lines in a voice it does not have installed. bin/sweep.sh runs a queue of evaluations there, and bin/align-audit.mjs asks whether the synthesizer's word onsets, which the core trusts to slice audio onto notes, are accurate per voice.