21a42cb3e6
The two feature models are frozen and pretrained; only the 100KB head was trained here. The shapes were measured rather than assumed: 2.0s of 16kHz audio gives 197 mel frames, and 76-frame windows at stride 8 give exactly the 16 embeddings the head was fitted on. This file knows tensors and nothing about the 80ms cadence, which is why the scaling openWakeWord applies between the two feature models lives here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN