SOEN Conditional Audio Generator
A requested word/timing plan drives SOEN phenomenological dynamics. A learned decoder predicts a log-mel spectrogram, then Griffin-Lim reconstructs a waveform. These clips are generated by the model, not copied from Speech Commands.
Generated Waveform
Predicted Log-Mel Spectrogram
SOEN Trace Groups