SOEN Conditional Audio Generator

A requested word/timing plan drives SOEN phenomenological dynamics. A learned decoder predicts a log-mel spectrogram, then Griffin-Lim reconstructs a waveform. These clips are generated by the model, not copied from Speech Commands.

Generated Clip

Generated Waveform

Predicted Log-Mel Spectrogram

SOEN Trace Groups