isolating regions of human speech from silence or background noise, generating the initial temporal masks