Voices Obscured in Complex Enrivonmental Settings (VOiCES)

machine learning automatic speech recognition speaker identification denoising speech processing

Description

VOiCES is a speech corpus recorded in acoustically challenging settings, using distant microphone recording. Speech was recorded in real rooms with various acoustic features (reverb, echo, HVAC systems, outside noise, etc.). Adversarial noise, either television, music, or babble, was concurrently played with clean speech. Data was recorded using multiple microphones strategically placed throughout the room. The corpus includes audio recordings, orthographic transcriptions, and speaker labels.

Update Frequency

Data from two additional rooms will be added to the corpus Fall 2018.

License

Creative Commons BY 4.0 (see here for more details)

Documentation

https://voices18.github.io/

Contact

https://github.com/voices18/utilities/issues

Usage Examples

Resources on AWS

  • Description
    wav audio files, orthographic transcriptions, and speaker ID
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::lab41openaudiocorpus
    AWS Region
    us-east-1

Edit this dataset entry on GitHub

Home