The Registry of Open Data on AWS is now available on AWS Data Exchange
All datasets on the Registry of Open Data are now discoverable on AWS Data Exchange alongside 3,000+ existing data products from category-leading data providers across industries. Explore the catalog to find open, free, and commercial data sets. Learn more about AWS Data Exchange

YouTube 8 Million - Data Lakehouse Ready

computer vision labeled machine learning parquet video

Description

This both the original .tfrecords and a Parquet representation of the YouTube 8 Million dataset. YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. It comes with precomputed audio-visual features from billions of frames and audio segments, designed to fit on a single hard disk. This dataset also includes the YouTube-8M Segments data from June 2019. This dataset is 'Lakehouse Ready'. Meaning, you can query this data in-place straight out of the Registry of Open Data S3 bucket. Deploy this dataset's corresponding CloudFormation template to create the AWS Glue Catalog entries into your account in about 30 seconds. That one step will enable you to interact with the data with AWS Athena, AWS SageMaker, AWS EMR, or join into your AWS Redshift clusters. More detail in (the documentation)[https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/README.md.

Update Frequency

Google Research has not updated the dataset since 2019.

License

https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/docs/roda_attributions.txt

Documentation

https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/docs/roda_install.md

Managed By

See all datasets managed by Amazon Web Services.

Contact

https://github.com/aws-samples/data-lake-as-code/issues

How to Cite

YouTube 8 Million - Data Lakehouse Ready was accessed on DATE from https://registry.opendata.aws/yt8m.

Usage Examples

Tutorials
Publications

Resources on AWS

  • Description
    Original YT8M *.tfrecords. File structure info can be found here.
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake/yt8m/
    AWS Region
    us-west-2
    AWS CLI Access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake/yt8m/
  • Description
    Lakehouse ready YT8M as Glue Parquet files. Install instructions here.
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake/yt8m_ods/
    AWS Region
    us-west-2
    AWS CLI Access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake/yt8m_ods/
  • Description
    Replica of the two locations above in us-east-1.
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake-us-east-1/
    AWS Region
    us-east-1
    AWS CLI Access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake-us-east-1/

Edit this dataset entry on GitHub

Tell us about your project

Home