The Registry of Open Data on AWS is now available on AWS Data Exchange
All datasets on the Registry of Open Data are now discoverable on AWS Data Exchange alongside 3,000+ existing data products from category-leading data providers across industries. Explore the catalog to find open, free, and commercial data sets. Learn more about AWS Data Exchange

1000 Genomes Phase 3 Reanalysis with DRAGEN 3.5 - Data Lakehouse Ready

bioinformatics biology genetic genomic Homo sapiens life sciences parquet population genetics vcf

Description

The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. There were a total of 3202 individuals sequenced as part of Phase 3 of this project. The high coverage samples were processed using the Illumina DRAGEN v3.5.7b pipeline and are available at s3://1000genomes-dragen/. This dataset contains the VCFs transformed to Parquet/ORC in 3 different schemas - partitioned by samples, partitioned by chromosome and a nested data format. These representations of the 1000 Genomes DRAGEN data are stored in Parquet/ORC format and can be queried through Amazon Athena. To add these tables to your Glue Data Catalog and for sample queries on this dataset, please refer to the link in our Documentation.

Update Frequency

Not updated

License

Data from the 1000 Genomes Project is now available without embargo, following the final publication from the project. Use of the data should be cited in the usual way, with current details available at http://www.internationalgenome.org/faq/how-do-i-cite-1000-genomes-project.

Documentation

https://github.com/aws-samples/data-lake-as-code/tree/roda#readme

Managed By

See all datasets managed by Amazon Web Services.

Contact

https://github.com/aws-samples/data-lake-as-code/issues

How to Cite

1000 Genomes Phase 3 Reanalysis with DRAGEN 3.5 - Data Lakehouse Ready was accessed on DATE from https://registry.opendata.aws/1000-genomes-data-lakehouse-ready.

Usage Examples

Tutorials

Resources on AWS

  • Description
    Parquet representations of 1000 Genomes VCF outputs from DRAGEN, ready for enrollment into Data Lake as Code.
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-hcls-datalake/thousandgenomes_dragen
    AWS Region
    us-east-1
    AWS CLI Access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-hcls-datalake/thousandgenomes_dragen/

Edit this dataset entry on GitHub

Tell us about your project

Home