biotech blueprint chemistry genetic genomic life sciences parquet
Amazon is no longer hosting this Data Lakehouse Ready dataset
ClinVar is a freely accessible, public archive of reports of the relationships among human variations and phenotypes, with supporting evidence. ClinVar thus facilitates access to and communication about the relationships asserted between human variation and observed health status, and the history of that interpretation. ClinVar processes submissions reporting variants found in patient samples, assertions made regarding their clinical significance, information about the submitter, and other supporting data. The alleles described in submissions are mapped to reference sequences, and reported according to the HGVS standard. ClinVar then presents the data for interactive users as well as those wishing to use ClinVar in daily workflows and other local applications. ClinVar works in collaboration with interested organizations to meet the needs of the medical genetics community as efficiently and effectively as possible. This representation of ClinVar is stored in Parquet format and most easily utilized through Amazon Athena. Follow the documentation link for install instructions (< 2 minute install).
Every Sunday at 1AM UTC
https://github.com/aws-samples/data-lake-as-code/blob/roda/docs/roda_attributions.txt
https://github.com/aws-samples/data-lake-as-code/blob/roda/docs/roda_install.md
See all datasets managed by Amazon Web Services.
https://github.com/aws-samples/data-lake-as-code/issues
ClinVar - Data Lakehouse Ready was accessed on DATE
from https://registry.opendata.aws/clinvar.
arn:aws:s3:::aws-roda-hcls-datalake/clinvar_summary_variants/
us-east-1
aws s3 ls --no-sign-request s3://aws-roda-hcls-datalake/clinvar_summary_variants/