The Registry of Open Data on AWS is now available on AWS Data Exchange
All datasets on the Registry of Open Data are now discoverable on AWS Data Exchange alongside 3,000+ existing data products from category-leading data providers across industries. Explore the catalog to find open, free, and commercial data sets. Learn more about AWS Data Exchange

LongBench - cross-platform reference dataset profiling cancer cell lines with bulk and single-cell approaches

bam benchmark bioinformatics cancer fastq life sciences long read sequencing short read sequencing single-cell transcriptomics vcf

Description

LongBench is a comprehensive benchmark dataset of the latest long-read transcriptomics technologies from Oxford Nanopore (ON) and Pacific Biosciences, alongside a comparison with next-generation sequencing from Illumina. We generated bulk and single-cell libraries from lung cancer cell lines which include different cancer subtypes to capture real biological variation. To further compare and assess sequencing platform performance, Sequins and SIRVs (Set 4) synthetic spike-ins have been included.

Update Frequency

New data will be added as soon as they are available.

License

CC BY-4.0

Documentation

https://github.com/mritchielab/LongBench.io

Managed By

Richie Lab, Walter and Eliza Hall Institute of Medical Research

See all datasets managed by Richie Lab, Walter and Eliza Hall Institute of Medical Research.

Contact

mritchie@wehi.edu.au

How to Cite

LongBench - cross-platform reference dataset profiling cancer cell lines with bulk and single-cell approaches was accessed on DATE from https://registry.opendata.aws/longbench.

Usage Examples

Tutorials

Resources on AWS

  • Description
    Bulk, single-cell, and single-nucleus RNA-seq data from the LongBench project, covering eight human lung cancer cell lines. Bulk sequencing (FASTQ) was performed on ONT PCR-cDNA, ONT direct RNA (including pod5 files for RNA modification analysis), PacBio Kinnex, and Illumina platforms. Single-cell and single-nucleus sequencing (FASTQ) was performed on ONT PCR-cDNA, PacBio Kinnex, and Illumina platforms. Aligned reads (BAM), variant calls (VCF), and processed gene expression data are also provided, along with reference genome annotations (GTF and FASTA).
    Resource type
    S3 Bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::longbench-data
    AWS Region
    ap-southeast-2
    AWS CLI Access (No AWS account required)
    aws s3 ls --no-sign-request s3://longbench-data/

Edit this dataset entry on GitHub

Tell us about your project

Home