VariantSpark Notebook
AEHRC
Price: Typical Total Price$0.005/hrTotal pricing per instance for services hosted on t3a.nano in US East (N. Virginia). View Details
Product Overview
VariantSpark is a scalable toolkit for genome-wide association studies optimized for GWAS like datasets. Machine learning methods and, in particular, random forests (RFs) are a promising alternative to standard single SNP analyses in genome-wide association studies (GWAS) and from scalable to rare variants from whole genome sequence data. RFs provide variable importance measures to rank genomic locations according to their predictive power to the disease or phenotype. Although there are a number of existing random forest implementations, some even parallel or distributed such as: Random Jungle, ranger or SparkML, none are optimized to deal with modern whole genome datasets, containing thousands of samples and millions of variables. Implemented directly on Apache Spark core, VariantSpark builds random forest models and estimates variable importance using the mean decrease gini method, processing VCF and CSV files. The package also includes a Jupyter notebook with examples to perform Quality Control and data manipulation tasks using HAIL.is (included in the package) as well as for visualizing the results.VariantSpark can process 200 samples with 20M variables in 1 hour consuming $3 of AWS resources. VariantSpark compute time increases linearly with both variables and samples.
Version: 1.1.2
By
AEHRC
Categories: Education & Research
High Performance Computing
Healthcare & Life Sciences
Operating System
Linux/Unix, Ubuntu 20.04
Delivery Methods
CloudFormation Template
Highlights: VariantSpark can work directly with the VCF data, without the costly pre-processing required by other tools due to its novel approach of building random forest models. VariantSpark is implemented directly on top of Apache Spark - a modern distributed framework for big data processing, which gives VariantSpark the ability to scale horizontally to process even whole genome sequence data. More information available in our peer-reviewed publication O'Brien et al. VariantSpark: population scale clustering of genotype information BMC Genomics 2015 and our most recent pre-print Bayat et al. VariantSpark, A Random Forest Machine Learning Implementation for Ultra High Dimensional Data BioRxiv 2019.
Category: Infrastructure Software
Delivery Method: CloudFormation Template
Company Information
Company Name: AEHRC
About Company: The Australian e-Health Research Centre (AEHRC) is the leading national research facility applying information and communication technology to improve health services and clinical treatment for Australians.
AEHRC is a joint venture between CSIRO and the Queensland Government, through Queensland Health. AEHRC is developing and deploying digital innovations to improve service delivery in the Queensland and Australian health systems, generate commercialisation revenue, and increase the pool of world-class e-health expertise in Australia.