200 billion sequences and counting: analysis, discovery and exploration of datasets with EBI Metagenomics

EBI metagenomics (EMG, https://www.ebi.ac.uk/metagenomics/) is a freely available hub for the analysis and exploration of metagenomic, metatranscriptomic, amplicon and assembly data. The resource provides rich functional and taxonomic analyses of user-submitted sequences, as well as analysis of publicly available metagenomic datasets held within the European Nucleotide Archive (ENA). EMG has recently undergone rapid expansion, with an over 10-fold increase in data volumes in the first 5 months of 2016. It now houses ~ 50k publicly available data sets, and represents one of the largest collections of analysed metagenomic data. As its data content has grown, EMG has increasingly become a platform for data discovery. To support this process, we have made a series of user-interface improvements, including the classification of projects by biome, presentation of results data for better visualisation and more convenient download, and provision of project level summary files. More recently, we have indexed project metadata for use with the EBI search engine, enabling exploration across different datasets. For example, users are able to search with a particular taxonomic lineage or protein function and discover the projects, samples and sequencing runs in which that lineage or function is found. This functionality allows users to explore associations between biomes, environmental conditions and organisms and functions (e.g., discovering protein coding sequences that correspond to certain enzyme families found in aquatic environments at a given temperature range). Here, we give an overview of the EMG data analysis pipeline and web site, and illustrate the use of the new search facility for data discovery.

Licence: Creative Commons Attribution Non Commercial No Derivatives 4.0 International

Keywords: metagenomics

Remote created date: 2016-12-16

Remote updated date: 2017-01-11


Activity log