December 18, 2014 --------------------------------------------------------------------------------- Multi-tissue eQTLs computed for GTEx Pilot Phase data using two different methods --------------------------------------------------------------------------------- In addition to single tissue eQTL analysis conducted within the GTEx consortium for the pilot phase data, two methods were developed for joint eQTL analysis across multiple tissues by two groups in the GTEx consortium, one at the University of Chicago (UC) (method called eQTL-BMA) and the other at the University of North Carolina at Chapel Hill (UNC) (method called MT-eQTL). Such multi-tissue eQTL methods increase power to detect eQTLs that are active in more than one tissue, compared to eQTL methods that examine each tissue separately. ------------- Publications ------------- 1. UC Multi-tissue eQTL Method (eQTL-BMA configuration model): Flutre T, Wen X, Pritchard J, Stephens M (2013) "A Statistical Framework for Joint eQTL Analysis in Multiple Tissues". PLoS Genet 9(5): e1003486. doi:10.1371/journal.pgen.1003486 2. UNC Multi-tissue eQTL Method (MT-eQTL): Gen Li, Andrey A. Shabalin, Ivan Rusyn, Fred A. Wright, Andrew B. Nobel (2013) "An Empirical Bayes Approach for Multiple Tissue eQTL Analysis". http://arxiv.org/abs/1311.2948 --------- Software --------- 1. eQTL-BMA config model (UC method): The eqtlbma package can be downloaded here: https://github.com/timflutre/eqtlbma/wiki 2. MT-eQTL (UNC metho)d: MT-eQTL method can be downloaded at: http://www.bios.unc.edu/research/genomic_software/Multi-Tissue-eQTL/ -------------------------------------------------------------------------------------------------- All Results files are packaged in the tar archive "Multi_tissue_eQTL_GTEx_Pilot_Phase_datasets.tar" -------------------------------------------------------------------------------------------------- ## To extract (untar) the archive use the following tar command: tar -xvf Multi_tissue_eQTL_GTEx_Pilot_Phase_datasets.tar ----------------------------------------- Summary and Description of Results Files ----------------------------------------- We present the consensus eQTL profiles from "joint" models for each of the 9 tissues analyzed in the pilot phase (Adipose_subcutaneous, Artery_Tibal, Whole_Blood, Heart_Left_Ventricle, Lung, Muscle_Skeletal, Nerve_Tibial, Skin_Lower_Leg_Sun_Exposed, Thyroid) in the file "res_final_amean_com_genes_com_snps.txt.gz". These values may be interpreted as Pr(SNP is eQTL in tissue s | data). The file "res_final_amean_com_genes_com_snps.txt.gz" represents this consensus, which is an average of the values from the UC model and the UNC model. 9875 eGenes are presented, with the "top" (most significant) SNP in each gene used. The eQTL posterior probabilities for each method separately are given in these files: res_final_uc_com_genes_com_snps.txt.gz res_final_unc_com_genes_com_snps.txt.gz Consensus built on Intersection, Not a Union --------------------------------------------- We note that, although the UC (Flutre et al, 2013) and UNC (Gen Li et al. on arXiv) models are different, the overlap among significant genes and significant SNPs was very high. For example, 98.45% of the UC-declared significant genes were also significant by the UNC model. Thus, we determined that a consensus based on a union of results (instead of an intersection) would add little value. Moreover, because the two models are initially approached from different points of view (the UC model has both gene and snp levels, while the UNC approach focuses on gene-SNP pairs), using an intersection is conceptually simpler than a union. Details ------- The list of genes was obtained by taking the intersection between each group's list. UNC controls for local FDR at the SNP-level and at a 5% threshold, finds 1,284,028 significant SNPs, which happens to be in cis of 17,247 genes. The UC approach controls for the Bayesian FDR at the gene-level and finds 10,030 significant genes at a 5% threshold. The intersection of the two list gives 9875 genes (>98% of UC's eGenes). For each common gene, the top SNP among all common SNPs for this gene was chosen based on the UC posterior of being "the" eQTL given that the gene has one eQTL. For each tissue, the activity probability is the arithmetic average of the UNC posterior of the SNP being an eQTL in this tissue given the data and the UC posterior of the SNP being an eQTL in this tissue given the data. Both posteriors are "unconditional". More precisely, the UC posterior is computed as: Pr(SNP is eQTL in tissue s | SNP is the eQTL, data) x Pr(gene contains an eQTL | data). The UNC posterior is Pr(SNP is eQTL in tissue s). ---------------------------------------------------------------------------------------------------------------------------------------- All the results (for all gene-snp pairs, not only the significant ones as in the files above) are also available in the following files ---------------------------------------------------------------------------------------------------------------------------------------- For both the average bewteen the two methods ('amean') and for each method separately ('uc' and 'unc'): res_final_amean_com_genes_com_snps_all.txt.gz res_final_uc_com_genes_com_snps_all.txt.gz res_final_unc_com_genes_com_snps_all.txt.gz -------------------------------------- Tissue-specificity Configuration Files -------------------------------------- Both UC and UNC models represent eQTL sharing among the 9 tissues in terms of configurations. A configuration is a binary vector of length 9 in which the s-th element is 1 if the SNP is an eQTL in the s-th tissue, and 0 otherwise. There is hence a total of 2^9=512 possible configurations. Both UC and UNC models can compute, for each gene-SNP pair, its (marginal) posterior probability to be in a given configuration, i.e. P(SNP p is eQTL for gene g according to config c | data). We provided 3 files (one from UC, one from UNC and one with the mean between the two methods) with such posteriors for all gene-SNP pairs (~10 millions) for each tissue-specific configurations (i.e. 100000000 for Adipose-specific, 010000000 for Artery-specific, ..., 000000001 for Thyroid-specific) as well as the tissue-consistent configuration (i.e. 111111111). Since the UC and UNC models are different, the absolute values of the posteriors differ between these two models. However, their rankings are very similar for the tissue-specific configurations (Spearman rank correlation coefficient higher than 0.9). The 3 files are: ---------------- res_final_amean_com_genes_com_snps_configs_all.txt.gz res_final_uc_com_genes_com_snps_configs_all_sorted.txt.gz res_final_unc_com_genes_com_snps_configs_all_sorted.txt.gz ---------------------------------------- Dealing with Linkage Disequilibrium (LD) ---------------------------------------- The file "res_final_uc_all_genes_best_snps.txt.gz" was generated only with the UC model, in attempt to account for LD in reporting the most likely eQTLs per gene. The file contains a list of variants for each gene such that the list has a pre-defined probability (here, 95%) of containing "the" eQTL for that gene. This was generated only with the UC model, since this method, as opposed to the UNC method, models the gene and SNP levels explicitly. To handle LD, it assumes that each gene has at most one eQTL (reasonable assumption for the pilot project given that the sample size is low) and that all cis-SNPs are equally likely a priori to be this eQTL. With this, we can give a list of SNPs for each gene so that the list has some pre-defined probability (95%) of containing "the" eQTL for that gene. The main advantage this has over marginal analysis of each SNP is that it competes the SNPs against one another. For example, suppose you have 3 SNPs that are in partial LD with one another, with -log10(p) values of 100.1, 100.0 and 50. All 3 are clearly strongly associated, so would count as "eQTL" in any marginal analysis. But the first two are clearly by far the best candidates for being the eQTN: probably the last one is only associated due to LD with the others (especially if we assume only one eQTL per gene). So this approach has the advantage that it substantially reduces "noise" in eQTL lists that are simply due to SNPs being in LD with an actually functional variant. That is, this approach hence provides a much shorter, more targeted list of "eQTLs" than simply listing every SNP that is significant when looked at on its own. Such a list is available in the file "res_final_uc_all_genes_best_snps.txt.gz". ------------------------------------------------------------------------ Description of the columns in "res_final_uc_all_genes_best_snps.txt.gz": ------------------------------------------------------------------------ This file contains all the 22,286 genes, as well as the best SNPs per gene, for a total of 4,728,367 gene-snp pairs. The column "gene.post" corresponds to the posterior of the gene being an "eGene" (meaning that the gene has an eQTL in at least one tissue). The column "snp.post.the" corresponds to the conditional posterior of the SNP being "the" eQTL given that the gene is an eGene. You can see that summing the column "snp.post.the" for a gene gives a probability just above 0.95. Of most interest for downstream analyses, the columns "snp.post.Adipose", etc, correspond to the marginal posterior of the SNP being active in Adipose, etc. These columns are the same as in the file "res_final_uc_com_genes_com_snps.txt.gz" (described above), which only contains the best SNP per gene, as well as the file "res_final_uc_com_genes_com_snps_all.txt.gz" (described above) which contains all SNPs per gene. The last two columns, "best.config" and "post.best.config" contain the best tissue configuration per gene and posterior probability of the corresponding configuraion, respectively. The tissue numbers inthe configuration correspond to the tissue order in the header. ----------------- Additional Files: ----------------- 1. The files below contain the prior probabilities of any eQTL being in each of the 512 different eQTL configurations across 9 tissues, with the UC or UNC methods. res_final_uc_config_probas.txt res_final_unc_config_probas.txt 2. The files below contain the probability of an eQTL being active in X tissues (X from 1 to 9) by summing config probabilities from the above file, for the UC or UNC methods separately (for UC method see R function "calcActivityProbasPerSubgroup"). In other words, the numbers correspond to the sum of prior probabilities for configurations with eQTL in X tissues. For instance, there are 9 configurations with eQTL only in one tissue. The sum of the 9 prior probabilities using the UNC method is 0.32032. It can also be interpreted as the integrated prior for eQTL configurations with Hamming weight 1. res_final_uc_activity_probas.txt res_final_unc_activity_probas.txt 3. The files below contain information on eQTL sharing between tissues. 3.1. From the UC model: "res_final_uc_pairwise-eqtl-sharing_all-tissues.txt": A matrix for which element i,j corresponds to Pr(being an eQTL in subgroup j | being an eQTL in subgroup i); data obtained by fitting the whole 9 tissues together (see R function calcMarginalPairwiseEqtlSharing) "res_final_uc_pairwise-eqtl-sharing_tissue-pairs.txt": A matrix for which element i,j corresponds to Pr(being an eQTL in subgroup j | being an eQTL in subgroup i); data obtained by fitting each pair of tissues. One can "easily" retrieve these files using the eqtlbma package hosted online: https://github.com/timflutre/eqtlbma/wiki (see program "eqtlbma_hm" as well as R functions in the script "utils_eqtlbma.R"). 3.2. From the UNC model: "res_final_unc_pairwise-eqtl-sharing_all-tissues.txt": This file, similiar to the "res_final_uc_pairwise-eqtl-sharing_all-tissues.txt", above contains the Pi1 analysis results based on the prior probabilities (Pi1 analysis based on conditional analysis in Nica AC, Parts L, Glass D, Nisbet J, Barrett A, et al. (2011) The Architecture of Gene Regulatory Variation across Multiple Human Tissues: The MuTHER Study. PLoS Genet 7(2): e1002003. doi:10.1371/journal.pgen.1002003) . In particular, for each tissue pair, the value in the upper triangle is the conditional probability of a common eQTL given it is an eQTL in the column tissue; the value in the lower triangle is the conditional probability of a common eQTL given it is an eQTL in the row tissue. For instance, P(eQTL in both Adipose and Artery | eQTL in Artery)=0.9313.