HSG journal club discussion summary - 2018/08/31
Official link - https://zork5.wustl.edu/hsg/bio5489/summary_8-31-2018.html
Paper choice: Khera et al. (2018), Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations, Nat Genet 30:1219-1224.
Link to the paper: https://www.nature.com/articles/s41588-018-0183-z
This year, the Human Genetics Journal Club at Washington University will be writing summaries of the discussions that arise during our journal club. Our goal is to capture some of the salient discussion points that may be interesting or useful to researchers both at WashU and beyond.
Our organizer, Dr. Nancy Saccone, presents the first journal club of each school year. This year, Dr. Saccone selected the recent article “Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations” by Khera et al. in Nature Genetics. The primary claim of this paper is, as the title states, that polygenic risk scores can help capture individuals at actionable risk (similar to the risk from monogenic mutations) for common diseases. Dr. Saccone had a couple of reasons for choosing this paper. Polygenic risk scores, referred to as GPS (genome-wide polygenic scores) in this paper and PRS in other articles, attempt to measure the risk conferred by a person’s genetic background towards a specific disease. GPSes are currently a highly active area of human genetics research and efforts are underway to identify their utility. If GPSs prove to be useful, they could find application in the clinic and in directing public health efforts.
To briefly summarize the paper, the authors computed a set of putative polygenic risk scores for five common diseases using the results of previous large genome-wide association studies. Next, the authors picked the “best-performing” GPS for each disease, as measured by the Area under the Receiver-Operator Curve, and applied it to an independent test set of UK Biobank participants. This helped the authors predict study participants with three, four, and five-fold increased risk of contracting the disease when compared to the rest of the population. The authors show with the actual prevalence of disease (obtained from the participants’ medical records) that an elevated GPS does predict high-risk individuals. Interestingly, the high risk groups identified by GPS were not necessarily identified by conventional risk factors; we would have liked to see this point explored further, e.g. by comparing the high-GPS group to the union of individuals with at least one conventional risk factor. This study, to our knowledge, is the largest application of GPSs for the study of common disease to date.
One of the first points raised during our journal club discussion was: how meaningful is a 3-fold increase in risk for a given disease? For example, is an increase in risk from 1 in a million to 3 in a million still relevant to public health or clinical efforts? The general consensus was that, in order understand the answer to this question, we need good estimates of the general baseline risk for these diseases. Table 1 shows the prevalence of the five diseases in the training and testing datasets, but this was not emphasized in the paper. Because the diseases studied here are common, the public health relevance is strong.
The authors used Area Under the Receiver-Operator Curve (AUC) as a metric to choose between the candidate GPSs. The improvement in AUC when switching from a GPS with approximately 100 SNPs to a GPS with > 6 million SNPs is tiny, as seen in Supplementary Table 1. We mused about what sort of an advantage this small increase in AUC confers and whether it would be equally effective to choose a model with fewer SNPs. From Supplementary Table 1, the Odds Ratio (OR) per Standard Deviation (SD) of the polygenic score increases as the prediction model includes more SNPs; does this confer increased specificity for binning the population into high-risk groups? Again, an extended discussion related to model selection, whether in the main paper or the supplement, would have been helpful in the interpretation of model choice.
Another discussion centered around what action could (or should) be taken once we can identify high-risk groups of individuals. It was suggested that if we could show that GPSs are equally or more helpful than the current clinical calculators, doctors might consider treatment options like early statin treatments for some members of the population. Insurance companies seem to gain the most from such scores currently since it would give them an alternative, and possibly more accurate, method to evaluate risk. However, will insurance companies cover the cost of sequencing and GPS calculation or will these early intervention tools only be accessible to those who can afford to pay out-of-pocket? Naturally, this led to a brief discussion of whether health insurance and privacy laws have caught up with the recent advances in genomic technologies.
We also discussed the transferability of using GPSs to diagnose potential patients of all ethnicities. This paper focuses on GPSs that were computed by using phenotype and genotype data from the UK BioBank, where participants are primarily of European ancestry; the source GWAS studies were also primarily of European ancestry. Several of us were reminded of another paper presented in our journal club last year (“Human Demographic History Impacts Genetic Risk Prediction across Diverse Populations”, Martin et al., AJHG 2017) that showed polygenic risk scores derived from single-ancestry GWAS were not able to consistently predict disease risk in individuals from other populations. Everyone seemed to agree that it is only a matter of time before GPS become diagnostic tools. However, great care must be taken to ensure that they are not misapplied (i.e. using data from an European-ancestry GWAS to assess risk for an individual of non-European descent, potentially leading to a misdiagnosis). Our discussion also re-emphasized the authors’ point that the benefits of GPS need to quickly become available to all populations, and that sufficiently large and high-quality datasets from non-European populations need to continue to be generated.
A few other points came up as well. It was not immediately clear whether or not the BRCA1 SNP and similarly monogenic SNPs for other diseases were excluded from the polygenic risk score calculation. Also, the variance explained by the GPS is only around 2 - 4% across the five diseases studied here. This is not very high. What are the implications of explaining this low amount of variance, and how would the accuracy of the GPS change with more variance explained? Will it be possible to integrate environmental factors into the risk score calculation, if data such as smoking status is available? As a minor point, we also noted that some of the tables in the main paper were inefficient in their use of space.
Time will tell whether GPSs make it to the clinical or epidemiological mainstream, but the indications from this paper are generally positive, at least for the diseases tackled here.
Summarizing author: Avinash Ramu
Summary team: Yize Li, Ju Heon Maeng, Alina Schmidt, Celine St. Pierre, Lijun Yao
Editor: Dr. Nancy Saccone