I find biobanks fascinating
Human genetic association studies have been getting larger, and some of the meta-analysis use millions of individuals. Consumer genetic databases such as 23andMe are approaching(have already reached) 5-10 million consumers. These are massive studies unimaginable a few decades ago. The massive sample-sizes allow detection of variants with tiny effect. How do the variants add up to produce different phenotypes, that’s a question for the current century. We’re slowly making progress in that area. Today I want to talk about biobanks.
In my Epi class we learnt about a couple of different study designs. One common study design is the case-control study. Most of the disease GWAS use this study desing. Researchers identify a set of individuals with the desired trait as cases and another set of individuals without the trait as controls. These individuals’ genetic information is then collected and analyzed to find differences between the two groups. The phenotypic information available is whatever is available at the time of recruitment, usually a broad survey of the individual.
The second form of study design is the cohort study. Here a group of individuals is followed over time. At some point genetic information can be collected from the participants. But the key is that the individuals are followed over time. So any new phenotype or changes in phenotype can be tracked. 23andMe, Framingham heart study, Geisinger, the UK biobank are all examples of this sort of a study design. I think cohort studies are going to be super useful as we approach the next generation of genetic studies.
Right now, the Covid pandemic is on. Many individuals show mild symptoms but some individuals in various age-groups show serious symptoms. Why is this? Imagine you had healthcare and genetic data about these individuals for a long time, perhaps in say a Kaiser Permanante health system or as part of Icahn’s BioME biobank. You could start identifying risk factors, genetic or others that might correlate with disease severity. There’s no need to recruit and assemble another case-control cohort and spend millions trying to genotype them. Instead all the information is essentially already collected in the health records and can be repurposed to a new research question.
I’ve been trying to keep up with biobanks around the world. Biobanks seem to work well with nationalised healthcare systems, for example in the UK and Finland. In the US, Kaiser Permanante has a genetic research wing. Other biobanks in the US that I know of are Vanderbilt, Geisinger and BioME in NYC.
It makes sense for all large private hospitals with an interest in research to start this form of centralised information and specimen collection. This will help identify genotypic and phenotypic outliers right away.