Science is a global practice. Most papers I read today involve collaborations, a lot of the collaborations are across countries. Many(all?) of the papers use datasets and computational methods/tools that originated across the globe. It’s sometimes interesting to take back and look at a country’s contribution to a field. I’ve worked in the field of human genetics and the contribution of the UK stands out to me.

Let’s start from the origins. Mendel wasn’t English but there’s Darwin, Wallace and Huxley. However, I’ll stick the quants for this post. One could argue that the roots of modern human genetics started in the UK. Galton introduced the idea of regression, an everyday(every hour?) word in statistical genetics as well as in science. Fisher’s contributions to statistics and population genetics will span several books. And how could we ever forget Haldane? I’d love to claim Haldane as Indian due to his later life but that would be a travesty. There’s a whole list of other population geneticists from the UK who have made seminal contributions in the 1900s.

Let’s move to technology. There’s Sanger sequencing developed by the inimitable Fred Sanger. Solexa started out in England and became today’s sequencing behemoth Illumina. Oxford Nanopore is arguably the most promising alternative to Illumina. Ion Torrent which uses current to decode nucleotides is also derived from a British company. That’s most of the sequencing companies that I know. The Sanger Institute is still one of the leading sequencing centers around the world, constantly reinventing itself as the focus of genetic studies have kept changing over the past few decades.

The role of the Sanger institute in the human genome project led by John Sulston has been well documented. Adjusted for population size/GDP the UK’s role in the project has to be the biggest. Following the human genome project, several large human genetic studies came out of the UK and have set the stage for today. The Welcome Trust Case Control Consortium was one of the first large GWAS tackling many common diseases at once. The 1000 genomes project catalogued genetic variation around the world first with low-depth sequencing followed by deep sequencing. The easy access and clear description of the data and methods from the 1000 genomes project has been unsurpassed by other genetic studies. In fact there have been much larger genetic studies in the US and elsewhere, but the data from some of these studies have not gone past a few investigators. Most of these studies result in a flagship publication for the consortium and then the data is buried in a solid state drive. The 1000 genomes project has really set the standards for sequencing studies, and sadly that benchmark is often not met.

The UK has also led the world in computational and statistical methods development for the sequencing era. Several of the most widely used aligners including BWA came out of the UK. The format specifications for files dealing with NGS also originated in the UK (SAM, BAM, VCF). This is a crucial development to Bioinformatics else we would be spending more time than we do now in the madness of converting files from one format to another. The initial round of variant callers also originated mostly in the UK. A lot of these methods formed the theoretical basis for the next generation of variant callers that sprung up elsewhere with time. Again most of the tools developed in the UK were open access, had clear documentation and well maintained mailing lists. Some of the prominent tools that came out elsewhere like the GATK were initially not open-source.

The current focus of genetic studies is on capturing diverse populations. Here too the UK set the stage and led the initial efforts. The sequencing of British Pakistanis to identify homozygous LOFs was one of the early studies. The East London project is also another promninent study in this direction.

I’ve written about biobanks and the important role they play in genetic studies in a previous post. The best established biobank today is clearly the UK biobank. The data is used by thousands of investigators around the world. This sort of data sharing is sadly an exception hardly the norm when these studies are conducted in other countries. Apart from the genotypes, the repeated phenotyping of these individuals makes this data invaluable. I would argue that the UK biobank data might be used hundreds of years down the line to look back at the state of human populations in the 2000s.

Look, I’m not a fanboy of the UK, in fact I’m pretty far from that. I’m sure we can find flaws in science from the UK just like any other country. However, the repeated excellence of UK based genetic studies particularly in terms of data sharing and open-source methods development cannot just be a coincidence. There’s something different about the priorities of the studies designed there. There is not a tendency to hide behind privacy concerns but a willingness to balance these concerns with scientific progress. The vested interests of a few investigators seems to be often set aside. I’ll drink a cup of tea to celebrate that.