Doing More with Less: A Simpler Way to Predict Genetic Risk

Three blue DNA strands against dark blue background

Even as recently as a decade ago, scientists were drowning in data. Today, powerful and efficient data science methods can devour data to extract meaning from it. But more data brings a different challenge: figuring out which information actually matters.

Genomics is one example. Modern genomic datasets include millions to billions of genetic differences across hundreds of thousands of people. How do you find the genetic signals that really matter among so many possibilities?

It’s kind of like trying to identify individual voices in a stadium filled with a million people talking at once. Many are saying similar things, but simply listening to more doesn’t necessarily make the message clearer. Could a computer model make equally good predictions by listening to fewer genomic “voices?”

Robert Tibshirani, PhD, Trevor Hastie, PhD, and Manuel Rivas, PhD – all faculty in the Department of Biomedical Data Science (DBDS) – joined lead author Joshua Richland, a master’s student in the Department of Statistics, to answer that question.

Tibshirani and Hastie had recently developed a machine-learning method called univariate-guided sparse regression, or uniLasso. Working with Rivas, a genomics expert, the team adapted uniLasso to polygenic risk scores, which combine information from many genetic variants to predict a person’s likelihood of having a particular disease, or a trait like height.

They put uniLasso to the test using genetic data from more than 336,000 UK Biobank participants, analyzing more than 1 million genetic variants to build polygenic risk scores for body mass index (BMI), asthma, and coronary heart disease. The results, published in PLOS Genetics, show that uniLasso can do more with less: using far fewer variants to make predictions with nearly the same accuracy.

UniLasso addresses a problem that can arise when a model considers many related factors at once. Consider age and Alzheimer’s disease. On its own, older age clearly predicts greater Alzheimer’s risk. But when age is combined with other factors closely related to it, the model can sometimes get the relationship backward – suggesting, incorrectly, that getting older lowers the risk of Alzheimer’s disease.

UniLasso is designed to put a stop to those errors. It first looks at each factor on its own, then uses that information when considering all the factors together.

It’s sort of like installing a governor on a company van that limits it from speeding. — Robert Tibshirani.

“It’s sort of like installing a governor on a company van that limits it from speeding,” said Tibshirani.

This approach can be especially useful in genomics, because nearby genetic variants are often inherited together and provide much of the same information. UniLasso can sort through those variants and make predictions using far fewer of them.

Polygenic risk scores can draw from enormous numbers of genetic variants, and Lasso methods are already used in genomics to build more focused prediction models. The freely available uniLasso software gives scientists another tool for building simpler genetic prediction models.

Compared with standard Lasso, uniLasso used 38% fewer variants to predict height and 43% fewer for BMI, with little impact on predictive performance. For coronary heart disease, it used fewer than half as many variants with nearly identical accuracy.

The researchers also developed a version called uniLasso ES that can draw on results from other studies without accessing individual participants’ genetic data. For example, using summary results from FinnGen, a health study of about 500,000 people in Finland, uniLasso ES improved predictions across all four traits. This could be particularly useful when privacy restrictions make raw genetic data difficult to share.

What’s next? Larger, more diverse genomic datasets could provide an important test of uniLasso. The UK Biobank participants analyzed in this study were of European ancestry.

Testing this approach in larger and more diverse populations…could help us…identify genetic signals that have been missed in less diverse datasets. — Manuel Rivas

“Testing this approach in larger and more diverse populations is an important next step,” Rivas said. “It could help us understand whether these models work as well across populations and identify genetic signals that have been missed in less diverse datasets.”

The recently released All of Us dataset could provide one opportunity to do that, since 86% of it comes from communities historically underrepresented in biomedical research, including older adults, women, people with disabilities, people of all races and ethnicities, and residents of rural and non-metropolitan areas.

Statistical machine learning is part of the broader artificial intelligence (AI) landscape, but uniLasso does not involve large language models or generative AI. The work instead demonstrates the staying power of relatively simple statistical ideas as tools for cutting-edge discovery: even as newer forms of AI evolve rapidly.

The same lesson applies to genomics: More data doesn’t necessarily require more complexity. By doing more with less, uniLasso could help researchers cut through the complexity of the genome, focus on the signals that matter most, and open new paths toward understanding human health and disease.

 

UniLasso is freely available at https://github.com/jrich129/PRS-Lasso/releases/tag/v1.0.0

 

Additional co-authors on the publication include Tuomo Kiiskinen, postdoctoral scholar in biomedical data science; William Wang, doctoral student in computer science; Wenhui Sophia Lu, doctoral student in statistics; Balasubramanian Narasimhan, senior research scientist in statistics.