blood-proteins-predict-diabetes-and-heart-disease-years-before-symptoms-in-obesity
Blood Proteins Predict Diabetes and Heart Disease Years Before Symptoms in Obesity

Blood Proteins Predict Diabetes and Heart Disease Years Before Symptoms in Obesity

Obesity has long been treated as a single, monolithic risk factor, a red flag that automatically signals diabetes, hypertension, and heart disease somewhere down the road. But clinicians have known for decades that the label conceals enormous variability: some people with high body fat develop metabolic disease within a few years, while others remain free of complications for decades. A new study published in Genome Medicine offers a way to tell these trajectories apart years in advance, using nothing more exotic than a blood sample and a carefully constructed statistical model. By measuring nearly three thousand circulating proteins in thousands of people with obesity who were still clinically healthy, researchers built risk scores that forecast who would develop type 2 diabetes, high blood pressure, or dyslipidemia, often long before any conventional test would sound an alarm.

The research, led by Yang Liu and Dewen Dong of the Health Sciences Institute at China Medical University, together with colleagues including John R. Speakman and senior author Deliang Wen, drew on the UK Biobank, one of the largest biomedical databases in the world. The team focused on 13,931 participants who met the study’s definition of pre-clinical obesity: they had abnormal adiposity, established through body mass index and related anthropometric measures, but had not yet been diagnosed with any obesity-related cardiometabolic disease at baseline. This population is precisely where early prediction matters most, because it represents a window of opportunity in which intervention could still prevent disease rather than merely manage it.

The biological raw material came from the UK Biobank Pharma Proteomics Project, which used affinity-based assays to quantify 2,919 proteins circulating in each participant’s plasma. Proteins are the workhorses of physiology, and their concentrations in blood reflect the real-time state of tissues and organs in ways that genetic data cannot. A person’s genome is essentially fixed at conception, but the proteome is dynamic, shifting with inflammation, insulin resistance, liver fat, vascular stress, and countless other processes that precede overt disease. That dynamism is exactly what makes plasma proteins attractive candidates for early warning systems.

The first analytical pass was deliberately simple: the researchers tested each protein individually for associations with future diagnoses of type 2 diabetes, hypertension, and hyperlipidemia using Cox proportional hazards models, the standard statistical framework for time-to-event data in epidemiology. To guard against false positives across such a massive number of tests, they applied Bonferroni correction, one of the most stringent thresholds in statistics. Even so, the screen yielded 973 significant protein-disease associations involving 653 unique proteins. The sheer volume of signal confirmed what many smaller studies had hinted at: the blood proteome carries a deep reservoir of information about cardiometabolic futures.

Individual proteins, however, are noisy predictors on their own. The team’s real innovation was to combine them into composite proteomic risk scores, which they called ProteinScores, using elastic-net Cox regression. Elastic net is a machine learning technique that performs variable selection and regularization simultaneously, shrinking the coefficients of less informative proteins toward zero while retaining a parsimonious set of features. This approach tames the multicollinearity that plagues proteomic data, where hundreds of proteins travel together in correlated clusters reflecting shared biological pathways. The researchers built separate ProteinScores for three prediction horizons, five, ten, and fifteen years, recognizing that the molecular signatures of disease emerging within five years may differ from those portending trouble a decade or more out.

Evaluating the scores required rigorous benchmarking against increasingly demanding comparison models. The ProteinScores clearly improved discrimination beyond simple age- and sex-adjusted models and beyond baseline covariate models incorporating standard clinical variables across all prediction horizons. The tougher test came from extended clinical models that already included disease-specific indicators such as glycated hemoglobin, blood pressure, and lipid measurements. Against these, the incremental gains were more outcome- and horizon-dependent, with statistically significant improvements in the area under the receiver operating characteristic curve, or AUC, appearing mainly for selected longer-term predictions. That pattern makes physiological sense: clinical indicators capture the disease process once it is already underway, whereas proteins can reveal subclinical perturbations that precede measurable changes in glucose or blood pressure by years.

Discrimination alone does not make a risk score clinically useful; a model must also be well calibrated, meaning its predicted probabilities match observed event rates, and it must meaningfully reclassify patients into more accurate risk categories. The team assessed both. Calibration analyses and category-based net reclassification improvement, a metric quantifying how many people are moved into correct risk strata, showed that the ProteinScores provided additional risk-stratification information in several disease-horizon combinations, though the magnitude of benefit varied. The researchers also pitted their proteomic scores against two rival omics platforms available in the same cohort: polygenic risk scores from UK Biobank fields and metabolomic scores built from circulating metabolites. In this framework the ProteinScores showed favorable performance overall, but the relative ranking shifted by outcome and horizon, underscoring that no single molecular layer holds a monopoly on predictive power.

Beyond prediction, the study delivered biological insight. Pathway enrichment analyses of the proteins selected into the scores implicated lipid transport, lipoprotein remodeling, extracellular matrix organization, peptide hormone processing, and immune-inflammatory signaling. These are not random associations; they map onto the canonical axes of cardiometabolic disease biology. Lipid transport and lipoprotein remodeling connect directly to atherosclerosis and dyslipidemia, extracellular matrix organization to the fibrotic remodeling of adipose tissue that accompanies metabolic dysfunction, peptide hormone processing to the endocrine signaling disrupted in insulin resistance, and immune-inflammatory pathways to the chronic low-grade inflammation that characterizes unhealthy obesity. In effect, the risk scores distill the molecular conversation between fat tissue, liver, vasculature, and immune system into a single number.

The clinical implications are considerable. Current guidelines increasingly recognize that obesity is heterogeneous, distinguishing metabolically healthy from metabolically unhealthy obesity, but these categories are crude snapshots rather than forecasts. A proteomic risk score could transform pre-clinical obesity from a vague state of heightened risk into a stratified condition where patients know, with quantified probability, what their five-, ten-, and fifteen-year futures hold. High-risk individuals could be prioritized for intensive lifestyle intervention, pharmacotherapy such as incretin-based drugs, or heightened surveillance, while lower-risk individuals might be spared unnecessary treatment and anxiety. Because the scores rely on plasma proteins measurable with existing affinity-based platforms, the path from research to clinical deployment is shorter than for many emerging omics technologies, though cost, standardization, and prospective validation in diverse populations remain hurdles.

The study’s limitations are those inherent to its design. It was conducted within the UK Biobank, a cohort that skews toward middle-aged British participants and may not generalize to other ancestries or health systems, and the authors note that the magnitude of improvement over extended clinical models varied across outcomes and horizons rather than being uniformly large. The scores were also developed and evaluated within the same data framework, so external validation in independent cohorts will be essential before routine clinical use. Nevertheless, the work represents a substantive step toward precision prevention in obesity medicine. Rather than waiting for diabetes or hypertension to declare itself, physicians may one day read the proteomic tea leaves in a single blood draw, identifying which of their patients with obesity are quietly heading toward disease and intervening while prevention is still possible. In a world where more than a billion people live with obesity, that kind of foresight could reshape public health on a global scale.

Subject of Research: Proteomic risk scores for early prediction of cardiometabolic diseases in pre-clinical obesity

Article Title: Proteomic risk scores facilitate early prediction of cardiometabolic diseases in pre-clinical obesity

Article References: Liu, Y., Dong, D., Cheskin, L. J., Jiang, L., Cao, G., Speakman, J. R., & Wen, D. (2026). Proteomic risk scores facilitate early prediction of cardiometabolic diseases in pre-clinical obesity. Genome Medicine. https://doi.org/10.1186/s13073-026-01780-1

Image Credits: AI Generated

DOI: 10.1186/s13073-026-01780-1

Keywords: proteomics, pre-clinical obesity, type 2 diabetes, hypertension, hyperlipidemia, UK Biobank, risk prediction, elastic-net regression, polygenic risk scores, metabolomics, precision prevention, cardiometabolic disease