Gene expression integration and similarity score–based modeling improve risk stratification in idiopathic venous thrombophilia

publication
bioinformatics
transcriptomics
thrombosis
A new study in the Journal of Thrombosis and Haemostasis led by Pol Ezquerra integrates whole-blood transcriptomics, genetic variants, and clinical risk factors with machine learning to better stratify patients with idiopathic venous thromboembolism — including long non-coding RNAs not previously linked to thrombosis.
Author

Pol Ezquerra

Published

August 31, 2026

Modified

August 31, 2026

When blood clots without an obvious reason

Some people get a venous clot — deep vein thrombosis, sometimes a pulmonary embolism — with no surgery, no long flight, no cancer, no obvious trigger behind it. Clinicians call these events idiopathic, and they are the frustrating ones. Two patients can receive the same diagnosis and yet face very different futures, and the tools we normally use do not always tell them apart.

What we did

In a study just published in Journal of Thrombosis and Haemostasis, we asked whether the blood’s own gene expression could help sort out who is really at risk. We worked with data from the GAIT2 project: 790 individuals, of whom 70 had already suffered an idiopathic venous thromboembolism. For every person we brought together three layers of information — whole-blood RNA sequencing, the genetic variants we already know are involved in thrombosis, and plain clinical variables such as body mass index, age, and ABO blood group.

Then we let two supervised models, Elastic Net and XGBoost, look for patterns across all of it at once.

What the models saw

Both models agreed on the top player: von Willebrand factor, a protein we know sits right at the centre of clot formation. After that came the clinical variables, and then something more interesting — the expression of 494 genes, among them STS, FAM13A, GPRIN1, FLVCR2 and FAM177B, and several long non-coding RNAs that had never been connected to thrombosis before. It was reassuring that the models also recovered UQCRC2 and PRKRA, genes that have already appeared in thrombosis studies; the signatures were not completely alien.

The enrichment analyses pointed to cardiomyopathy-related KEGG pathways and renal HPA terms, which suggests that some of the story lies outside the coagulation cascade as we usually draw it.

Controls that look like cases

The part I find most interesting is what came next. We combined the two models into a single similarity score and asked, for every control in the cohort, a slightly uncomfortable question: how close is this healthy person — at the molecular and clinical level — to the patients who did have a clot?

The score put 74% of the VTE cases inside a “risk zone”, but also 23% of the controls. Those controls were never diagnosed with anything, yet they walk around with a profile that looks a lot like the patients’ own. They are exactly the kind of people we might want to keep an eye on.

Where this leaves us

I would not claim that gene expression profiles will replace clinical judgement tomorrow. What the study does show is that transcriptomic information adds something real on top of what BMI, age, ABO or the usual scores can tell us. A small panel of blood-expressed genes, lncRNAs included, may eventually help detect people with a high baseline predisposition before the first clot happens. That is, in the end, what personalised prevention is supposed to be about.

This is the outcome of a long collaboration with the groups behind the GAIT2 project, and it fits the line of work we have been pursuing at B2SLab: integrating omics data and machine learning with an eye on clinical translation.


Reference: Ezquerra-Condeminas P, Martinez-Perez A, Howald C, Brown AA, Souto JC, Viñuela A, Buil A, Perera-Lluna A, Soria JM. Gene expression integration and similarity score–based modeling improve risk stratification in idiopathic venous thrombophilia. Journal of Thrombosis and Haemostasis (2026) 24(8):3002–3016. https://doi.org/10.1016/j.jtha.2026.05.015