Introduction

Breast cancer is the most common malignancy among women and remains the leading cause of cancer-related mortality worldwide, with an estimated 665,684 deaths in 2022 [1]. Clinically, breast cancer is categorized into estrogen receptor-positive (ER+) and estrogen receptor-negative (ER−) subtypes based on estrogen receptor expression [2]. Key pathological indicators play a critical role in diagnosis, prognostic assessment, and therapeutic decision-making, such as tumor subtype, histological grade, and ER/PR/HER2 status [3]. However, substantial inter-patient heterogeneity and variability in treatment response limit the prognostic accuracy of these conventional markers [4]. Thus, there is an urgent need for more precise prognostic biomarkers to guide individualized therapy [5].

Circulating proteins have emerged as promising biomarkers for disease diagnosis, prognosis, and therapeutic targeting [6, 7]. In breast cancer, proteomic studies have identified key candidates such as PRKDC, which shows elevated phosphorylation in the Basal-I subtype and may serve as a subtype-specific target [8]. However, studies specifically linking circulating proteins to breast cancer prognosis remain limited [9]. Mendelian randomization (MR) uses genetic variants as instrumental variables (IVs) to infer causal relationships between exposures (e.g., circulating proteins) and disease outcomes, minimizing confounding and reverse causation [10]. Integrating MR with large-scale plasma proteomics provides novel insights into the genetic basis of cancer progression [11].

In this study, we applied colocalization analysis, summary-based MR (SMR), heterogeneity in dependent instruments (HEIDI) tests, and two-sample MR (TSMR) to systematically identify circulating proteins associated with breast cancer survival. Additionally, single cell type expression analysis and druggability assessments were performed to explore their potential as targets for improving breast cancer prognosis [12]. This study aimed to identify circulating proteins associated with breast cancer prognosis, providing insights for therapeutic development.

Material and methods

The study design is presented in Figure 1. We sequentially applied Bayesian colocalization, SMR, HEIDI tests, and TSMR to validate the potential causal relationships between protein biomarkers and breast cancer survival, with additional validation using GTEx and eQTLGen data. To determine the tumor-specific expression patterns of the corresponding genes, we conducted single-cell RNA-seq analysis to explore cell type-specific enrichment in breast cancer tissues. Finally, we performed protein-protein interaction (PPI) and druggability analyses to evaluate their therapeutic potential.

Figure 1

Study design

UKBPPP – UK Biobank Pharma Proteomics Project, BCAC – Breast Cancer Association Consortium, GWAS – genome-wide association study, HEIDI – heterogeneity in dependent instrument, MR – Mendelian randomization.

https://www.archivesofmedicalscience.com/f/fulltexts/208246/AMS-22-3-208246-g001_min.jpg

Proteomic data source

Summary statistics of genetic associations with plasma proteins were extracted from four large proteomic studies: Alexander (2091 proteins) [13], deCODE (4907 proteins) [14], Fenland (3892 proteins) [15], and UKBPPP (1478 proteins) [16]. For external validation, eQTLGen (whole blood expression data from > 31,000 individuals) and GTEx (multi-tissue expression data) were used.

Outcome data sources

The study included 37,954 breast cancer patients of European ancestry, with 2,900 deaths recorded. Subtype-specific analyses included 6,881 ER- patients (920 deaths) and 23,059 ER+ patients (1,333 deaths). Details on study populations, genotyping, and imputation methods are available in prior publications [17]. Ethics approvals and informed consent were obtained. Appendix lists the sources and corresponding information of all aggregated statistical datasets used in this study.

Bayesian colocalization analysis

For each locus, the Bayesian method assessed the support for the following five exclusive hypotheses: 1) no association with either trait; 2) association with trait 1 only; 3) association with trait 2 only; 4) both traits are associated, but with distinct causal variants for each trait; and 5) both traits are associated and share the same causal variant. The analysis provides posterior probabilities for each tested hypothesis (H0, H1, H2, H3, and H4). We used the following prior probabilities: p1 = 10−4, p2 = 10−4, and p12 = 10−5. Colocalization was defined as PP4 > 0.5.

SMR analysis

MR analysis, treating plasma proteins as exposures and breast cancer survival as outcomes, used Bonferroni correction to adjust for multiple testing. Specifically, proteins were categorized into three groups: (1) no colocalization evidence, (2) moderate colocalization evidence (e.g., PP4 between 0.5 and 0.8), and (3) high colocalization evidence (e.g., PP4 ≥ 0.8). The Bonferroni correction set the significance threshold at p < 0.05.

The HEIDI test, applied when ≥3 SNPs were available, excluded associations with pleiotropy (pHEIDI < 0.05). SNPs with high (r2 > 0.9) or weak (r2 < 0.05) linkage disequilibrium (LD) were excluded.

Selection of genetic instruments

In TSMR analysis, cis-pQTLs (± 1 MB of the gene) with p < 5 × 10–8 were used as IVs. Furthermore, SNPs with allele frequency differences greater than 0.2 between the pQTL data and the GWAS data were excluded. We permitted the exclusion of a maximum of 5% of SNPs based on allele frequency differences. IV strength was assessed using the F-statistic (F = β2/SE2), and SNPs with F < 10 were considered weak and excluded. Finally, the top-associated SNP with gene expression was selected as the genetic instrument [10, 12].

TSMR analysis

TSMR analysis was further conducted to verify the causal associations between proteins and breast cancer survival. The following criteria were used to select instruments and proteins: (i) SNPs associated with any protein were selected (p < 5 × 10−8); (ii) SNPs and proteins within the major histocompatibility complex (MHC) region (chr6: 25.5–34.0Mb) were excluded due to their complex LD structure; (iii) LD clumping was then conducted to identify independent pQTLs for each protein (r2 < 0.01); (iv) the R2 and F-statistic (R2 = 2 × EAF × (1 – EAF) × beta2; F = R2 × (N – 2)/(1 – R2)) were used to estimate the strength of genetic instruments, where R2 was the proportion of the variability of the protein levels explained by each genetic instrument.

We performed sensitivity analyses using Cochran’s Q, MR-Egger intercept, and Steiger filtering, with significance determined by corresponding p-values.

Single cell-type expression analysis

To explore cell type-specific expression, we analyzed single-cell RNA-seq data (GSE176078) from breast tumor tissues, focusing on genes encoding proteins with potential causal effects on breast cancer. Low-quality cells were filtered out, and the remaining data were log-normalized. To assess whether breast cancer survival–associated genes are preferentially expressed in specific cell types within breast tumor tissue, differential expression analysis was performed using the Wilcoxon rank-sum test. The genes with an average log2 fold change (log2FC) more than 0.5 and a false discovery rate (FDR) adjusted p-value less than 0.05 were identified as enrichment genes in a cell type.

Based on colocalization analysis, MR analysis, and single-cell specificity analysis, we classified proteins into three distinct target groups. Those that met all MR-based prioritization criteria and exhibited cell type-specific enrichment were assigned to Tier 1, those with posterior probability of hypothesis 4 (PPH4) > 0.8 and cell type-specific enrichment were assigned to Tier 2, and proteins lacking single-cell expression or with moderate colocalization evidence were assigned to Tier 3.

Immune infiltration analysis

To assess the relationship between prognostic protein-coding gene expression and immune cell infiltration in breast cancer, we employed the TIMER web tool (Tumor Immune Estimation Resource, http://cistrome.org/TIMER/) [18]. In this study, the “gene” module was used to evaluate the relationship between gene expression and immune cell infiltration.

PPI and druggability evaluation

PPI networks were constructed using the STRING database (https://string-db.org/). To assess the druggability of identified proteins, we searched identified proteins in DrugBank, DGIdb, the ChEMBL and Dependency Map databases [19]. For proteins identified in drug databases, information on the drug name and the process of drug development was documented. To assess the potential druggability, we classified these proteins into four categories: 1) approved; 2) in clinical trials; 3) investigational; 4) experimental.

Statistical analysis

MR analyses applied the Wald ratio (single SNP) and inverse-variance weighted (IVW) (≥ 2 SNPs) methods, complemented by MR-Egger and weighted median approaches to account for pleiotropy and ensure robust causal estimates [20, 21]. The results were presented as odds ratios per standard deviation increase in genetically determined plasma proteins. The above analyses were performed using “coloc”, “TwoSampleMR”, “Seurat”, “SingleR”, and other necessary packages in R (version 4.3.2) [22, 23].

Results

Colocalization analysis

We conducted a colocalization analysis to evaluate whether the observed associations between proteins and breast cancer survival or its subtypes were driven by shared genetic signals (Supplementary Table SI). Seven proteins showed strong colocalization evidence: ARG2, RPL14, and ACBD7 with overall survival; OPCML and DRAXIN with ER- survival; and NFU1 and TXNL4B with ER+ survival. Additionally, 41 proteins demonstrated moderate evidence of colocalization. The remaining proteins showed no evidence of colocalization with breast cancer survival.

SMR and HEIDI tests verified seven causal proteins

To validate the effect of proteins on breast cancer survival, we performed SMR and HEIDI analyses on 7522 proteins using data from four large cohorts. Among the 40 proteins that passed SMR analysis, only 4 of them (LDLRAP1, SKAP1, CSF2, SCLY) failed the HEIDI test (p < 0.05). Among the remaining proteins, 20 were identified as potentially associated with overall breast cancer survival (Supplementary Table SII). Subtype analysis revealed that 7 proteins might be associated with ER- breast cancer survival, while 9 proteins might be associated with ER+ breast cancer survival.

TSMR analysis

We identified 131 significant SNPs as IVs (p < 5 × 10–8), all with F-statistics > 10 (Supplementary Table SIII). Eight proteins did not pass the TSMR analysis (p adj > 0.05) and were excluded from further analysis. Using the Wald ratio or IVW methods with Bonferroni correction, we found that nine proteins (ADAM15, ARG2, CD83, CEP85, GORASP2, HAPLN1, LEFTY2, SH3BGRL3, and SNCG) were associated with improved survival, and five with poorer outcomes (IL36A, PGM1, RPL14, SERPINB5, and UBE2F). In analyses stratified by breast cancer subtype, ALDH2, HAPLN1, MTHFD2, and TXNL4B were associated with improved ER+ survival, whereas MUC16 and NFU1 predicted worse outcomes. For ER- breast cancer, ANXA1 was linked to improved survival, while ALOX15B, CPA2, GRHPR, KLK14, and OPCML were associated with reduced survival (Supplementary Table SIV). These associations were generally consistent in the weighted median and MR-Egger analyses. The results of the four main TSMR methods are shown in Supplementary Table SV. No heterogeneity or horizontal pleiotropy was found (p-value for Cochran’s Q test under the inverse-variance weighted method > 0.05, p-value for Egger intercept > 0.05) (Supplementary Table SVI).

In the external validation phase, we successfully replicated the causal association of GORASP2 and UBE2F with breast cancer survival, as well as MTHFD2 with ER+ breast cancer survival, using eQTLGen data (Figure 2). However, validation was not possible for nine proteins due to data unavailability, and several others did not replicate their causal associations in external datasets. These discrepancies may reflect dataset-specific differences such as sample size, or population characteristics. Colocalization, SMR, and TSMR results are summarized for display in Supplementary Table SVII. Furthermore, we linked genetic effects to protein function and assessed the expression levels of predicted proteins across various tissues using the GTEx database (Supplementary Figure S1).

Figure 2

Forest plot for SMR and TSMR analysis based on 27 proteins

https://www.archivesofmedicalscience.com/f/fulltexts/208246/AMS-22-3-208246-g002_min.jpg

Cell-type specificity expression in breast cancer tissue

To investigate whether the 27 genes exhibited cell type-specific enrichment in breast cancer tissues, we conducted a single-cell expression analysis using single-cell RNA-seq data from the GEO database. Cells were clustered into 19 clusters and subsequently categorized into eight cell types: epithelial cells, cycling cells, T cells, myeloid cells, B cells, plasmablasts, endothelial cells, and mesenchymal cells (Figure 3 A). Figures 3 B and C shows the single-cell expression of these 27 genes in each cluster. Notably, IL6A was not included in the dataset, and neither KLK14 nor OPCML was expressed in any cell population. Eight genes demonstrated cell type-specific enrichment in breast cancer tissues, characterized by an average log2FC > 0.5 and FDR < 0.05 (Figure 3 D).

Figure 3

Single-cell type expression of 27 genes in breast cancer tissues. A – A total of 19 cell clusters and 8 cell types were identified. B, C – Expression of protein-coding genes in each cluster. D – Eight protein-coding genes showed evidence of cell type-specific enrichment at average log2FC > 0.5 and FDR < 0.05

https://www.archivesofmedicalscience.com/f/fulltexts/208246/AMS-22-3-208246-g003_min.jpg

Finally, guided by our colocalization analysis, MR analysis, and single-cell specificity analysis, we categorized the proteins into three distinct target groups, summarized in Supplementary Table SVIII. Tier 1 includes eight proteins that passed all tests (ADAM15, CD83, SH3BGRL3, SNCG, ANXA1, GRHPR, ALDH2, MTHFD2), and Tier 2 includes four proteins (ARG2, RPL14, NFU1, TXNL4B). Proteins lacking single-cell expression or supported by moderate colocalization evidence were assigned to Tier 3. Figure 4 shows the supporting evidence for colocalization between the 27 proteins and the results.

Figure 4

Supporting evidence for colocalization between proteins and outcomes. Circle size indicates the colocalization

p-value for H4 (colocalization analysis), and the color of the circle indicates the classification of the evidence

https://www.archivesofmedicalscience.com/f/fulltexts/208246/AMS-22-3-208246-g004_min.jpg

To explore the potential immunological role of proteins, we used TIMER to analyze the correlation between protein expression levels and tumor-infiltrating immune cell levels. Notably, in BRCA, Tier 1 of proteins showed a significant association with immune infiltration (Supplementary Figure S2). These findings support the hypothesis that proteins may influence the tumor microenvironment by regulating the recruitment or activation of immune cells.

PPI and druggability evaluation of therapeutic targets

PPI analyses revealed limited interactions between identified potentially pathogenic proteins, with only eight proteins interacting (Supplementary Figure S3). Several of these proteins are targeted by existing drugs approved for other indications, suggesting potential for repurposing in breast cancer. For instance, sulfasalazine (targeting CD83), kaempferol (ALOX15B), and eflornithine (ARG2) demonstrate anti-inflammatory or anti-tumor properties. Additionally, cardiovascular agents such as acetylsalicylic acid (ALOX15B) and nitroglycerin (ALDH2) may warrant further investigation. Lastly, hydrocortisone (ANXA1 target) could modulate the tumor microenvironment via metabolic and immune regulation. Further studies and clinical trials would be needed to confirm their applicability for breast cancer. A summary of investigational and approved drugs targeting the identified proteins is provided in Supplementary Table SIX.

Discussion

We systematically examined causal relationships between 17,267 circulating proteins and breast cancer survival using Bayesian colocalization, SMR, HEIDI tests, and TSMR, identifying 27 potential prognostic biomarkers. Six proteins (SH3BGRL3, GRHPR, ARG2, RPL14, NFU1, TXNL4B) were linked to breast cancer survival for the first time. Among these, eight proteins (ADAM15, CD83, SH3BGRL3, SNCG, ANXA1, GRHPR, ALDH2, and MTHFD2) showed the strongest overall survival associations, while four proteins (ARG2, RPL14, NFU1, and TXNL4B) demonstrated strong but slightly less robust associations. Subtype-stratified analyses revealed distinct patterns: ALDH2, HAPLN1, MTHFD2, and TXNL4B were associated with improved survival in ER+ breast cancer, whereas MUC16 and NFU1 were linked to worse prognosis. For ER- breast cancer, ANXA1 correlated with better survival, while ALOX15B, CPA2, GRHPR, KLK14, and OPCML were linked to poorer survival. To further explore clinical potential of these findings, we evaluated druggability and identified 13 proteins with approved or investigational therapeutic agents. These findings provide valuable insights into the molecular mechanisms underlying breast cancer prognosis and suggest potential therapeutic targets for improving patient outcomes.

ADAM15, particularly its isoform ADAM15-C, has been linked to improved survival after lymph node metastasis, likely through its effects on tumor growth and angiogenesis [24]. CD83, highly expressed in mature dendritic cells and axillary lymph nodes, enhances anti-tumor immunity and may serve as a novel prognostic marker for early metastasis [25, 26]. ANXA1, involved in both tumor growth and immune response, was associated with improved survival, possibly by reducing inflammation and activating M1 macrophages [27]. ALDH2, an alcohol metabolism enzyme, showed improved ER+ survival, likely through reduced oxidative stress and enhanced myeloid cell function in the tumor microenvironment [28, 29]. MTHFD2, as a key enzyme in folate metabolism, helps maintain cellular homeostasis, limit harmful senescence-associated effects, and thereby contribute to improved breast cancer prognosis [30]. Unfortunately, the relationship between SNCG expression and prognosis in our MR analysis appears inconsistent with previous studies and may be partly due to potential confounding or intermediate factors [31, 32]. Collectively, these findings not only highlight their potential as prognostic biomarkers but also provide insights into breast cancer progression and immune interactions, offering promising avenues for therapeutic intervention.

In addition to previously recognized prognostic proteins, we also identified several new biomarkers associated with breast cancer survival, including SH3BGRL3, GRHPR, ARG2, RPL14, NFU1, and TXNL4B, with SH3BGRL3 and GRHPR providing the most compelling evidence (Tier 1). SH3BGRL3 significantly correlates with epidermal growth factor receptor (EGFR) expression (p < 0.0001), suggesting involvement in EGFR-mediated oncogenic pathways, making it a promising therapeutic target [33].

GRHPR, a cytoplasmic glyoxylate-metabolizing enzyme, is negatively associated with survival in ER-negative breast cancer, indicating subtype-specific roles [34]. Its species-specific regulation, especially the lack of PPARα control in humans, calls for further study of its metabolic role in breast cancer [35]. Beyond these findings, ARG2, RPL14, NFU1, and TXNL4B also emerged as novel survival-associated proteins, warranting further research to determine their biological functions and potential therapeutic implications. These newly identified biomarkers expand our understanding of breast cancer prognosis, offering new avenues for both biomarker development and therapeutic intervention.

A key strength of this study is the subtype-stratified analysis, which revealed distinct biomarker-survival associations across molecular subtypes and uncovered overlooked prognostic factors. Notably, we observed that not all circular proteins directly influence survival outcomes during tumor development and progression. This underscores the importance of incorporating prognostic data into MR studies and the need for further mechanistic and clinical validation to improve biomarker-based prognostication. Given that gene function can vary by cell type, we leveraged single-cell transcriptomic datasets to examine gene expression patterns at the cellular level. This analysis enhances our understanding of biomarker relevance in breast cancer and supports the development of more precise targeted therapies. Furthermore, TIMER-based immune infiltration analysis showed that these genes correlate with specific immune cells, suggesting that they may shape the tumor microenvironment and have prognostic or therapeutic value.

However, several limitations of this study should also be acknowledged. Firstly, caution is required when interpreting the posterior probability for hypothesis 4 in colocalization analysis (PPH4). A low PPH4 value may not indicate a lack of co-localization evidence if PPH3 is also low due to insufficient power. Secondly, some causal associations failed to replicate in eQTLGen, likely due to differences in population structure, limited power, or tissue origin. Notably, eQTLGen is based on whole blood, which may not capture regulatory effects relevant to breast tissue. Nonetheless, the successful replication of several key proteins, including GORASP2, UBE2F, and MTHFD2, supports the credibility of our main findings. Future studies using larger, multi-tissue eQTL resources are warranted to improve replication accuracy. Thirdly, although our study provides preliminary evidence linking certain drug targets to breast cancer, these associations may be biased if the genetic instruments influence outcomes through pathways other than protein levels. Moreover, the biological mechanisms by which these proteins affect tumor progression remain unclear and require further validation through in vivo and in vitro studies. Unfortunately, the mechanisms by which these proteins affect tumor progression are not yet fully understood and require validation through in vivo and in vitro studies.

In conclusion, in this study, we identified 27 genes encoding proteins associated with overall and subtype-specific breast cancer survival across eight distinct cell types. These genes display varying effect sizes and unique associations with breast cancer, offering promising targets for both screening biomarkers and therapeutic drug development.