❌

Normal view

Machine learning-based identification of key genes underlying sex differences in hepatocellular carcinoma and targeted drug screening

Biomed Rep. 2026 Apr 24;24(6):74. doi: 10.3892/br.2026.2147. eCollection 2026 Jun.

ABSTRACT

Hepatocellular carcinoma (HCC) shows a marked predominance in men, yet the molecular basis for this sex disparity remains unclear. The present study leveraged multi-omics data and machine learning algorithms to identify key genes associated with sex-specific differences in HCC and to screen for putative candidate compounds, aiming to provide new insights for sex-specific therapy. The mRNA expression data of male and female patients with HCC and paracancerous tissues were obtained from the GEO and TCGA databases. To mitigate overfitting, data were partitioned into independent training and testing sets. Candidate genes were screened by differential expression analysis and weighted gene co-expression network analysis. A total of four complementary algorithms, random forest, support vector machines, generalized linear models and extreme gradient boosting were used to identify key genes with high predictive capability. CYP17A1 and IRX3 were identified as the top differentially expressed core genes associated with HCC in men. Pan-cancer analysis showed that CYP17A1 was lowly expressed in the majority of tumors, but significantly highly expressed in HCC, rectal adenocarcinoma and gastric cancer (P<0.001). Functional cell-based assays showed that knockout of CYP17A1 inhibited the proliferation, migration and invasion ability of HCC cells (P<0.001). Immunohistochemistry showed that CYP17A1 protein expression was significantly increased in HCC tissues from male patients when compared with that in paracancerous tissues (P<0.001), whereas there was no significant difference in female patient tissues (P>0.05). Notably, while IRX3 was identified computationally, its functional role remains to be experimentally validated. Molecular docking predicted a potential interaction between the natural compound Saikosaponin A and the CYP17A1 protein, and cellular assays revealed that it dose-dependently inhibits HCC cell malignant phenotypes. The present study suggests that CYP17A1 is associated with sex differences in HCC, potentially via the androgen signaling axis. Furthermore, IRX3 emerges as a novel hypothesis-generating candidate gene. Finally, the findings of the present study highlight Saikosaponin A as a putative therapeutic candidate for male patients with HCC, warranting further target-dependency investigations.

PMID:42125766 | PMC:PMC13158723 | DOI:10.3892/br.2026.2147

Predictive Value of Machine Learning for Poststroke Mortality Risk: Systematic Review and Meta-Analysis

Background: People with stroke face a high mortality risk, and an accurate prediction model is essential to the guidance of clinical decision-making in this population. Recently, with growing attention paid to machine learning (ML) in stroke care, some researchers have investigated the effectiveness of ML in predicting the mortality risk in stroke. However, systematic evidence is still lacking for its effectiveness. Objective: This systematic review aims to evaluate the value of ML in predicting the stroke mortality risk. The findings are expected to offer an evidence-based basis for developing and assessing clinical risk prediction tools. Methods: A search was made in Cochrane Library, PubMed, Embase, and Web of Science up to June 23, 2025, and studies that reported a complete performance of ML in predicting stroke mortality were included. Studies with only risk factors analyzed were excluded. The risk of bias of the included studies was assessed using PROBAST (Prediction model Risk of Bias Assessment Tool). Pooled risk ratios with 95% CIs and prediction intervals (PIs) were derived using the Hartung-Knapp-Sidik-Jonkman method under a random-effects model. Subgroup analyses were also conducted by model type, stroke type, patient source, and treatment background. Moreover, a metaregression was conducted on the C-index for out-of-hospital mortality at different time points to explore the influence of time factors on the model’s predictive performance. Results: Sixty-eight studies were included (23 predicting in-hospital mortality and 45 predicting out-of-hospital mortality), describing the development of 75 prediction models and 43 external validations. The follow-up period was 1 month to 15 years. For predicting in-hospital mortality, the external validation set had a pooled C-index of 0.727 (95% CI 0.677-0.781, 95% PI 0.521-1.000), with sensitivity and specificity of 0.64 (95% CI 0.57-0.70) and 0.74 (95% CI 0.70-0.77), respectively. For predicting out-of-hospital mortality, the pooled C-index was 0.847 (95% CI 0.808-0.887, 95% PI 0.750-0.956) in the external validation set, with sensitivity and specificity of 0.71 (95% CI 0.55-0.82) and 0.76 (95% CI 0.74-0.78), respectively. Comparatively, the overall pooled C-indexes were 0.788 (95% CI 0.766-0.810, 95% PI 0.621-0.999) and 0.812 (95% CI 0.798-0.826, 95% PI 0.693-0.952), respectively. The metaregression revealed a gradual decline in the predictive performance of the overall model and logistic regression model alone, whereas a random forest model maintained sustained performance. Age, National Institutes of Health Stroke Scale score, and stroke-related complications were the most frequently used variables for modeling. Conclusions: This is the first meta-analysis to demonstrate that ML-based prediction of stroke mortality is feasible. The performance of ML supports its role as an auxiliary tool for identifying high-risk populations, thereby optimizing clinical monitoring and resource allocation. However, due to substantial heterogeneity and a relatively high risk of bias in available studies, caution is warranted in real-world application. The effectiveness of ML may vary across settings, and external validation is recommended before broader implementation. Trial Registration: PROSPERO CRD420251086321; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251086321

ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search

arXiv:2509.23519v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models by grounding their outputs in external documents. These systems, however, remain vulnerable to attacks on the retrieval corpus, such as prompt injection. RAG-based search systems (e.g., Google's Search AI Overview) present an interesting setting for studying and protecting against such threats, as defense algorithms can benefit from built-in reliability signals -- like document ranking -- and represent a non-LLM challenge for the adversary due to decades of work to thwart SEO. Motivated by, but not limited to, this scenario, this work introduces ReliabilityRAG, a framework for adversarial robustness that explicitly leverages reliability information of retrieved documents. Our first contribution adopts a graph-theoretic perspective to identify a "consistent majority" among retrieved documents to filter out malicious ones. We introduce a novel algorithm based on finding a Maximum Independent Set (MIS) on a document graph where edges encode contradiction. Our MIS variant explicitly prioritizes higher-reliability documents and provides provable robustness guarantees against bounded adversarial corruption under natural assumptions. Recognizing the computational cost of exact MIS for large retrieval sets, our second contribution is a scalable weighted sample and aggregate framework. It explicitly utilizes reliability information, preserving some robustness guarantees while efficiently handling many documents. We present empirical results showing ReliabilityRAG provides superior robustness against adversarial attacks compared to prior methods, maintains high benign accuracy, and excels in long-form generation tasks where prior robustness-focused methods struggled. Our work is a significant step towards more effective, provably robust defenses against retrieved corpus corruption in RAG.
❌