❌

Reading view

Machine learning-based identification of key genes underlying sex differences in hepatocellular carcinoma and targeted drug screening

Biomed Rep. 2026 Apr 24;24(6):74. doi: 10.3892/br.2026.2147. eCollection 2026 Jun.

ABSTRACT

Hepatocellular carcinoma (HCC) shows a marked predominance in men, yet the molecular basis for this sex disparity remains unclear. The present study leveraged multi-omics data and machine learning algorithms to identify key genes associated with sex-specific differences in HCC and to screen for putative candidate compounds, aiming to provide new insights for sex-specific therapy. The mRNA expression data of male and female patients with HCC and paracancerous tissues were obtained from the GEO and TCGA databases. To mitigate overfitting, data were partitioned into independent training and testing sets. Candidate genes were screened by differential expression analysis and weighted gene co-expression network analysis. A total of four complementary algorithms, random forest, support vector machines, generalized linear models and extreme gradient boosting were used to identify key genes with high predictive capability. CYP17A1 and IRX3 were identified as the top differentially expressed core genes associated with HCC in men. Pan-cancer analysis showed that CYP17A1 was lowly expressed in the majority of tumors, but significantly highly expressed in HCC, rectal adenocarcinoma and gastric cancer (P<0.001). Functional cell-based assays showed that knockout of CYP17A1 inhibited the proliferation, migration and invasion ability of HCC cells (P<0.001). Immunohistochemistry showed that CYP17A1 protein expression was significantly increased in HCC tissues from male patients when compared with that in paracancerous tissues (P<0.001), whereas there was no significant difference in female patient tissues (P>0.05). Notably, while IRX3 was identified computationally, its functional role remains to be experimentally validated. Molecular docking predicted a potential interaction between the natural compound Saikosaponin A and the CYP17A1 protein, and cellular assays revealed that it dose-dependently inhibits HCC cell malignant phenotypes. The present study suggests that CYP17A1 is associated with sex differences in HCC, potentially via the androgen signaling axis. Furthermore, IRX3 emerges as a novel hypothesis-generating candidate gene. Finally, the findings of the present study highlight Saikosaponin A as a putative therapeutic candidate for male patients with HCC, warranting further target-dependency investigations.

PMID:42125766 | PMC:PMC13158723 | DOI:10.3892/br.2026.2147

  •  

Spatial multi-omics unveils the monoclonal origin, neuroendocrine plasticity, and microenvironment niches in combined small-cell lung cancer

Cell Rep Med. 2026 Apr 10:102741. doi: 10.1016/j.xcrm.2026.102741. Online ahead of print.

ABSTRACT

Combined small-cell lung cancer (cSCLC) is an aggressive subtype of SCLC with mixed histologic components. Despite heterogeneity and poorer prognosis than de novo SCLC, cSCLC is managed as SCLC because molecular insight into biology, lineage plasticity, and tumor microenvironment (TME) is limited. We perform spatial whole-exome sequencing, spatial transcriptomics, and single-nucleus RNA sequencing across 19 treatment-naive cSCLC tumors. Different histologic components share a monoclonal origin, whereas divergence associates with distinct mutation and copy-number alteration patterns. Our results define spatially exclusive or interspersed tumor domains with distinct TME and immune landscapes; fibroblast-rich boundaries enriched for an aggressive fibroblast subtype may shape TME and treatment responses. We identify lineage plasticity, including adenocarcinoma-to-SCLC transdifferentiation and SCLC-subtype coexistence, and develop cSCLC Detector, a sensitive mutation-based assay improving cSCLC detection in tissue and liquid biopsies. These findings illuminate cSCLC evolution and heterogeneity, underscoring the need for tailored diagnostic and therapeutic strategies for this aggressive subtype.

PMID:41966692 | DOI:10.1016/j.xcrm.2026.102741

  •  

Spatial multi-omics unveils the monoclonal origin, neuroendocrine plasticity, and microenvironment niches in combined small-cell lung cancer

Cell Rep Med. 2026 Apr 10:102741. doi: 10.1016/j.xcrm.2026.102741. Online ahead of print.

ABSTRACT

Combined small-cell lung cancer (cSCLC) is an aggressive subtype of SCLC with mixed histologic components. Despite heterogeneity and poorer prognosis than de novo SCLC, cSCLC is managed as SCLC because molecular insight into biology, lineage plasticity, and tumor microenvironment (TME) is limited. We perform spatial whole-exome sequencing, spatial transcriptomics, and single-nucleus RNA sequencing across 19 treatment-naive cSCLC tumors. Different histologic components share a monoclonal origin, whereas divergence associates with distinct mutation and copy-number alteration patterns. Our results define spatially exclusive or interspersed tumor domains with distinct TME and immune landscapes; fibroblast-rich boundaries enriched for an aggressive fibroblast subtype may shape TME and treatment responses. We identify lineage plasticity, including adenocarcinoma-to-SCLC transdifferentiation and SCLC-subtype coexistence, and develop cSCLC Detector, a sensitive mutation-based assay improving cSCLC detection in tissue and liquid biopsies. These findings illuminate cSCLC evolution and heterogeneity, underscoring the need for tailored diagnostic and therapeutic strategies for this aggressive subtype.

PMID:41966692 | DOI:10.1016/j.xcrm.2026.102741

  •  

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

arXiv:2604.01840v1 Announce Type: new Abstract: While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks suffer from a foundational methodological flaw: by distributing identical advantages across all generated tokens, these methods inherently dilute the learning signals essential for optimizing the critical, visually-grounded steps of multimodal reasoning. To bridge this gap, we formulate \textit{Token Visual Dependency}, quantifying the causal information gain of visual inputs via the Kullback-Leibler (KL) divergence between visual-conditioned and text-only predictive distributions. Revealing that this dependency is highly sparse and semantically pivotal, we introduce Perception-Grounded Policy Optimization (PGPO), which is a novel fine-grained credit assignment framework that dynamically reshapes advantages at the token level. Through a threshold-gated, mass-conserving mechanism, PGPO actively amplifies learning signals for visually-dependent tokens while suppressing gradient noise from linguistic priors. Extensive experiments based on the Qwen2.5-VL series across seven challenging multimodal reasoning benchmarks demonstrate that PGPO boosts models by 18.7% on average. Both theoretical and empirical analyses confirm that PGPO effectively reduces gradient variance, prevents training collapse, and acts as a potent regularizer for robust, perception-grounded multimodal reasoning. Code will be published on https://github.com/Yzk1114/PGPO.
  •  

Semantic Refinement with LLMs for Graph Representations

arXiv:2512.21106v2 Announce Type: replace-cross Abstract: Graph-structured data exhibit substantial heterogeneity in where their predictive signals originate: in some domains, node-level semantics dominate, while in others, structural patterns play a central role. This structure-semantics heterogeneity implies that no graph learning model with a fixed inductive bias can generalize optimally across diverse graph domains. However, most existing methods address this challenge from the model side by incrementally injecting new inductive biases, which remains fundamentally limited given the open-ended diversity of real-world graphs. In this work, we take a data-centric perspective and treat node semantics as a task-adaptive variable. We propose a Graph-Exemplar-guided Semantic Refinement (GES) framework for graph representation learning which -- unlike existing LLM-enhanced methods that generate node descriptions without graph context -- leverages structurally and semantically similar nodes from the graph itself to guide semantic refinement. Specifically, a GNN is first trained to produce predictive states, which along with structural and semantic similarity are used to retrieve in-graph exemplars that inform an LLM in refining node descriptions. We evaluate our approach on both text-rich and text-free graphs. Results show consistent improvements on semantics-rich and structure-dominated graphs, demonstrating the effectiveness of data-centric semantic refinement under structure-semantics heterogeneity.
  •  

A Very Big Video Reasoning Suite

arXiv:2602.20159v1 Announce Type: cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/ .
  •  
❌