❌

Normal view

USP4-Dependent CHAF1B Stabilization Regulates Distinct SETDB1 Ubiquitin States Linked to AKT T308 Signaling and Lipogenic Remodeling in HCC

Adv Sci (Weinh). 2026 Sep 29:e78039. doi: 10.1002/advs.78039. Online ahead of print.

ABSTRACT

Durable responses to current therapies remain limited in hepatocellular carcinoma (HCC), highlighting the need to identify regulators of malignant progression. By integrating multi-omics analyses, spatial transcriptomics, clinical specimens, and multiple models, we identified chromatin assembly factor 1B (CHAF1B) as a functional regulator of HCC phenotypes. Gain- and loss-of-function of CHAF1B altered proliferative, migratory, clonogenic, and tumorigenic phenotypes. LC-MS/MS, DIA proteomics, and cell-based assays revealed CHAF1B-associated lipogenic remodeling characterized by SREBP1C nuclear localization, lipogenic gene/protein induction, and lipid-droplet accumulation. Mechanistically, the WD40 repeat-containing region of CHAF1B contributed to its association with UHRF1 and SETDB1, supporting UHRF1-associated K63-linked ubiquitination and CRM1/exportin-1-dependent cytoplasmic redistribution of SETDB1. Conversely, CHAF1B depletion enhanced SETDB1 association with VHL and favored a predominantly K11-associated degradative ubiquitin state linked to proteasomal SETDB1 loss. SETDB1 redistribution and catalytic activity were associated with AKT T308-linked signaling. A focused CRISPR-based screen of deubiquitinases identified USP4 as an upstream regulator of CHAF1B protein homeostasis. USP4 depletion or Akebia saponin D (ASD) increased K48-linked ubiquitination of CHAF1B, reduced CHAF1B protein abundance, attenuated AKT T308-linked signaling, and suppressed malignant and lipogenic phenotypes. These findings reveal distinct ubiquitin-dependent states governing SETDB1 stability and identify USP4-dependent CHAF1B stabilization as an upstream regulatory node in HCC.

PMID:42811544 | PMC:PMC13624420 | DOI:10.1002/advs.78039

ACTB promotes ESCC progression by regulating the AKT–mTOR signaling pathway through an m<sup>6</sup>A-dependent mechanism

Oncogene, Published online: 30 September 2026; doi:10.1038/s41388-026-03995-3

ACTB promotes ESCC progression by regulating the AKT–mTOR signaling pathway through an m6A-dependent mechanism

Transformer-Based Multitask Framework Integrating Habitat and Deep Learning for Predicting Early Disease Control and Survival in Immunotherapy-Treated Hepatocellular Carcinoma

Adv Sci (Weinh). 2026 Sep 27:e78005. doi: 10.1002/advs.78005. Online ahead of print.

ABSTRACT

Hepatocellular carcinoma (HCC) patients show heterogeneous responses to immune checkpoint inhibitors (ICIs). This study developed ECOS-Net, a transformer-based multitask network integrating CT-derived habitat and 2.5-dimensional (2.5D) deep learning features for simultaneously predicting early disease control (DC) and overall survival (OS). Of 1,234 patients with HCC enrolled from eight institutions and public databases, 832 ICI-treated patients were used for model development. ECOS-Net fused features using multi-head attention and generated early DC probabilities and OS risk scores. ECOS-DC achieved AUCs of 0.836, 0.822, and 0.817 in training, internal validation, and external test sets, outperforming clinical models (all p values < 0.05). ECOS-OS yielded C-indices of 0.730, 0.722, and 0.720, respectively. Integrated models also showed favorable external performance (early DC AUC: 0.825; OS C-index: 0.741). Patients with higher ECOS-DC probabilities had a higher likelihood of early DC, whereas those with higher ECOS-OS risk had shorter OS, with directionally consistent associations across most subgroups. Exploratory biological analyses suggested that the higher ECOS-DC probability and lower ECOS-OS risk groups were associated with immune-active tumor microenvironment features. Therefore, ECOS-Net shows potential as a non-invasive imaging-based risk stratification framework for simultaneously predicting early DC and OS in ICI-treated HCC patients.

PMID:42801546 | PMC:PMC13616327 | DOI:10.1002/advs.78005

Generative AI Assisted Workflows in Architectural Conceptual Design: Performance, Creative Self-Efficacy, and Cognitive Load

arXiv:2601.10696v2 Announce Type: replace Abstract: Generative AI (GenAI) is increasingly adopted in design education, yet evaluating its educational value through final outcomes provides an incomplete picture. This study compares two ecologically plausible workflows in an architectural conceptual design task: GenAI-assisted image generation and ArchDaily-based precedent search. The comparison concerns complete workflows rather than the isolated contributions. Thirty-six students completed a two-phase design task, first designing independently and then revising with their assigned workflow. Eight judges rated design performance, while participants reported task-specific and general creative self-efficacy and cognitive load after each phase. Difference-in-differences analyses showed no significant overall differences between the GenAI and precedent-search workflows in design performance, cognitive workload, or task-specific creative self-efficacy. Beyond these null overall effects, three patterns were observed. General creative self-efficacy showed a significant relative decline under the GenAI workflow. A subgroup analysis suggested higher revision-phase performance among novice students using GenAI than among those using precedent search (F (1,32) = 4.303, p = 0.046). However, this exploratory interaction should be interpreted cautiously due to low rating reliability, small subgroup cells, and imprecise estimation. Third, exploratory prompt analyses suggested that iterative, task-specific prompting strategies (CD3, CD6) were associated with cognitive load reductions at the uncorrected level, but neither association survived multiple-comparison correction. Overall, the GenAI workflow did not produce uniform gains. Its educational value may depend on pedagogical framing, learner characteristics, and human-AI interaction structure, underscoring the need to preserve creative agency and develop prompt literacy.

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

npj Digital Medicine, Published online: 12 September 2026; doi:10.1038/s41746-026-03238-5

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

Position: Science of AI Evaluation Requires Item-level Benchmark Data

arXiv:2604.03244v1 Announce Type: new Abstract: AI evaluations have become the primary evidence for deploying generative AI systems across high-stakes domains. However, current evaluation paradigms often exhibit systemic validity failures. These issues, ranging from unjustified design choices to misaligned metrics, remain intractable without a principled framework for gathering validity evidence and conducting granular diagnostic analysis. In this position paper, we argue that item-level AI benchmark data is essential for establishing a rigorous science of AI evaluation. Item-level analysis enables fine-grained diagnostics and principled validation of benchmarks. We substantiate this position by dissecting current validity failures and revisiting evaluation paradigms across computer science and psychometrics. Through illustrative analyses of item properties and latent constructs, we demonstrate the unique insights afforded by item-level data. To catalyze community-wide adoption, we introduce OpenEval, a growing repository of item-level benchmark data designed supporting evidence-centered AI evaluation.

Hypoxia-related and immune phenotype-related fusion model for non-invasive prognostication of hepatocellular carcinoma treated by TACE: a multicentre study

Gut. 2026 Mar 30:gutjnl-2025-337938. doi: 10.1136/gutjnl-2025-337938. Online ahead of print.

ABSTRACT

BACKGROUND: Survival outcomes after transarterial chemoembolisation (TACE) vary in hepatocellular carcinoma (HCC) patients, and existing prognostic scores and imaging models often lack generalisability and biological interpretability.

OBJECTIVE: To develop and validate a multimodal prognostication model for HCC that allows for a precise assessment of survival outcomes of HCC patients receiving TACE therapy.

DESIGN: This study enrolled 1448 HCC patients, including a TACE cohort (n=1349), a biomarker subset from a randomised trial (n=41), a single-cell RNA sequencing cohort and The Cancer Genome Atlas (TCGA) HCC cohort (n=50). Pre-treatment contrast-enhanced CT images were used to construct deep learning and conventional radiomic models. The early-fusion and late-fusion models (LFMs) were compared, and a clinical-radiologic model (CRM) was formed by integrating the better-performing LFM with clinical variables. Using TCGA data and single-cell transcriptomic profiles, the differences between high-score and low-score groups in tumour immune microenvironment, cellular functional states and key signalling pathways were investigated.

RESULTS: The CRM effectively stratified patients' survival across multiple independent cohorts and achieved more granular risk stratification than the existing clinical models. Multi-omic analyses revealed that in the LFM high-score group, myelocytomatosis oncogene was activated, epithelial-mesenchymal transition enhanced, glycolysis upregulated and hypoxia pathway activated. Single-cell transcriptomic data confirmed that virtually all cell types in high-risk patients scored high in hypoxia, and cytotoxic T cells had a reduced cytotoxic activity.

CONCLUSION: The CRM model can non-invasively predict the prognosis of HCC patients treated by TACE therapy.

PMID:41856522 | DOI:10.1136/gutjnl-2025-337938

Divergent tumor immunity determined by bacteria-cancer cell engagement

In a preclinical breast cancer metastasis model, the same bacteria strain, when present intracellularly versus extracellularly, exerts opposing effects on tumor immunity by inducing divergent neutrophil states, highlighting the intricacy in bacterial-host engagement for shaping tumor immunity.

daVinci-Env: Open SWE Environment Synthesis at Scale

arXiv:2603.13023v1 Announce Type: cross Abstract: Training capable software engineering (SWE) agents demands large-scale, executable, and verifiable environments that provide dynamic feedback loops for iterative code editing, test execution, and solution refinement. However, existing open-source datasets remain limited in scale and repository diversity, while industrial solutions are opaque with unreleased infrastructure, creating a prohibitive barrier for most academic research groups. We present OpenSWE, the largest fully transparent framework for SWE agent training in Python, comprising 45,320 executable Docker environments spanning over 12.8k repositories, with all Dockerfiles, evaluation scripts, and infrastructure fully open-sourced for reproducibility. OpenSWE is built through a multi-agent synthesis pipeline deployed across a 64-node distributed cluster, automating repository exploration, Dockerfile construction, evaluation script generation, and iterative test analysis. Beyond scale, we propose a quality-centric filtering pipeline that characterizes the inherent difficulty of each environment, filtering out instances that are either unsolvable or insufficiently challenging and retaining only those that maximize learning efficiency. With $891K spent on environment construction and an additional $576K on trajectory sampling and difficulty-aware curation, the entire project represents a total investment of approximately $1.47 million, yielding about 13,000 curated trajectories from roughly 9,000 quality guaranteed environments. Extensive experiments validate OpenSWE's effectiveness: OpenSWE-32B and OpenSWE-72B achieve 62.4% and 66.0% on SWE-bench Verified, establishing SOTA among Qwen2.5 series. Moreover, SWE-focused training yields substantial out-of-domain improvements, including up to 12 points on mathematical reasoning and 5 points on science benchmarks, without degrading factual recall.

MindfulAgents: Personalizing Mindfulness Meditation via an Expert-Aligned Multi-Agent System

arXiv:2603.06926v1 Announce Type: cross Abstract: Mindfulness meditation is a widely accessible and evidence-based method for supporting mental health. Despite the proliferation of mindfulness meditation apps, sustaining user engagement remains a persistent challenge. Personalizing the meditation experience is a promising strategy to improve engagement, but it often requires costly and unscalable manual effort. We present MindfulAgents, a multi-agent system powered by large language models that (1) generates guided meditation scripts based on an expert-established mindfulness framework, (2) encourages users' reflection on emotional states and mindfulness skills, and (3) enables real-time personalization of the mindfulness meditation experience for each user. In a formative lab study (N=13), MindfulAgents significantly improved in-session engagement (p = 0.011) and self-awareness (p = 0.014), and reduced momentary stress (p = 0.020). Furthermore, a four-week deployment study (N=62) demonstrated a notable increase in long-term engagement (p = 0.002) and level of mindfulness (p = 0.023). Participants reported that MindfulAgents offered more relevant meditation sessions personalized to individual needs in various contexts, supporting sustained practice. Our findings highlight the potential of LLM-driven personalization for enhancing user engagement in digital mindfulness meditation interventions.

Preference Leakage: A Contamination Problem in LLM-as-a-judge

arXiv:2502.01534v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annotation methods in model development. While their combination significantly enhances the efficiency of model training and evaluation, little attention has been given to the potential contamination brought by this new model development paradigm. In this work, we expose preference leakage, a contamination problem in LLM-as-a-judge caused by the relatedness between the synthetic data generators and LLM-based evaluators. To study this issue, we first define three common relatednesses between the data generator LLM and the judge LLM: being the same model, having an inheritance relationship, and belonging to the same model family. Through extensive experiments, we empirically confirm the bias of judges towards their related student models caused by preference leakage across multiple LLM baselines and benchmarks. Further analysis suggests that preference leakage is a pervasive and real-world problem that is harder to detect compared to previously identified biases in LLM-as-a-judge scenarios. All of these findings imply that preference leakage is a widespread and challenging problem in the area of LLM-as-a-judge. We release all codes and data at: https://github.com/David-Li0406/Preference-Leakage.
❌