❌

Normal view

Metabolic dysregulation: Its role in diabetes mellitus and cancers

Mol Aspects Med. 2026 Feb 23;108:101461. doi: 10.1016/j.mam.2026.101461. Online ahead of print.

ABSTRACT

Diabetes mellitus (DM) is a significant risk factor for several cancers, particularly cancers of the liver, pancreas, and endometrium. This review aims to understand the connections between diabetic pathophysiology and cancer biology. We synthesize how core metabolic disturbances-hyperinsulinemia, hyperglycemia, and inflammation-promote tumorigenesis by dysregulating canonical oncogenic pathways such as IGF-1 signaling, DNA damage repair, and immunometabolism. Subsequently, we focus on how key molecular integrators-such as p38 MAPK, Wnt/β-catenin, and the AGEs-RAGE axis-mediate metabolic stress to confer proliferative and invasive advantages to tumor cells. However, a direct translational application of these mechanisms, particularly in the context of repurposing antidiabetic drugs for cancer therapy, remains challenging due to inconsistent clinical outcomes. To address this gap, we suggest that a fundamental shift in approach is required. We propose that future research must move beyond simple pathway categorization and instead utilize spatial analysis techniques to reveal how diabetic metabolites reshape the tumor microenvironment (TME). By integrating single-cell and spatial omics technologies, the field can begin to map the precise cellular niches within tumors where diabetic metabolites exacerbate malignant progression and foster treatment resistance. This perspective is essential for developing targeted strategies to mitigate cancer risk and improve outcomes for the expanding population of patients with DM and cancer.

PMID:41734405 | DOI:10.1016/j.mam.2026.101461

Multi-objective fluorescent molecule design with a data-physics dual-driven generative framework

arXiv:2601.13564v1 Announce Type: cross Abstract: Designing fluorescent small molecules with tailored optical and physicochemical properties requires navigating vast, underexplored chemical space while satisfying multiple objectives and constraints. Conventional generate-score-screen approaches become impractical under such realistic design specifications, owing to their low search efficiency, unreliable generalizability of machine-learning prediction, and the prohibitive cost of quantum chemical calculation. Here we present LUMOS, a data-and-physics driven framework for inverse design of fluorescent molecules. LUMOS couples generator and predictor within a shared latent representation, enabling direct specification-to-molecule design and efficient exploration. Moreover, LUMOS combines neural networks with a fast time-dependent density functional theory (TD-DFT) calculation workflow to build a suite of complementary predictors spanning different trade-offs in speed, accuracy, and generalizability, enabling reliable property prediction across diverse scenarios. Finally, LUMOS employs a property-guided diffusion model integrated with multi-objective evolutionary algorithms, enabling de novo design and molecular optimization under multiple objectives and constraints. Across comprehensive benchmarks, LUMOS consistently outperforms baseline models in terms of accuracy, generalizability and physical plausibility for fluorescence property prediction, and demonstrates superior performance in multi-objective scaffold- and fragment-level molecular optimization. Further validation using TD-DFT and molecular dynamics (MD) simulations demonstrates that LUMOS can generate valid fluorophores that meet various target specifications. Overall, these results establish LUMOS as a data-physics dual-driven framework for general fluorophore inverse design.

Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?

arXiv:2512.15769v1 Announce Type: cross Abstract: The increasing use of generative models such as diffusion models for synthetic data augmentation has greatly reduced the cost of data collection and labeling in downstream perception tasks. However, this new data source paradigm may introduce important security concerns. This work investigates backdoor propagation in such emerging generative data supply chains, namely Data-Chain Backdoor (DCB). Specifically, we find that open-source diffusion models can become hidden carriers of backdoors. Their strong distribution-fitting ability causes them to memorize and reproduce backdoor triggers during generation, which are subsequently inherited by downstream models, resulting in severe security risks. This threat is particularly concerning under clean-label attack scenarios, as it remains effective while having negligible impact on the utility of the synthetic data. Furthermore, we discover an Early-Stage Trigger Manifestation (ESTM) phenomenon: backdoor trigger patterns tend to surface more explicitly in the early, high-noise stages of the diffusion model's reverse generation process before being subtly integrated into the final samples. Overall, this work reveals a previously underexplored threat in generative data pipelines and provides initial insights toward mitigating backdoor risks in synthetic data generation.

AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential

npj Digital Medicine, Published online: 11 December 2025; doi:10.1038/s41746-025-02198-6

AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential

KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening

arXiv:2512.05994v1 Announce Type: cross Abstract: With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.

AI Deception: Risks, Dynamics, and Controls

arXiv:2511.22619v2 Announce Type: replace Abstract: As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.

MedCondDiff: Lightweight, Robust, Semantically Guided Diffusion for Medical Image Segmentation

arXiv:2512.00350v1 Announce Type: cross Abstract: We introduce MedCondDiff, a diffusion-based framework for multi-organ medical image segmentation that is efficient and anatomically grounded. The model conditions the denoising process on semantic priors extracted by a Pyramid Vision Transformer (PVT) backbone, yielding a semantically guided and lightweight diffusion architecture. This design improves robustness while reducing both inference time and VRAM usage compared to conventional diffusion models. Experiments on multi-organ, multi-modality datasets demonstrate that MedCondDiff delivers competitive performance across anatomical regions and imaging modalities, underscoring the potential of semantically guided diffusion models as an effective class of architectures for medical imaging tasks.

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

arXiv:2502.07299v3 Announce Type: replace-cross Abstract: The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-trained models have achieved great success in analyzing these macromolecules individually, their interconnected nature remains underexplored. This paper follows the guidance of the central dogma to redesign both the data and model pipeline and offers a comprehensive framework, Life-Code, that spans different biological functions. As for data flow, we propose a unified pipeline to integrate multi-omics data by reverse-transcribing RNA and reverse-translating amino acids into nucleotide-based sequences. As for the model, we design a codon tokenizer and a hybrid long-sequence architecture to encode the interactions between coding and non-coding regions through masked modeling pre-training. To model the translation and folding process with coding sequences, Life-Code learns protein structures of the corresponding amino acids by knowledge distillation from off-the-shelf protein language models. Such designs enable Life-Code to capture complex interactions within genetic sequences, providing a more comprehensive understanding of multi-omics with the central dogma. Extensive experiments show that Life-Code achieves state-of-the-art results on various tasks across three omics, highlighting its potential for advancing multi-omics analysis and interpretation.

Rethinking Bias in Generative Data Augmentation for Medical AI: a Frequency Recalibration Method

arXiv:2511.12301v1 Announce Type: cross Abstract: Developing Medical AI relies on large datasets and easily suffers from data scarcity. Generative data augmentation (GDA) using AI generative models offers a solution to synthesize realistic medical images. However, the bias in GDA is often underestimated in medical domains, with concerns about the risk of introducing detrimental features generated by AI and harming downstream tasks. This paper identifies the frequency misalignment between real and synthesized images as one of the key factors underlying unreliable GDA and proposes the Frequency Recalibration (FreRec) method to reduce the frequency distributional discrepancy and thus improve GDA. FreRec involves (1) Statistical High-frequency Replacement (SHR) to roughly align high-frequency components and (2) Reconstructive High-frequency Mapping (RHM) to enhance image quality and reconstruct high-frequency details. Extensive experiments were conducted in various medical datasets, including brain MRIs, chest X-rays, and fundus images. The results show that FreRec significantly improves downstream medical image classification performance compared to uncalibrated AI-synthesized samples. FreRec is a standalone post-processing step that is compatible with any generative model and can integrate seamlessly with common medical GDA pipelines.

From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models

arXiv:2511.10788v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made reasoning a central benchmark for evaluating intelligence. While prior surveys focus on efficiency by examining how to shorten reasoning chains or reduce computation, this view overlooks a fundamental challenge: current LLMs apply uniform reasoning strategies regardless of task complexity, generating long traces for trivial problems while failing to extend reasoning for difficult tasks. This survey reframes reasoning through the lens of {adaptivity}: the capability to allocate reasoning effort based on input characteristics such as difficulty and uncertainty. We make three contributions. First, we formalize deductive, inductive, and abductive reasoning within the LLM context, connecting these classical cognitive paradigms with their algorithmic realizations. Second, we formalize adaptive reasoning as a control-augmented policy optimization problem balancing task performance with computational cost, distinguishing learned policies from inference-time control mechanisms. Third, we propose a systematic taxonomy organizing existing methods into training-based approaches that internalize adaptivity through reinforcement learning, supervised fine-tuning, and learned controllers, and training-free approaches that achieve adaptivity through prompt conditioning, feedback-driven halting, and modular composition. This framework clarifies how different mechanisms realize adaptive reasoning in practice and enables systematic comparison across diverse strategies. We conclude by identifying open challenges in self-evaluation, meta-reasoning, and human-aligned reasoning control.

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

arXiv:2510.19755v3 Announce Type: replace-cross Abstract: Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existing acceleration techniques have made progress, they still face challenges such as limited applicability, high training costs, or quality degradation. Against this backdrop, \textbf{Diffusion Caching} offers a promising training-free, architecture-agnostic, and efficient inference paradigm. Its core mechanism identifies and reuses intrinsic computational redundancies in the diffusion process. By enabling feature-level cross-step reuse and inter-layer scheduling, it reduces computation without modifying model parameters. This paper systematically reviews the theoretical foundations and evolution of Diffusion Caching and proposes a unified framework for its classification and analysis. Through comparative analysis of representative methods, we show that Diffusion Caching evolves from \textit{static reuse} to \textit{dynamic prediction}. This trend enhances caching flexibility across diverse tasks and enables integration with other acceleration techniques such as sampling optimization and model distillation, paving the way for a unified, efficient inference framework for future multimodal and interactive applications. We argue that this paradigm will become a key enabler of real-time and efficient generative AI, injecting new vitality into both theory and practice of \textit{Efficient Generative Intelligence}.

BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation

arXiv:2509.24204v2 Announce Type: replace-cross Abstract: Vision foundation models like the Segment Anything Model (SAM), pretrained on large-scale natural image datasets, often struggle in medical image segmentation due to a lack of domain-specific adaptation. In clinical practice, fine-tuning such models efficiently for medical downstream tasks with minimal resource demands, while maintaining strong performance, is challenging. To address these issues, we propose BALR-SAM, a boundary-aware low-rank adaptation framework that enhances SAM for medical imaging. It combines three tailored components: (1) a Complementary Detail Enhancement Network (CDEN) using depthwise separable convolutions and multi-scale fusion to capture boundary-sensitive features essential for accurate segmentation; (2) low-rank adapters integrated into SAM's Vision Transformer blocks to optimize feature representation and attention for medical contexts, while simultaneously significantly reducing the parameter space; and (3) a low-rank tensor attention mechanism in the mask decoder, cutting memory usage by 75% and boosting inference speed. Experiments on standard medical segmentation datasets show that BALR-SAM, without requiring prompts, outperforms several state-of-the-art (SOTA) methods, including fully fine-tuned MedSAM, while updating just 1.8% (11.7M) of its parameters.

Prospective proteomics for discovering biomarkers in lung adenocarcinoma: a literature review

Transl Cancer Res. 2025 Sep 30;14(9):6102-6117. doi: 10.21037/tcr-2025-1092. Epub 2025 Sep 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung adenocarcinoma (LUAD), as the main subtype of non-small cell lung cancer (NSCLC), faces clinical challenges including molecular heterogeneity, late diagnosis, and aggressive growth, leading to a low 5-year survival rate. Biomarkers are critical for early detection, accurate differentiation of benign/malignant lesions, and guiding personalized treatment strategies. Proteomic technologies using liquid biopsy show potential by analyzing protein changes and post-translational modifications (PTMs) to identify novel biomarkers and unravel cancer mechanisms. This review examines proteomic advances in LUAD, compares platform strengths, lists validated protein markers, and discusses challenges like specificity and regulations. It aims to develop a precision medicine framework by integrating multi-omics data for improved diagnosis and treatment.

METHODS: This study conducted a literature review by searching the PubMed and Web of Science databases for original articles written in English from 2002 to 2025, using the keywords "lung adenocarcinoma" OR "LUAD" AND "biomarkers" AND "proteomics" OR "SomaScan" OR "spatial proteomics" to identify the latest research findings in the field of proteomics technology and LUAD biomarkers. The included studies mainly focused on the current landscape of biomarkers in the diagnosis, treatment, and prognosis of LUAD.

KEY CONTENT AND FINDINGS: This review discusses high-throughput methods for comprehensive protein profiling in accessible biospecimens (tissues, blood, urine) to identify biomarkers for LUAD. We systematically evaluate emerging proteomic strategies, including mass spectrometry (MS), proximity extension assays (PEAs), spatial proteomics techniques, and SomaScan platforms-coupled with innovative computational frameworks have revolutionized biomarkers discovery and their translational potential in developing precision diagnostics and targeted therapies. Additionally, the review addresses challenges in integrating proteomics with genomics, transcriptomics, and metabolomics, offering new methodologies and expanding research in life sciences. As technological advancements continue, it is anticipated that more potential biomarkers will be conducted to validate the broader application in LUAD treatment, addressing early-stage disease complexities and aiding in selecting more effective treatment strategies.

CONCLUSIONS: By synthesizing cutting-edge evidence on proteome-driven LUAD biomarkers, this review elucidates actionable strategies to refine early detection protocols and mechanism-informed personalized treatment frameworks, directly advancing precision oncology initiatives for this prevalent malignancy through biomarker-guided clinical decision-making and multi-omics integration.

PMID:41158224 | PMC:PMC12554480 | DOI:10.21037/tcr-2025-1092

Multi-omics analyses inform mechanisms of immunotherapy response in pancreatic cancer

Front Immunol. 2025 Oct 2;16:1673098. doi: 10.3389/fimmu.2025.1673098. eCollection 2025.

ABSTRACT

INTRODUCTION: Pancreatic ductal adenocarcinoma (PDAC) continues to exhibit resistance to immunotherapy. In this study, we evaluated the efficacy of combining immunotherapy with chemotherapy for the treatment of advanced pancreatic cancer. Additionally, we employed a multimodal analytical approach to elucidate the immune landscape and conduct transcriptomic profiling in PDAC.

METHODS: A retrospective analysis was conducted on the clinical data of 52 patients diagnosed with advanced PDAC who underwent a combined treatment regimen of immunotherapy and chemotherapy. The study evaluated the objective response rate (ORR), disease control rate (DCR), and progression-free survival (PFS). To characterize the immune landscape in treatment-naive pancreatic ductal adenocarcinoma (PDAC) tumors and in the systemic circulation, flow cytometry, multiplex immunohistochemistry (mIHC), and whole transcriptome sequencing were employed.

RESULTS: The study reported an ORR of 32.7%, a DCR of 67.3%, and a 6-month PFS rate of 38.5%, with a median PFS of 5.5 months. Patients treated with a combination of immunotherapy and gemcitabine achieved the longest PFS. The first-line treatment cohort exhibited a significantly higher DCR (79.3% vs. 52.2%, P = 0.038) and a longer median PFS (6.6 vs. 3.5 months, P = 0.032) compared to the second-line treatment cohort. The efficacy of treatment varied depending on the drug combinations used. Flow cytometry analysis revealed a greater frequency of CD45- CD64+ cells in the peripheral blood of patients with progressive disease (PD) compared to those with a partial response (PR). Multiplex immunofluorescence (MIF) analysis indicated an increased intratumoral infiltration of CD8+ T cells and CD137+ CD8+ T cells in patients with PR. Whole transcriptome sequencing (WTSS) identified key genes involved in immune regulation, signal transduction, and digestive function. Hemopexin (HPX) and regulatory factor X-associated protein (RFXAP) were upregulated in PR patients and showed a positive correlation with survival, whereas Interleukin-6 (IL-6) expression was linked to poor prognosis.

CONCLUSIONS: These findings indicate that immunochemotherapy shows potential for the treatment of advanced PDAC. Our study elucidates the immune landscape associated with PDAC and provides critical insights for the identification of prospective therapeutic targets, which could guide the development of innovative combination immunotherapy strategies.

PMID:41112307 | PMC:PMC12528169 | DOI:10.3389/fimmu.2025.1673098

New advances in oral microbiology and tumor research

World J Clin Oncol. 2025 Jul 24;16(7):106981. doi: 10.5306/wjco.v16.i7.106981.

ABSTRACT

Cancer remains a major global health concern, with escalating incidence and mortality rates underscoring the urgent need for novel diagnostic and therapeutic strategies. Increasing evidence has identified the oral microbiota as a critical contributor to tumorigenesis, thereby expanding the understanding of cancer pathogenesis beyond conventional risk factors such as tobacco use and genetic predisposition. This review summarizes recent progress in elucidating the complex relationship between the oral microbiota and various malignancies, particularly oral squamous cell carcinoma, esophageal adenocarcinoma, and pancreatic ductal adenocarcinoma. Pathogenic bacteria, including Porphyromonas gingivalis and Fusobacterium nucleatum, have been implicated in promoting tumor progression through mechanisms involving chronic inflammation, the production of metabolic toxins, and immune evasion. The dysbiosis of the oral microbiota, often driven by lifestyle factors such as poor diet, tobacco use, and alcohol consumption, further exacerbates these carcinogenic processes. Emerging therapeutic approaches including probiotics, oral microbiota transplantation, and CRISPR-based bacterial editing are under investigation for their potential to restore microbial homeostasis and suppress pathogenic species. Additionally, saliva-based microbial biomarkers have shown promise for non-invasive cancer screening. The integration of multi-omics technologies and artificial intelligence-driven platforms is further advancing the development of precision oncology. This review aims to consolidate fragmented findings concerning the oral microbiota-cancer axis and address existing gaps in mechanistic understanding. The review's significance lies in the translational potential of microbial research to clinical applications, offering opportunities to reduce the global cancer burden through early detection and microbiota-targeted therapies.

PMID:40741186 | PMC:PMC12304933 | DOI:10.5306/wjco.v16.i7.106981

Complex genetic variation in nearly complete human genomes

Nature, Published online: 23 July 2025; doi:10.1038/s41586-025-09140-6

Using sequencing and haplotype-resolved assembly of 65 diverse human genomes, complex regions including the major histocompatibility complex and centromeres are analysed.
❌