❌

Normal view

Multi-Omics and Single-Cell Mendelian Randomization Reveal a Potential Role of VNN2 in Lung Adenocarcinoma in Resting Natural Killer Cells

World J Oncol. 2026 Mar 5;17(2):247-255. doi: 10.14740/wjon2689. eCollection 2026 Apr.

ABSTRACT

BACKGROUND: We aimed to evaluate the potential association between genetically predicted vanin-2 (VNN2) expression and lung adenocarcinoma (LUAD) risk, and to explore the immune cell subtype that may underlie this relationship.

METHODS: We integrated whole-blood expression quantitative trait loci (eQTL) data from eQTLGen, plasma protein quantitative trait loci (pQTL) data from deCODE, and LUAD genome-wide association study (GWAS) data from European-ancestry cohorts, together with differential expression analysis using GEPIA2, to identify candidate genes for subsequent single-cell eQTL (sc-eQTL) Mendelian randomization (MR) analysis. For the sc-eQTL analysis, VNN2-associated eQTLs from 14 immune cell types profiled in the OneK1K single-cell eQTL resource were tested for associations with LUAD risk.

RESULTS: Bulk-level MR analysis showed that genetically predicted increases in VNN2 expression and protein levels were significantly associated with a reduced risk of LUAD (eQTL-MR: odds ratio (OR) = 0.964, 95% confidence interval (95% CI), 0.934-0.995; P = 0.024; pQTL-MR: OR = 0.946, 95% CI, 0.921-0.970; P = 2.87 × 10-5). Transcriptomic analyses confirmed significant downregulation of VNN2 in LUAD tumors compared with normal lung tissues. sc-eQTL MR identified the strongest association in resting natural killer (rNK) cells (OR = 0.896, 95% CI, 0.829-0.967; P = 0.005).

CONCLUSIONS: Multi-omics and sc-eQTL MR analyses indicated that genetically predicted increases in VNN2 expression were associated with a reduced risk of LUAD, with the most pronounced effect observed in rNK cells. These findings suggest a potential cell type-specific role of VNN2 in LUAD susceptibility and warrant further studies to validate its biological relevance and clinical implications.

PMID:41822323 | PMC:PMC12978397 | DOI:10.14740/wjon2689

Profiling of the mycobiome and metabolome: a comparative study of benign pulmonary nodules and lung adenocarcinoma

Front Cell Infect Microbiol. 2026 Feb 23;16:1732958. doi: 10.3389/fcimb.2026.1732958. eCollection 2026.

ABSTRACT

INTRODUCTION: Lung adenocarcinoma (LUAD), the most common subtype of non-small cell lung cancer, is a form of malignant pulmonary nodule that requires clinical differentiation from benign pulmonary nodules (BPN). The mechanisms underlying the development of LUAD are complex, and effective non-invasive methods for differentiating BPN from LUAD are lacking. This study aimed not only to distinguish BPN from LUAD using gut fungi and serum metabolites, but also to establish an integrated network of gut fungi-metabolite-cytokine interactions.

METHODS: Fecal and serum samples from individuals with BPN and patients with LUAD were subjected to internal transcribed spacer sequencing, ultra-performance liquid chromatography-tandem mass spectrometry, and multiplex Luminex assays to quantify gut fungi, metabolites, and cytokines, respectively.

RESULTS: A significant difference in gut fungal communities was observed between the BPN and LUAD groups. Multiple genera and species were more abundant in LUAD than in BPN. Docosapentaenoic acid n-6 (DPAn-6), indole-3-propionic acid (IPA), and interferon-γ-induced protein 10 (IP-10) were significantly elevated in the LUAD group. The integrated model established using a combination of gut fungi and metabolites demonstrated excellent performance in distinguishing BPN from LUAD. A network of interactions was established among differentially abundant gut fungi, serum metabolites, and cytokines.

CONCLUSION: Our study identifies a novel panel of fungal and metabolite biomarkers for differentiating between BPN and LUAD, and constructs a multi-omics network that provides new insights into investigating the mechanistic role of gut mycobiota dysbiosis in LUAD.

PMID:41809995 | PMC:PMC12968269 | DOI:10.3389/fcimb.2026.1732958

Facile induction of immune tolerance by an interleukin-2–TGFβ surrogate agonist

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10208-0

A fusion protein designed to comprise IL-2 and a helminth-derived TGFβ mimic activates IL-2 and TGFβ signalling pathways in IL-2 receptor-expressing T cells and induces stable antigen-specific regulatory T cells in peripheral lymphoid organs.

Risk-adaptive therapy guided by dynamic ctDNA in nasopharyngeal carcinoma

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10244-w

A clinical trial testing whether monitoring ctDNA clearance during treatment for nasopharyngeal cancer could be used to inform decisions about an individual’s subsequent therapeutic programme shows promising results.

Breast Cancer Screening Knowledge and Sentiments in Singaporean Women: Mixed Methods Study Using Topic Modeling, Sentiment Analysis, and Structured Questionnaire Data

Background: Mammography screening uptake in Singapore remains below 40% despite campaigns and subsidies. Natural language processing (NLP) can extract nuanced attitudes from free text that fixed response options miss, revealing latent factors influencing breast cancer (BC) screening behavior. Objective: This study characterized women’s attitudes toward mammography using mixed methods data, examined associations between BC awareness and screening willingness, and identified barriers and facilitators through NLP of free-text responses. Methods: We conducted a cross-sectional study within the multicenter cohort in Singapore (October 2021-December 2023). In total, 4169 women aged 35‐59 years (median 48, IQR 43‐54) were recruited via convenience sampling (3 hospitals and 2 polyclinics). Participants completed online structured questionnaires on demographics and screening history, then a BC education quiz with feedback. Participants answering >80% correctly were classified as “BC-aware.” Posteducation, participants reported screening willingness (motivated or neutral) with optional free-text explanations. Logistic regression models (adjusted for study site, age, ethnicity, marital status, housing, and education) examined the associations with willingness. For 3819 English-language respondents, biterm topic modeling identified themes and sentiment analysis quantified emotional tone. Statistical significance: =.05. Results: Overall, 79% (3287/4169) were BC-aware, and 94% (3908/4169) reported increased motivation posteducation. BC-aware women had higher screening motivation than BC-unaware women (adjusted odds ratio [aOR] 2.88, 95% CI 2.19‐3.80;

Agentic Critical Training

arXiv:2603.08706v1 Announce Type: new Abstract: Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contrast successful actions against suboptimal alternatives and thus lack awareness of action quality. Recent approaches attempt to address this by introducing self-reflection supervision derived from contrasts between expert and alternative actions. However, the training paradigm fundamentally remains imitation learning: the model imitates pre-constructed reflection text rather than learning to reason autonomously. We propose Agentic Critical Training (ACT), a reinforcement learning paradigm that trains agents to identify the better action among alternatives. By rewarding whether the model's judgment is correct, ACT drives the model to autonomously develop reasoning about action quality, producing genuine self-reflection rather than imitating it. Across three challenging agent benchmarks, ACT consistently improves agent performance when combined with different post-training methods. It achieves an average improvement of 5.07 points over imitation learning and 4.62 points over reinforcement learning. Compared to approaches that inject reflection capability through knowledge distillation, ACT also demonstrates clear advantages, yielding an average improvement of 2.42 points. Moreover, ACT enables strong out-of-distribution generalization on agentic benchmarks and improves performance on general reasoning benchmarks without any reasoning-specific training data, highlighting the value of our method. These results suggest that ACT is a promising path toward developing more reflective and capable LLM agents.

GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning

arXiv:2603.06620v1 Announce Type: cross Abstract: The growing demand for automated graph algorithm reasoning has attracted increasing attention in the large language model (LLM) community. Recent LLM-based graph reasoning methods typically decouple task descriptions from graph data, generate executable code augmented by retrieval from technical documentation, and refine the code through debugging. However, we identify two key limitations in existing approaches: (i) they treat technical documentation as flat text collections and ignore its hierarchical structure, leading to noisy retrieval that degrades code generation quality; and (ii) their debugging mechanisms focus primarily on runtime errors, yet ignore more critical logical errors. To address them, we propose {\method}, an \textit{agentic hierarchical retrieval-augmented coding framework} that exploits the document hierarchy through top-down traversal and early pruning, together with a \textit{self-debugging coding agent} that iteratively refines code using automatically generated small-scale test cases. To enable comprehensive evaluation of complex graph reasoning, we introduce a new dataset, {\dataset}, covering small-scale, large-scale, and composite graph reasoning tasks. Extensive experiments demonstrate that our method achieves higher task accuracy and lower inference cost compared to baselines\footnote{The code is available at \href{https://github.com/FairyFali/GraphSkill}{\textcolor{blue}{https://github.com/FairyFali/GraphSkill}}.}.

MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation

arXiv:2603.08572v1 Announce Type: cross Abstract: Learning natural, stable, and compositionally generalizable whole-body control policies for humanoid robots performing simultaneous locomotion and manipulation (loco-manipulation) remains a fundamental challenge in robotics. Existing reinforcement learning approaches typically rely on a single monolithic policy to acquire multiple skills, which often leads to cross-skill gradient interference and motion pattern conflicts in high-degree-of-freedom systems. As a result, generated behaviors frequently exhibit unnatural movements, limited stability, and poor generalization to complex task compositions. To address these limitations, we propose MetaWorld-X, a hierarchical world model framework for humanoid control. Guided by a divide-and-conquer principle, our method decomposes complex control problems into a set of specialized expert policies (Specialized Expert Policies, SEP). Each expert is trained under human motion priors through imitation-constrained reinforcement learning, introducing biomechanically consistent inductive biases that ensure natural and physically plausible motion generation. Building upon this foundation, we further develop an Intelligent Routing Mechanism (IRM) supervised by a Vision-Language Model (VLM), enabling semantic-driven expert composition. The VLM-guided router dynamically integrates expert policies according to high-level task semantics, facilitating compositional generalization and adaptive execution in multi-stage loco-manipulation tasks.

Membership Inference Attacks on Tokenizers of Large Language Models

arXiv:2510.05699v3 Announce Type: replace-cross Abstract: Membership inference attacks (MIAs) are widely used to assess the privacy risks associated with machine learning models. However, when these attacks are applied to pre-trained large language models (LLMs), they encounter significant challenges, including mislabeled samples, distribution shifts, and discrepancies in model size between experimental and real-world settings. To address these limitations, we introduce tokenizers as a new attack vector for membership inference. Specifically, a tokenizer converts raw text into tokens for LLMs. Unlike full models, tokenizers can be efficiently trained from scratch, thereby avoiding the aforementioned challenges. In addition, the tokenizer's training data is typically representative of the data used to pre-train LLMs. Despite these advantages, the potential of tokenizers as an attack vector remains unexplored. To this end, we present the first study on membership leakage through tokenizers and explore five attack methods to infer dataset membership. Extensive experiments on millions of Internet samples reveal the vulnerabilities in the tokenizers of state-of-the-art LLMs. To mitigate this emerging risk, we further propose an adaptive defense. Our findings highlight tokenizers as an overlooked yet critical privacy threat, underscoring the urgent need for privacy-preserving mechanisms specifically designed for them.

Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization

arXiv:2603.04135v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the underlying sampling distribution, compromising theoretical rigor and convergence behavior. To address this limitation, we propose Dynamic Pruning Policy Optimization (DPPO), a framework that enables dynamic pruning while preserving unbiased gradient estimation through importance sampling-based correction. By incorporating mathematically derived rescaling factors, DPPO significantly accelerates GRPO training without altering the optimization objective of the full-batch baseline. Furthermore, to mitigate the data sparsity induced by pruning, we introduce Dense Prompt Packing, a window-based greedy strategy that maximizes valid token density and hardware utilization. Extensive experiments demonstrate that DPPO consistently accelerates training across diverse models and benchmarks. For instance, on Qwen3-4B trained on MATH, DPPO achieves 2.37$\times$ training speedup and outperforms GRPO by 3.36% in average accuracy across six mathematical reasoning benchmarks.

Quantum-Inspired Fine-Tuning for Few-Shot AIGC Detection via Phase-Structured Reparameterization

arXiv:2603.02281v1 Announce Type: cross Abstract: Recent studies show that quantum neural networks (QNNs) generalize well in few-shot regimes. To extend this advantage to large-scale tasks, we propose Q-LoRA, a quantum-enhanced fine-tuning scheme that integrates lightweight QNNs into the low-rank adaptation (LoRA) adapter. Applied to AI-generated content (AIGC) detection, Q-LoRA consistently outperforms standard LoRA under few-shot settings. We analyze the source of this improvement and identify two possible structural inductive biases from QNNs: (i) phase-aware representations, which encode richer information across orthogonal amplitude-phase components, and (ii) norm-constrained transformations, which stabilize optimization via inherent orthogonality. However, Q-LoRA incurs non-trivial overhead due to quantum simulation. Motivated by our analysis, we further introduce H-LoRA, a fully classical variant that applies the Hilbert transform within the LoRA adapter to retain similar phase structure and constraints. Experiments on few-shot AIGC detection show that both Q-LoRA and H-LoRA outperform standard LoRA by over 5% accuracy, with H-LoRA achieving comparable accuracy at significantly lower cost in this task.

A Deployment-Friendly Foundational Framework for Efficient Computational Pathology

arXiv:2602.14010v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have enabled robust generalization in computational pathology through large-scale datasets and expansive architectures, but their substantial computational cost, particularly for gigapixel whole slide images, limits clinical accessibility and scalability. Here, we present LitePath, a deployment-friendly foundational framework designed to mitigate model over-parameterization and patch level redundancy. LitePath integrates LiteFM, a compact model distilled from three large PFMs (Virchow2, H-Optimus-1 and UNI2) using 190 million patches, and the Adaptive Patch Selector (APS), a lightweight component for task-specific patch selection. The framework reduces model parameters by 28x and lowers FLOPs by 403.5x relative to Virchow2, enabling deployment on low-power edge hardware such as the NVIDIA Jetson Orin Nano Super. On this device, LitePath processes 208 slides per hour, 104.5x faster than Virchow2, and consumes 0.36 kWh per 3,000 slides, 171x lower than Virchow2 on an RTX3090 GPU. We validated accuracy using 37 cohorts across four organs and 26 tasks (26 internal, 9 external, and 2 prospective), comprising 15,672 slides from 9,808 patients disjoint from the pretraining data. LitePath ranks second among 19 evaluated models and outperforms larger models including H-Optimus-1, mSTAR, UNI2 and GPFM, while retaining 99.71% of the AUC of Virchow2 on average. To quantify the balance between accuracy and efficiency, we propose the Deployability Score (D-Score), defined as the weighted geometric mean of normalized AUC and normalized FLOP, where LitePath achieves the highest value, surpassing Virchow2 by 10.64%. These results demonstrate that LitePath enables rapid, cost-effective and energy-efficient pathology image analysis on accessible hardware while maintaining accuracy comparable to state-of-the-art PFMs and reducing the carbon footprint of AI deployment.

UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model

arXiv:2602.14178v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) require a visual representation that simultaneously supports high-fidelity reconstruction, complex semantic extraction, and generative suitability. However, existing visual tokenizers typically struggle to satisfy these conflicting objectives within a single framework. In this paper, we introduce UniWeTok, a unified discrete tokenizer designed to bridge this gap using a massive binary codebook ($\mathit{2^{128}}$). For training framework, we introduce Pre-Post Distillation and a Generative-Aware Prior to enhance the semantic extraction and generative prior of the discrete tokens. In terms of model architecture, we propose a convolution-attention hybrid architecture with the SigLu activation function. SigLu activation not only bounds the encoder output and stabilizes the semantic distillation process but also effectively addresses the optimization conflict between token entropy loss and commitment loss. We further propose a three-stage training framework designed to enhance UniWeTok's adaptability cross various image resolutions and perception-sensitive scenarios, such as those involving human faces and textual content. On ImageNet, UniWeTok achieves state-of-the-art image generation performance (FID: UniWeTok 1.38 vs. REPA 1.42) while requiring a remarkably low training compute (Training Tokens: UniWeTok 33B vs. REPA 262B). On general-domain, UniWeTok demonstrates highly competitive capabilities across a broad range of tasks, including multimodal understanding, image generation (DPG Score: UniWeTok 86.63 vs. FLUX.1 [Dev] 83.84), and editing (GEdit Overall Score: UniWeTok 5.09 vs. OmniGen 5.06). We release code and models to facilitate community exploration of unified tokenizer and MLLM.

Synergizing Foundation Models and Federated Learning: A Survey

arXiv:2406.12844v2 Announce Type: replace-cross Abstract: Over the past few years, the landscape of Artificial Intelligence (AI) has been reshaped by the emergence of Foundation Models (FMs). Pre-trained on massive datasets, these models exhibit exceptional performance across diverse downstream tasks through adaptation techniques like fine-tuning and prompt learning. More recently, the synergy of FMs and Federated Learning (FL) has emerged as a promising paradigm, often termed Federated Foundation Models (FedFM), allowing for collaborative model adaptation while preserving data privacy. This survey paper provides a systematic review of the current state of the art in FedFM, offering insights and guidance into the evolving landscape. Specifically, we present a comprehensive multi-tiered taxonomy based on three major dimensions, namely efficiency, adaptability, and trustworthiness. To facilitate practical implementation and experimental research, we undertake a thorough review of existing libraries and benchmarks. Furthermore, we discuss the diverse real-world applications of this paradigm across multiple domains. Finally, we outline promising research directions to foster future advancements in FedFM. Overall, this survey serves as a resource for researchers and practitioners, offering a thorough understanding of FedFM's role in revolutionizing privacy-preserving AI and pointing toward future innovations in this promising area. A periodically updated paper collection on FM-FL is available at https://github.com/lishenghui/awesome-fm-fl.
❌