❌

Normal view

Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

arXiv:2605.25091v1 Announce Type: new Abstract: As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces significant challenges due to high-dimensional state spaces, discrete action commands, and strongly adversarial dynamic environments. To overcome the limitations of existing multi-agent reinforcement learning (MARL) methods in such settings, namely insufficient exploration efficiency, low sample utilization, and poor policy generalization, we propose Adversarial Curriculum and Evolutionary-enhanced Multi-agent Proximal Policy Optimization (ACE-MAPPO), a hybrid learning framework that integrates evolutionary algorithms with MAPPO. Specifically, a genetic soft update mechanism is introduced to enhance population diversity and mitigate convergence to local optima. An evolutionary-augmented prioritized trajectory replay strategy is further employed to improve the utilization of sparse high-value samples. In addition, an adversarial evolutionary curriculum learning mechanism is designed to enable adaptive training with progressively increasing difficulty. Extensive experimental results demonstrate that the proposed method outperforms MAPPO and other baseline algorithms in terms of training stability, convergence speed, and win rate, validating its effectiveness in multi-aircraft cooperative air combat scenarios.

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing benchmarks prioritize reproducibility but are often limited to open-source apps or file-operation tasks for the difficulty of constructing rewards on real applications, leaving a gap between benchmark settings and real-world usage. Moreover, most benchmarks focus on basic grounding and navigation, with limited coverage of complex, long-horizon interactions. To address these limitations, we introduce SimuWoB, a fully synthetic benchmark for mobile GUI agents with 120 challenging tasks spanning diverse types and difficulty levels. We build a robust virtual environment generation framework that synthesizes high-fidelity tasks and environments, and automatically provides valid rewards for each task. Each environment is deployed as a backend-free webpage accessible via URL, enabling efficient and reproducible evaluation. We conduct comprehensive experiments on several state-of-the-art mobile GUI agents. The average success rate is only 27.92%, dropping to 17.82% on long-horizon tasks, which reveals substantial weaknesses in current agents under complex scenarios. Evaluation result comparison with real-world sample tasks demonstrate that agent assessments based on our synthetic environment generalize well. We further provide diagnostic insights across key capability dimensions and discuss implications for future mobile GUI agent development.

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm.

A World Model of Radiologist Reading for Medical Image Representation Learning

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

arXiv:2509.25339v3 Announce Type: replace-cross Abstract: Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

arXiv:2510.10921v3 Announce Type: replace-cross Abstract: Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. Our approach leverages rich fine-grained supervision, including region-text matching and long-caption modeling, alongside multiple discriminative objectives. We further introduce the Textual Intra-modal Contrastive (TIC) loss to better distinguish semantically similar captions. Trained on a carefully curated mixture of large-scale English and Chinese data, including a newly released 12M Chinese region-text dataset, FG-CLIP 2 achieves powerful bilingual performance. To enable rigorous evaluation, we present a new benchmark for Chinese multimodal understanding, featuring long-caption retrieval and bounding box classification. Extensive experiments on 29 datasets across 8 tasks show that FG-CLIP 2 outperforms existing methods, achieving state-of-the-art results in both languages. We release the model, code, and benchmark to facilitate future research on bilingual fine-grained vision-language alignment.

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

arXiv:2601.22709v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. Treating the teacher as a proxy for task-relevant information, we introduce confidence-gated decoupled distillation to filter unreliable supervision, relational centered kernel alignment to transfer visual token structures, and an adaptive controller via Lagrangian relaxation to balance fidelity against capacity constraints. Across extensive benchmarks on LLaVA and Qwen families, our INT4 models consistently outperform FP16 baselines (e.g., LLaVA-1.5-7B: 70.1 vs. 66.8 on SQA; Qwen2-VL-2B: 76.9 vs. 72.6 on MMBench), nearly matching teacher performance. Using real INT4 kernel, we achieve 3$\times$ throughput with 54% memory reduction. This principled framework significantly outperforms existing quantization methods, making GRACE a compelling solution for resource-constrained deployment. Code and data are available at: https://github.com/ForeverBlue816/GRACE.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

arXiv:2602.08499v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods utilize rollouts in an indiscriminate and short-horizon manner: responses of heterogeneous quality within each prompt are treated uniformly, and historical rollouts are discarded after a single use. This leads to noisy supervision, poor sample efficiency, and suboptimal policy updates. We address these issues by formulating rollout scheduling in RLVR as a contextual bandit problem and proposing a unified neural scheduling framework that adaptively selects high-value rollouts throughout training. Each rollout is treated as an arm whose reward is defined by the induced performance gain between consecutive optimization steps. The resulting scheduler supports both noise-aware intra-group selection and adaptive global reuse of historical rollouts within a single principled framework. We provide theoretical justification by deriving sublinear regret bounds and showing that enlarging the rollout buffer improves the achievable performance upper bound. Experiments on six mathematical reasoning benchmarks demonstrate consistent gains in performance and training efficiency across multiple RLVR optimization methods.

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

arXiv:2605.23440v2 Announce Type: replace-cross Abstract: Joint Entity and Relation Extraction (JERE) is highly susceptible to weak generalization due to low-quality training data. Data augmentation is a common strategy to enhance model generalization across different domains. However, existing data augmentation methods often overlook text relevance and may disrupt semantic structures and dependencies, making it difficult to generate effective augmented data for improving model generalization. In this paper, we propose Structured Semantic Data Augmentation (SSDAU), a novel method designed to preserve the semantic structure of text during augmentation. SSDAU segments text based on entity labels and employs an encoder to capture semantic features of entities through context awareness. It then performs entity semantic restructuring to generate augmented data. To distinguish semantically similar entities, SSDAU fuses contextualized embeddings with traditional similarity scores. To mitigate potential topic ambiguity and information loss, we apply the BERTTopic model to filter out irrelevant topics, ensuring topic consistency. We evaluate SSDAU on datasets with different annotation types and compare its performance on five representative JERE models against seven popular data augmentation baselines. Experiments demonstrate that SSDAU generates semantically consistent data with superior robustness against ambiguity (8.26% F1 decrease vs. 31.91% for baselines), significantly outperforming all existing methods across all metrics.

Clinical and translational roles of circulating tumor cells in non-small cell and small cell lung cancer: a narrative review

J Thorac Dis. 2026 Apr 30;18(4):415. doi: 10.21037/jtd-2026-1-0025. Epub 2026 Apr 24.

ABSTRACT

BACKGROUND AND OBJECTIVE: Circulating tumor cells (CTCs) are malignant cells shed into blood that enable noninvasive, longitudinal assessment of lung cancer. Increasing evidence frames CTCs within a circulating tumor microenvironment (cTME) and broader circulating tumor-associated cell (CTAC) ecosystems that include multicellular clusters and circulating tumor endothelial cells (CTECs). We summarize definitions, detection approaches, and clinical applications of CTC-centered liquid biopsy in non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC).

METHODS: A comprehensive literature search was conducted in PubMed, Embase, Web of Science, and Google Scholar using the terms "non-small cell lung cancer", "small cell lung cancer", and "circulating tumor cells". Relevant clinical, basic, and translational studies were selected and synthesized to outline current knowledge and future directions.

KEY CONTENT AND FINDINGS: CTCs can be enriched by immunoaffinity, size, or microfluidic platforms, enabling enumeration and downstream profiling. In both NSCLC and SCLC, CTC positivity and higher burden are associated with worse survival, with the strongest effects in SCLC and with circulating tumor emboli (CTE). Serial monitoring provides early signals of response or failure; and post-treatment supports minimal residual disease (MRD) detection and relapse prediction. Molecular and phenotypic profiling enables driver and resistance tracking, including epidermal growth factor receptor (EGFR) and anaplastic lymphoma kinase (ALK), while CTECs may add vascular and immune-relevant information.

CONCLUSIONS: CTC-based assays have the potential to complement imaging and tissue biopsy across screening research, prognostication, therapeutic monitoring, MRD assessment, and personalized care. Clinical translation requires standardized preanalytical workflows, harmonized thresholds, and prospective trials testing CTC-guided management.

PMID:42182710 | PMC:PMC13190041 | DOI:10.21037/jtd-2026-1-0025

Machine learning-driven multi-omics integration uncovers a senescence associated molecular axis in HCC

Front Immunol. 2026 May 8;17:1762222. doi: 10.3389/fimmu.2026.1762222. eCollection 2026.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) exhibits profound molecular heterogeneity and aberrant cellular senescence. This study systematically dissects the senescence-associated molecular landscape to identify key regulators driving HCC progression and immune evasion.

METHODS: Integrating multi-cohort transcriptomic datasets, we developed a robust prognostic signature using 101 machine-learning models, identifying prognostic signature. We employed preliminary proteomic, exploratory metabolomic, and single-cell RNA sequencing (scRNA-seq) analyses to explore multi-omics alterations. The functional senescence status and MCM7 were validated in a clinical HCC cohort by RT-qPCR, Western blotting, immunohistochemistry, and multiplex immunofluorescence (mIF). Causality was established using in vitro functional assays in HepG2 cells.

RESULTS: A 12-gene random survival forest (RSF) signature accurately predicted patient survival across independent cohorts. MCM7 emerged as a central senescence-associated driver. ScRNA-seq and mIF confirmed MCM7 characterizes a highly proliferative, clonally expanding subset of CD8+ T cells within the tumor microenvironment. In vitro, MCM7 knockdown significantly inhibited HepG2 cell proliferation and upregulated senescence enforcers p16 and p21, whereas overexpression facilitated evasion. Additionally, TIDE analysis revealed that high-risk patients exhibited elevated immune evasion potential, predicting poor immunotherapy response.

CONCLUSION: This integrative multi-omics framework uncovers an MCM7 MCM7-driven senescence-associated axis promising HCC progression and immune dysfunction, offering a robust tool for prognostic stratification and novel therapeutic insights.

PMID:42183188 | PMC:PMC13195000 | DOI:10.3389/fimmu.2026.1762222

Machine learning-driven multi-omics integration uncovers a senescence associated molecular axis in HCC

Front Immunol. 2026 May 8;17:1762222. doi: 10.3389/fimmu.2026.1762222. eCollection 2026.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) exhibits profound molecular heterogeneity and aberrant cellular senescence. This study systematically dissects the senescence-associated molecular landscape to identify key regulators driving HCC progression and immune evasion.

METHODS: Integrating multi-cohort transcriptomic datasets, we developed a robust prognostic signature using 101 machine-learning models, identifying prognostic signature. We employed preliminary proteomic, exploratory metabolomic, and single-cell RNA sequencing (scRNA-seq) analyses to explore multi-omics alterations. The functional senescence status and MCM7 were validated in a clinical HCC cohort by RT-qPCR, Western blotting, immunohistochemistry, and multiplex immunofluorescence (mIF). Causality was established using in vitro functional assays in HepG2 cells.

RESULTS: A 12-gene random survival forest (RSF) signature accurately predicted patient survival across independent cohorts. MCM7 emerged as a central senescence-associated driver. ScRNA-seq and mIF confirmed MCM7 characterizes a highly proliferative, clonally expanding subset of CD8+ T cells within the tumor microenvironment. In vitro, MCM7 knockdown significantly inhibited HepG2 cell proliferation and upregulated senescence enforcers p16 and p21, whereas overexpression facilitated evasion. Additionally, TIDE analysis revealed that high-risk patients exhibited elevated immune evasion potential, predicting poor immunotherapy response.

CONCLUSION: This integrative multi-omics framework uncovers an MCM7 MCM7-driven senescence-associated axis promising HCC progression and immune dysfunction, offering a robust tool for prognostic stratification and novel therapeutic insights.

PMID:42183188 | PMC:PMC13195000 | DOI:10.3389/fimmu.2026.1762222

PRXL2B facilitates the progression of hepatocellular carcinoma and the therapeutic efficacy of oncolytic adenovirus H101 through the PI3K/AKT/PD-L1 axis

Biosci Trends. 2026 May 21. doi: 10.5582/bst.2026.01000. Online ahead of print.

ABSTRACT

Oncolytic adenovirus H101 has shown antitumor activity in hepatocellular carcinoma (HCC), but the molecular determinants of treatment response remain unclear. In this study, a Hepa1-6 subcutaneous tumor model was established in C57BL/6 mice and treated with intratumoral H101, followed by integrated transcriptomic and proteomic analyses to identify candidate genes associated with H101 response. PRXL2B was selected for further investigation using public multi-omics datasets, tissue microarray-based immunohistochemistry, in vitro functional assays, mechanistic analyses, and in vivo validation experiments. Integrated multi-omics analyses identified PRXL2B as a candidate gene downregulated after H101 treatment. Public datasets and tissue-based validation further showed that PRXL2B was upregulated in HCC tissues. In MHCC97H and HCCLM3 cells, PRXL2B knockdown inhibited proliferation, migration, and invasion, promoted apoptosis and cell-cycle arrest, and enhanced the antitumor effect of H101. Mechanistically, PRXL2B silencing reduced AKT phosphorylation and PD-L1 expression. In vivo, PRXL2B knockdown suppressed tumor growth, and the combination of PRXL2B knockdown and H101 produced the strongest antitumor effect. These findings indicate that PRXL2B promotes malignant phenotypes in HCC and may modulate H101 efficacy through the PI3K/AKT/PD-L1 axis. Targeting PRXL2B may therefore represent a potential strategy to enhance the therapeutic efficacy of oncolytic virus therapy in HCC.

PMID:42161529 | DOI:10.5582/bst.2026.01000

A pathogen lncRNA secreted into rice sequesters a host miRNA for virulence

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10572-x

A fungal long non-coding RNA from Magnaporthe oryzae translocates into rice cells to sequester a host microRNA that normally represses PKR1, a negative immunity regulator, thereby facilitating infection and revealing a widespread RNA-based pathogen–host interaction mechanism.

High-salt diet in macrophage-associated metabolic disorders: Mechanisms and therapeutic implications

Chin Med J (Engl). 2026 May 19. doi: 10.1097/CM9.0000000000004098. Online ahead of print.

ABSTRACT

High-salt diet (HSD) has emerged as a prevalent environmental factor that exacerbates chronic inflammation and insulin resistance in obesity-associated type 2 diabetes (T2D) by modulating macrophage polarization, metabolic reprogramming, and epigenetic imprinting. Current evidence demonstrates that HSD activates p38/mitogen-activated protein kinase (MAPK), nuclear factor kappa-B (NF-κB), and NOD-like receptor family pyrin domain containing 3 (NLRP3) inflammasome signaling pathways, by which it drives macrophage polarization toward a proinflammatory M1 phenotype while inducing a glycolysis-dominant metabolic shift, thereby establishing a persistent "metabolic memory". Moreover, HSD orchestrates metabolic memory in macrophages through coordinated epigenetic machinery, including histone modifications (Trimethylation of histone H3 at lysine 4 [H3K4me3] and Acetylation of histone H3 at lysine 27 [H3K27ac]), DNA methylation, and noncoding RNAs (e.g., long non-coding RNA MALAT1 and miR-155), leading to sustained inflammatory phenotypes. In multiple metabolic organs (e.g., adipose tissue, liver, pancreas, and gut), the HSD-macrophage axis aggravates systemic insulin resistance through shared proinflammatory signaling and other tissue-specific mechanisms. Most importantly, therapeutic strategies targeting the NLRP3 inflammasome, metabolic pathways, and epigenetic alterations offer novel approaches for managing metabolic inflammation. Future investigations are encouraged to leverage lineage tracing, single-cell sequencing, and spatial multi-omics technologies to advance the development of precision medicine for macrophage-associated metabolic disorders.

PMID:42156155 | DOI:10.1097/CM9.0000000000004098

Imaging interface-controlled bulk oxygen spillover

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10324-x

In situ microscopic single-particle imaging demonstrates the significance of rationally engineered metal–support interfaces for activating the oxygen in bulk catalyst, helping elucidate reaction pathways in catalytic conversions.
❌