❌

Reading view

Author Correction: A base editor for the long-term restoration of auditory function in mice with recessive profound deafness

Nature Biomedical Engineering, Published online: 07 September 2026; doi:10.1038/s41551-026-01794-5

Author Correction: A base editor for the long-term restoration of auditory function in mice with recessive profound deafness
  •  

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

arXiv:2609.09187v1 Announce Type: cross Abstract: Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.
  •  

Reliable Near-Field Multi-User Positioning Informed by Two-Stage MUSIC

arXiv:2609.09409v1 Announce Type: cross Abstract: Near-field localization is a promising technique for high-resolution multi-user positioning in future wireless systems, but its performance is often degraded by scattering-induced coherent propagation. Existing near-field localization methods, which require separate parameter estimation and path/source association, suffer from high computation overhead and accumulated errors, and usually do not provide any guarantee on reliability. In this paper, we propose \emph{MUSIC-Net}, an end-to-end near-field positioning deep learning (DL) framework informed by two-stage MUltiple SIgnal Classification (MUSIC) in mixed line-of-sight (LoS) and non-LoS (NLoS) multi-path scenarios, which embeds the two-stage MUSIC objects into training to isolate the LoS-related signal subspace and to identify a surrogate distance. The proposed framework directly recovers multi-user positions without the need for involved NLoS parameter estimation or path/source association. Furthermore, we introduce split conformal prediction (SCP) to move beyond point-estimation-based positioning towards statistically guaranteed (confidence) set estimation for all users. Numerical results show that the proposed MUSIC-Net achieves lower mean positioning error (MPER) than existing benchmarks and yields tighter SCP-calibrated prediction regions, demonstrating both accurate LoS localization and efficient uncertainty quantification (UQ) in coherent multi-path environments.
  •  

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

arXiv:2609.04298v2 Announce Type: replace Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
  •  

RAU: Reference-based Anatomical Understanding with Vision Language Models

arXiv:2509.22404v2 Announce Type: replace-cross Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its generalization ability to unseen modalities makes it scalable to unseen datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.
  •  

An operational perturbation proteomics-based virtual cell model

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11001-9

Temporal protein-abundance measurements from systematically perturbed breast cancer cell lines were generated to develop ProteinTalks, a virtual cell model that functions as an operational tool for diverse drug discovery tasks.
  •  

CAFs shape the immunosuppressive microenvironment of pancreatic cancer through the Lin28b-STING Axis

Nat Commun. 2026 Aug 7;17(1):9491. doi: 10.1038/s41467-026-76495-3.

ABSTRACT

Cancer-associated fibroblasts comprise diverse functionally distinct cellular subsets, with certain subpopulations exerting pivotal influence in shaping the pancreatic cancer immune microenvironment. Here we show that Lin28b+ cancer-associated fibroblasts contribute to establishing an immunologically cold tumor microenvironment in pancreatic ductal adenocarcinoma. Mechanistically, Lin28b directly binds to STING mRNA and promotes its degradation, thereby suppressing STING expression and downstream type I interferon signaling. Loss of Lin28b in cancer-associated fibroblasts activates the cGAS-STING-interferon signaling cascade, enhancing dendritic cell antigen presentation and CD8+ T cell cytotoxic function. Importantly, genetic inhibition of Lin28b in cancer-associated fibroblasts enhances sensitivity to anti-PD-L1 immune checkpoint blockade therapy. These findings reveal that targeting the Lin28b-STING axis represents a promising therapeutic strategy for overcoming the intrinsic resistance of pancreatic ductal adenocarcinoma to immunotherapy.

PMID:42693143 | PMC:PMC13542369 | DOI:10.1038/s41467-026-76495-3

  •  

The RNA-binding protein La/SSB is associated with HNSCC progression and TFAP2C/FSCN1-linked transcriptional regulation

Oncogene, Published online: 03 September 2026; doi:10.1038/s41388-026-03959-7

The RNA-binding protein La/SSB is associated with HNSCC progression and TFAP2C/FSCN1-linked transcriptional regulation
  •  

Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

arXiv:2605.25091v1 Announce Type: new Abstract: As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces significant challenges due to high-dimensional state spaces, discrete action commands, and strongly adversarial dynamic environments. To overcome the limitations of existing multi-agent reinforcement learning (MARL) methods in such settings, namely insufficient exploration efficiency, low sample utilization, and poor policy generalization, we propose Adversarial Curriculum and Evolutionary-enhanced Multi-agent Proximal Policy Optimization (ACE-MAPPO), a hybrid learning framework that integrates evolutionary algorithms with MAPPO. Specifically, a genetic soft update mechanism is introduced to enhance population diversity and mitigate convergence to local optima. An evolutionary-augmented prioritized trajectory replay strategy is further employed to improve the utilization of sparse high-value samples. In addition, an adversarial evolutionary curriculum learning mechanism is designed to enable adaptive training with progressively increasing difficulty. Extensive experimental results demonstrate that the proposed method outperforms MAPPO and other baseline algorithms in terms of training stability, convergence speed, and win rate, validating its effectiveness in multi-aircraft cooperative air combat scenarios.
  •  

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing benchmarks prioritize reproducibility but are often limited to open-source apps or file-operation tasks for the difficulty of constructing rewards on real applications, leaving a gap between benchmark settings and real-world usage. Moreover, most benchmarks focus on basic grounding and navigation, with limited coverage of complex, long-horizon interactions. To address these limitations, we introduce SimuWoB, a fully synthetic benchmark for mobile GUI agents with 120 challenging tasks spanning diverse types and difficulty levels. We build a robust virtual environment generation framework that synthesizes high-fidelity tasks and environments, and automatically provides valid rewards for each task. Each environment is deployed as a backend-free webpage accessible via URL, enabling efficient and reproducible evaluation. We conduct comprehensive experiments on several state-of-the-art mobile GUI agents. The average success rate is only 27.92%, dropping to 17.82% on long-horizon tasks, which reveals substantial weaknesses in current agents under complex scenarios. Evaluation result comparison with real-world sample tasks demonstrate that agent assessments based on our synthetic environment generalize well. We further provide diagnostic insights across key capability dimensions and discuss implications for future mobile GUI agent development.
  •  

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm.
  •  

A World Model of Radiologist Reading for Medical Image Representation Learning

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.
  •  

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

arXiv:2509.25339v3 Announce Type: replace-cross Abstract: Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload
  •  

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

arXiv:2510.10921v3 Announce Type: replace-cross Abstract: Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. Our approach leverages rich fine-grained supervision, including region-text matching and long-caption modeling, alongside multiple discriminative objectives. We further introduce the Textual Intra-modal Contrastive (TIC) loss to better distinguish semantically similar captions. Trained on a carefully curated mixture of large-scale English and Chinese data, including a newly released 12M Chinese region-text dataset, FG-CLIP 2 achieves powerful bilingual performance. To enable rigorous evaluation, we present a new benchmark for Chinese multimodal understanding, featuring long-caption retrieval and bounding box classification. Extensive experiments on 29 datasets across 8 tasks show that FG-CLIP 2 outperforms existing methods, achieving state-of-the-art results in both languages. We release the model, code, and benchmark to facilitate future research on bilingual fine-grained vision-language alignment.
  •  

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

arXiv:2601.22709v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. Treating the teacher as a proxy for task-relevant information, we introduce confidence-gated decoupled distillation to filter unreliable supervision, relational centered kernel alignment to transfer visual token structures, and an adaptive controller via Lagrangian relaxation to balance fidelity against capacity constraints. Across extensive benchmarks on LLaVA and Qwen families, our INT4 models consistently outperform FP16 baselines (e.g., LLaVA-1.5-7B: 70.1 vs. 66.8 on SQA; Qwen2-VL-2B: 76.9 vs. 72.6 on MMBench), nearly matching teacher performance. Using real INT4 kernel, we achieve 3$\times$ throughput with 54% memory reduction. This principled framework significantly outperforms existing quantization methods, making GRACE a compelling solution for resource-constrained deployment. Code and data are available at: https://github.com/ForeverBlue816/GRACE.
  •  

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

arXiv:2602.08499v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods utilize rollouts in an indiscriminate and short-horizon manner: responses of heterogeneous quality within each prompt are treated uniformly, and historical rollouts are discarded after a single use. This leads to noisy supervision, poor sample efficiency, and suboptimal policy updates. We address these issues by formulating rollout scheduling in RLVR as a contextual bandit problem and proposing a unified neural scheduling framework that adaptively selects high-value rollouts throughout training. Each rollout is treated as an arm whose reward is defined by the induced performance gain between consecutive optimization steps. The resulting scheduler supports both noise-aware intra-group selection and adaptive global reuse of historical rollouts within a single principled framework. We provide theoretical justification by deriving sublinear regret bounds and showing that enlarging the rollout buffer improves the achievable performance upper bound. Experiments on six mathematical reasoning benchmarks demonstrate consistent gains in performance and training efficiency across multiple RLVR optimization methods.
  •  

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

arXiv:2605.23440v2 Announce Type: replace-cross Abstract: Joint Entity and Relation Extraction (JERE) is highly susceptible to weak generalization due to low-quality training data. Data augmentation is a common strategy to enhance model generalization across different domains. However, existing data augmentation methods often overlook text relevance and may disrupt semantic structures and dependencies, making it difficult to generate effective augmented data for improving model generalization. In this paper, we propose Structured Semantic Data Augmentation (SSDAU), a novel method designed to preserve the semantic structure of text during augmentation. SSDAU segments text based on entity labels and employs an encoder to capture semantic features of entities through context awareness. It then performs entity semantic restructuring to generate augmented data. To distinguish semantically similar entities, SSDAU fuses contextualized embeddings with traditional similarity scores. To mitigate potential topic ambiguity and information loss, we apply the BERTTopic model to filter out irrelevant topics, ensuring topic consistency. We evaluate SSDAU on datasets with different annotation types and compare its performance on five representative JERE models against seven popular data augmentation baselines. Experiments demonstrate that SSDAU generates semantically consistent data with superior robustness against ambiguity (8.26% F1 decrease vs. 31.91% for baselines), significantly outperforming all existing methods across all metrics.
  •  
  •  

Clinical and translational roles of circulating tumor cells in non-small cell and small cell lung cancer: a narrative review

J Thorac Dis. 2026 Apr 30;18(4):415. doi: 10.21037/jtd-2026-1-0025. Epub 2026 Apr 24.

ABSTRACT

BACKGROUND AND OBJECTIVE: Circulating tumor cells (CTCs) are malignant cells shed into blood that enable noninvasive, longitudinal assessment of lung cancer. Increasing evidence frames CTCs within a circulating tumor microenvironment (cTME) and broader circulating tumor-associated cell (CTAC) ecosystems that include multicellular clusters and circulating tumor endothelial cells (CTECs). We summarize definitions, detection approaches, and clinical applications of CTC-centered liquid biopsy in non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC).

METHODS: A comprehensive literature search was conducted in PubMed, Embase, Web of Science, and Google Scholar using the terms "non-small cell lung cancer", "small cell lung cancer", and "circulating tumor cells". Relevant clinical, basic, and translational studies were selected and synthesized to outline current knowledge and future directions.

KEY CONTENT AND FINDINGS: CTCs can be enriched by immunoaffinity, size, or microfluidic platforms, enabling enumeration and downstream profiling. In both NSCLC and SCLC, CTC positivity and higher burden are associated with worse survival, with the strongest effects in SCLC and with circulating tumor emboli (CTE). Serial monitoring provides early signals of response or failure; and post-treatment supports minimal residual disease (MRD) detection and relapse prediction. Molecular and phenotypic profiling enables driver and resistance tracking, including epidermal growth factor receptor (EGFR) and anaplastic lymphoma kinase (ALK), while CTECs may add vascular and immune-relevant information.

CONCLUSIONS: CTC-based assays have the potential to complement imaging and tissue biopsy across screening research, prognostication, therapeutic monitoring, MRD assessment, and personalized care. Clinical translation requires standardized preanalytical workflows, harmonized thresholds, and prospective trials testing CTC-guided management.

PMID:42182710 | PMC:PMC13190041 | DOI:10.21037/jtd-2026-1-0025

  •  

Machine learning-driven multi-omics integration uncovers a senescence associated molecular axis in HCC

Front Immunol. 2026 May 8;17:1762222. doi: 10.3389/fimmu.2026.1762222. eCollection 2026.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) exhibits profound molecular heterogeneity and aberrant cellular senescence. This study systematically dissects the senescence-associated molecular landscape to identify key regulators driving HCC progression and immune evasion.

METHODS: Integrating multi-cohort transcriptomic datasets, we developed a robust prognostic signature using 101 machine-learning models, identifying prognostic signature. We employed preliminary proteomic, exploratory metabolomic, and single-cell RNA sequencing (scRNA-seq) analyses to explore multi-omics alterations. The functional senescence status and MCM7 were validated in a clinical HCC cohort by RT-qPCR, Western blotting, immunohistochemistry, and multiplex immunofluorescence (mIF). Causality was established using in vitro functional assays in HepG2 cells.

RESULTS: A 12-gene random survival forest (RSF) signature accurately predicted patient survival across independent cohorts. MCM7 emerged as a central senescence-associated driver. ScRNA-seq and mIF confirmed MCM7 characterizes a highly proliferative, clonally expanding subset of CD8+ T cells within the tumor microenvironment. In vitro, MCM7 knockdown significantly inhibited HepG2 cell proliferation and upregulated senescence enforcers p16 and p21, whereas overexpression facilitated evasion. Additionally, TIDE analysis revealed that high-risk patients exhibited elevated immune evasion potential, predicting poor immunotherapy response.

CONCLUSION: This integrative multi-omics framework uncovers an MCM7 MCM7-driven senescence-associated axis promising HCC progression and immune dysfunction, offering a robust tool for prognostic stratification and novel therapeutic insights.

PMID:42183188 | PMC:PMC13195000 | DOI:10.3389/fimmu.2026.1762222

  •  
❌