❌

Normal view

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

arXiv:2609.13082v1 Announce Type: new Abstract: Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step construction produces dependent intermediate artifacts that are often passed downstream without artifact-specific verification, allowing local defects to propagate into the final benchmark. We present Embodied-BenchForge, an agentic framework that transforms user-specified evaluation intents into complete embodied benchmark artifacts. It formulates construction as Closed-Loop Benchmark Synthesis, integrating forward artifact synthesis with backward verification and repair. Skill-Orchestrated Artifact Synthesis composes typed and reusable skills into executable workflows, while an artifact dependency graph records intermediate outputs and their dependencies. Requirement-Guided Verification and Repair applies artifact-specific contracts throughout construction and uses provenance to trigger local re-execution or upstream rollback when verification fails. Embodied-BenchForge constructs six benchmarks covering diverse embodied scenarios in the Offline EQA Track, together with one interactive benchmark containing 220 executable tasks in the Interactive Embodied Track. Evaluations of representative MLLMs and embodied agents show that the benchmarks distinguish model capabilities in both observation-based understanding and closed-loop execution. Quality assessment and ablations validate benchmark quality and the effectiveness of verification and repair, while repair and skill-reuse analyses demonstrate efficient localized recovery and cross-benchmark reusability.

Author Correction: A base editor for the long-term restoration of auditory function in mice with recessive profound deafness

Nature Biomedical Engineering, Published online: 07 September 2026; doi:10.1038/s41551-026-01794-5

Author Correction: A base editor for the long-term restoration of auditory function in mice with recessive profound deafness

Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations

arXiv:2605.14175v2 Announce Type: replace Abstract: In a long conversation, an LLM can produce a plausible continuation that rests on premises the conversation has already abandoned. No runtime check ties its output to what the conversation has established, a gap that context-manipulation attacks on deployed agents exploit. We close this gap with a runtime verifier: an LLM Interpreter classifies each utterance into one of eight epistemic operations, and a symbolic engine applies them to a dependency map that records what every claim rests on and whether it still stands. Whether a continuation is grounded reduces to a walk over the map, linear in its size, with no LLM call. Retraction propagates through the same map with a conflict-free guarantee, flagging exactly the conclusions that lose support. On ReviseQA for belief revision and MemoryAgentBench's fact-consolidation split, two third-party benchmarks where earlier premises are superseded, the verifier leads a budget-matched retrieval baseline across five QA models and lifts MemoryAgentBench single-hop accuracy from 0.46--0.95 to 0.93--0.98. With the verifier, even the 7B model overtakes unaided GPT-4o. These runs feed the engine the benchmarks' own structured updates. When a GPT-4o Interpreter extracts every update from raw text instead, accuracy is statistically unchanged. Per-query cost is flat in conversation length, prompts staying near 0.8k tokens where full context reaches 114k and retraction queries under a microsecond at 2000 turns.

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

arXiv:2605.25620v1 Announce Type: new Abstract: World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited semantic structure or inherited from frozen visual foundation models with excessive task-irrelevant detail, yielding state spaces that are poorly matched to downstream planning and control. This is especially challenging in reward-free offline settings, where the model must learn from fixed trajectories without reward supervision or online interaction. To address this, we propose TC-WM, a framework for turning foundation-model embeddings into compact, task-sufficient world representations. The key design is to treat the pretrained embedding space as a semantic scaffold rather than as the final state space: TC-WM linearly projects high-dimensional visual embeddings into a compact latent as the dynamic space, aligns a subspace with the agent's physical state via contrastive learning, and reconstructs embeddings to preserve useful visual structure. This combines the generality of foundation features with the controllability of task-centric dynamics. Theoretically, we show that TC-WM suffices to identify the underlying task-centric latent factors up to a simple transformation. Empirically, TC-WM enables test-time planning across diverse environments (e.g., Robomimic and D4RL), achieving better world-modeling quality and more precise control than state-of-the-art approaches.

Human pancreatic progenitor organoids define genetic and epigenetic barriers to early PDAC transformation

Dev Cell. 2026 May 19:S1534-5807(26)00159-0. doi: 10.1016/j.devcel.2026.04.012. Online ahead of print.

ABSTRACT

The lack of accurate human models that recapitulate pancreatic ductal adenocarcinoma (PDAC) initiation has hindered therapeutic development. Using pluripotent stem cell-derived pancreatic progenitor organoids, we established a human PDAC model that faithfully reproduces the genetic, epigenetic, and transcriptomic trajectories of tumor initiation and progression, validated against clinical datasets and tumor histopathology. We demonstrate that CDKN2A loss, which is nearly universal in patients but dispensable in mouse models, is essential for neoplastic transformation when combined with KRAS and TP53 mutations, whereas SMAD4 loss promotes tumor progression. Multi-omics profiling reveals epigenetic repression of the pancreatic lineage program during PDAC initiation, alongside AP-1-driven chromatin remodeling. We identify TET1 suppression as a mechanistic link between oncogenic ERK signaling and hypermethylation of essential pancreatic transcription factors. This model captures genetic and epigenetic determinants of human PDAC, reveals antagonism between oncogenic and lineage restriction programs, and supports TET-based lineage restoration as a potential early intervention strategy.

PMID:42161274 | PMC:PMC13196429 | DOI:10.1016/j.devcel.2026.04.012

De novo design of quasisymmetric two-component protein cages

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10464-0

Researchers designed two-component proteins forming quasisymmetric cages via geometric frustration, enabling tunable virus-like assemblies for cargo delivery, cellular uptake and studying intracellular diffusion and protein localization.

A Genetically Engineered Human Organoid Model Reveals Distinct Genetic and Epigenetic Barriers of Lineage Plasticity in Early PDAC Transformation

bioRxiv [Preprint]. 2026 Mar 11:2026.03.09.710586. doi: 10.64898/2026.03.09.710586.

ABSTRACT

The lack of accurate, human-based models recapitulating early-stage pancreatic ductal adenocarcinoma (PDAC) has hindered therapeutic development. Using pluripotent stem cell-derived pancreatic progenitor organoids, we established a human PDAC model that faithfully reproduces the genetic, epigenetic, and transcriptomic trajectory of tumor initiation and progression in vitro , validated against clinical datasets and histopathology. We demonstrate that CDKN2A loss, nearly universal in patients but dispensable in mouse models, is essential for neoplastic transformation when combined with KRAS and TP53 mutations, while SMAD4 loss promotes tumor progression. Multi-omics profiling reveals epigenetic repression of pancreatic lineage program during PDAC initiation, alongside oncogenic AP-1-driven chromatin remodeling. Notably, we identify TET1 suppression as a mechanistic link between oncogenic ERK signaling and the hypermethylation and silencing of essential pancreatic transcription factors. This model captures the genetic and epigenetic determinants of human PDAC, reveals antagonism between oncogenic and lineage restriction programs, and supports TET-based lineage restoration as a promising early intervention strategy for high-risk individuals.

PMID:41959451 | PMC:PMC13060829 | DOI:10.64898/2026.03.09.710586

Low-Bitrate Video Compression through Semantic-Conditioned Diffusion

arXiv:2512.00408v2 Announce Type: replace-cross Abstract: Traditional video codecs optimized for pixel fidelity collapse at ultra-low bitrates and produce severe artifacts. This failure arises from a fundamental misalignment between pixel accuracy and human perception. We propose a semantic video compression framework named DiSCo that transmits only the most meaningful information while relying on generative priors for detail synthesis. The source video is decomposed into three compact modalities: a textual description, a spatiotemporally degraded video, and optional sketches or poses that respectively capture semantic, appearance, and motion cues. A conditional video diffusion model then reconstructs high-quality, temporally coherent videos from these compact representations. Temporal forward filling, token interleaving, and modality-specific codecs are proposed to improve multimodal generation and modality compactness. Experiments show that our method outperforms baseline semantic and traditional codecs by 2-10X on perceptual metrics at low bitrates.

Multi-ancestry transcriptome prediction with functionally informed variants in TOPMed MESA improves performance of transcriptome-wide association studies

Am J Hum Genet. 2026 Apr 2;113(4):828-841. doi: 10.1016/j.ajhg.2026.03.008.

ABSTRACT

Reliable reference transcriptome prediction models are key to accurate multi-ancestry transcriptome-wide association studies (TWASs). We propose three methods leveraging functionally informed variants (FIVs) for transcriptome prediction models to improve multi-ancestry TWASs. We trained models on 1,287 multi-ancestry participants from the Trans-Omics for Precision Medicine (TOPMed) program Multi-Ethnic Study of Atherosclerosis (MESA) with RNA sequencing (RNA-seq) data from peripheral blood mononuclear cells (PBMCs). We validated models' prediction accuracy on two external independent datasets, Geuvadis and Jackson Heart Study. To test robustness of our methods for TWASs, we integrated models with three multi-ancestry GWASs from blood cell, lipid, and pulmonary function traits, respectively. Our methods presented similar prediction accuracy while using a smaller and functionally informed set of variants compared to the benchmark method, elastic net (EN). Overall, our methods achieved higher power and accuracy (with average improved accuracy of 24% over EN) for TWASs. However, no single proposed method outperformed all GWAS traits. To further improve TWAS performance, we propose an omnibus approach that aggregates TWAS summary statistics from our methods. The omnibus approach yielded the highest number of Bonferroni-significant TWAS genes for all GWAS traits, and it further improved TWAS power and accuracy for blood cell traits. Additionally, the omnibus approach detected some trait-relevant important genes that the EN missed. Our study demonstrates the value of including FIVs in multi-ancestry transcriptome prediction models for improving TWAS performance. Further, the observed TWAS improvement depends on the GWAS trait's relevance to the PBMCs used to build our transcriptome prediction models.

PMID:41932314 | DOI:10.1016/j.ajhg.2026.03.008

Improving Ensemble Forecasts of Abnormally Deflecting Tropical Cyclones with Fused Atmosphere-Ocean-Terrain Data

arXiv:2603.29200v2 Announce Type: replace-cross Abstract: Deep learning-based tropical cyclone (TC) forecasting methods have demonstrated significant potential and application advantages, as they feature much lower computational cost and faster operation speed than numerical weather prediction models. However, existing deep learning methods still have key limitations: they can only process a single type of sequential trajectory data or homogeneous meteorological variables, and fail to achieve accurate forecasting of abnormal deflected TCs. To address these challenges, we present two groundbreaking contributions. First, we have constructed a multimodal and multi-source dataset named AOT-TCs for TC forecasting in the Northwest Pacific basin. As the first dataset of its kind, it innovatively integrates heterogeneous variables from the atmosphere, ocean, and land, thus obtaining a comprehensive and information-rich meteorological dataset. Second, based on the AOT-TCs dataset, we propose a forecasting model that can handle both normal and abnormally deflected TCs. This is the first TC forecasting model to adopt an explicit atmosphere-ocean-terrain coupling architecture, enabling it to effectively capture complex interactions across physical domains. Extensive experiments on all TC cases in the Northwest Pacific from 2017 to 2024 show that our model achieves state-of-the-art performance in TC forecasting: it not only significantly improves the forecasting accuracy of normal TCs but also breaks through the technical bottleneck in forecasting abnormally deflected TCs.

Improving Ensemble Forecasts of Abnormally Deflecting Tropical Cyclones with Fused Atmosphere-Ocean-Terrain Data

arXiv:2603.29200v1 Announce Type: cross Abstract: Deep learning-based tropical cyclone (TC) forecasting methods have demonstrated significant potential and application advantages, as they feature much lower computational cost and faster operation speed than numerical weather prediction models. However, existing deep learning methods still have key limitations: they can only process a single type of sequential trajectory data or homogeneous meteorological variables, and fail to achieve accurate forecasting of abnormal deflected TCs. To address these challenges, we present two groundbreaking contributions. First, we have constructed a multimodal and multi-source dataset named AOT-TCs for TC forecasting in the Northwest Pacific basin. As the first dataset of its kind, it innovatively integrates heterogeneous variables from the atmosphere, ocean, and land, thus obtaining a comprehensive and information-rich meteorological dataset. Second, based on the AOT-TCs dataset, we propose a forecasting model that can handle both normal and abnormally deflected TCs. This is the first TC forecasting model to adopt an explicit atmosphere-ocean-terrain coupling architecture, enabling it to effectively capture complex interactions across physical domains. Extensive experiments on all TC cases in the Northwest Pacific from 2017 to 2024 show that our model achieves state-of-the-art performance in TC forecasting: it not only significantly improves the forecasting accuracy of normal TCs but also breaks through the technical bottleneck in forecasting abnormally deflected TCs.

Self-Improving Code Generation via Semantic Entropy and Behavioral Consensus

arXiv:2603.29292v1 Announce Type: cross Abstract: Improving the code generation capabilities of large language models (LLMs) typically relies on supervised fine-tuning or preference optimization, both of which require costly external resources such as powerful teacher models or reliable test units. However, in real-world scenarios, it is much harder to obtain reference solutions and test oracles than problem descriptions and test inputs. In this paper, we tackle a challenging yet realistic question: Can a code language model improve itself without access to a superior teacher and a test oracle? To answer this, we propose ConSelf, a self-improving approach built upon two key ideas. First, we introduce code semantic entropy, a novel metric that measures problem-level uncertainty by assessing the functional diversity of program behaviors, enabling a curriculum construction with the most learnable problems. Second, we present consensus-driven direct preference optimization (Con-DPO), a preference-based fine-tuning method that weights each preference pair by its behavioral consensus, thereby mitigating the impact of noisy self-generated supervision. Experiments on various benchmarks and backbone LLMs demonstrate that ConSelf significantly outperforms baselines, validating the effectiveness of semantic entropy-based curriculum construction and consensus-driven optimization in improving code generation without external supervision.

KEditVis: A Visual Analytics System for Knowledge Editing of Large Language Models

arXiv:2603.29689v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate exceptional capabilities in factual question answering, yet they sometimes provide incorrect responses. To address this issue, knowledge editing techniques have emerged as effective methods for correcting factual information in LLMs. However, typical knowledge editing workflows struggle with identifying the optimal set of model layers for editing and rely on summary indicators that provide insufficient guidance. This lack of transparency hinders effective comparison and identification of optimal editing strategies. In this paper, we present KEditVis, a novel visual analytics system designed to assist users in gaining a deeper understanding of knowledge editing through interactive visualizations, improving editing outcomes, and discovering valuable insights for the future development of knowledge editing algorithms. With KEditVis, users can select appropriate layers as the editing target, explore the reasons behind ineffective edits, and perform more targeted and effective edits. Our evaluation, including usage scenarios, expert interviews, and a user study, validates the effectiveness and usability of the system.

LPNSR: Prior-Enhanced Diffusion Image Super-Resolution via LR-Guided Noise Prediction

arXiv:2603.21045v3 Announce Type: replace-cross Abstract: Diffusion-based image super-resolution (SR), which aims to reconstruct high-resolution (HR) images from corresponding low-resolution (LR) observations, faces a fundamental trade-off between inference efficiency and reconstruction quality. The state-of-the-art residual-shifting diffusion framework achieves efficient 4-step inference, yet suffers from severe performance degradation in compact sampling trajectories. This is mainly attributed to two core limitations: the inherent suboptimality of unconstrained random Gaussian noise in intermediate steps, which leads to error accumulation and insufficient LR prior guidance, and the initialization bias caused by naive bicubic upsampling. In this paper, we propose LPNSR, a prior-enhanced efficient diffusion framework to address these issues. We first mathematically derive the closed-form analytical solution of the optimal intermediate noise for the residual-shifting diffusion paradigm, and accordingly design an LR-guided multi-input-aware noise predictor to replace random Gaussian noise, embedding LR structural priors into the reverse process while fully preserving the framework's core efficient residual-shifting mechanism. We further mitigate initial bias with a high-quality pre-upsampling network to optimize the diffusion starting point. With a compact 4-step trajectory, LPNSR can be optimized in an end-to-end manner. Extensive experiments demonstrate that LPNSR achieves state-of-the-art perceptual performance on both synthetic and real-world datasets, without relying on any large-scale text-to-image priors. The source code of our method can be found at https://github.com/Faze-Hsw/LPNSR.

Cellular Senescence in Gastric Cancer: Molecular Mechanisms, Microenvironment Remodeling and Therapeutic Implications

Aging Dis. 2026 Mar 19. doi: 10.14336/AD.2025.1571. Online ahead of print.

ABSTRACT

Gastric cancer (GC) remains a leading cause of cancer-related morbidity and mortality worldwide, with poor prognosis for advanced-stage patients. Therefore, in-depth exploration of the mechanisms underlying GC initiation and progression, as well as the development of novel therapeutic strategies, is of crucial importance. Cellular senescence is a stable cell cycle arrest program that plays a dual role in GC. It exerts tumor-suppressive effects via growth arrest but also promotes tumor progression and immune evasion by remodeling the tumor microenvironment (TME) through senescence-associated secretory phenotype (SASP). This review comprehensively elucidates the molecular mechanisms of cellular senescence in GC and the core regulatory networks involving gene regulation, epigenetic modifications, metabolic reprogramming, and cell cycle arrest. Additionally, the review highlights how senescent cells foster an immunosuppressive microenvironment via SASP, forming a self-reinforcing feed-forward loop. Regarding therapeutic strategies, we summarize potential approaches targeting cellular senescence, including senescence induction, senescent cell clearance, SASP modulation, and multi-target synergistic therapy by integrating epigenetic regulation, metabolic intervention, and immune microenvironment modulation. Despite progress, numerous challenges remain. Future studies should leverage multi-omics technologies, novel models' development, and large-scale clinical trials to advance the clinical translation of GC cellular senescence research, providing new insights for improving prognosis.

PMID:41910653 | DOI:10.14336/AD.2025.1571

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

arXiv:2603.22844v2 Announce Type: new Abstract: Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-based desmoking approaches rely on scarce paired supervision and deterministic restoration pipelines, making it difficult to perform exploration or reinforcement-driven refinement under real surgical conditions. We propose PhySe-RPO, a diffusion restoration framework optimized through Physics- and Semantics-Guided Relative Policy Optimization. The core idea is to transform deterministic restoration into a stochastic policy, enabling trajectory-level exploration and critic-free updates via group-relative optimization. A physics-guided reward imposes illumination and color consistency, while a visual-concept semantic reward learned from CLIP-based surgical concepts promotes smoke-free and anatomically coherent restoration. Together with a reference-free perceptual constraint, PhySe-RPO produces results that are physically consistent, semantically faithful, and clinically interpretable across synthetic and real robotic surgical datasets, providing a principled route to robust diffusion-based restoration under limited paired supervision.

LPNSR: Prior-Enhanced Diffusion Image Super-Resolution via LR-Guided Noise Prediction

arXiv:2603.21045v2 Announce Type: replace-cross Abstract: Diffusion-based image super-resolution (SR), which aims to reconstruct high-resolution (HR) images from corresponding low-resolution (LR) observations, faces a fundamental trade-off between inference efficiency and reconstruction quality. The state-of-the-art residual-shifting diffusion framework achieves efficient 4-step inference, yet suffers from severe performance degradation in compact sampling trajectories. This is mainly attributed to two core limitations: the inherent suboptimality of unconstrained random Gaussian noise in intermediate steps, which leads to error accumulation and insufficient LR prior guidance, and the initialization bias caused by naive bicubic upsampling. In this paper, we propose LPNSR, a prior-enhanced efficient diffusion framework to address these issues. We first mathematically derive the closed-form analytical solution of the optimal intermediate noise for the residual-shifting diffusion paradigm, and accordingly design an LR-guided multi-input-aware noise predictor to replace random Gaussian noise, embedding LR structural priors into the reverse process while fully preserving the framework's core efficient residual-shifting mechanism. We further mitigate initial bias with a high-quality pre-upsampling network to optimize the diffusion starting point. With a compact 4-step trajectory, LPNSR can be optimized in an end-to-end manner. Extensive experiments demonstrate that LPNSR achieves state-of-the-art perceptual performance on both synthetic and real-world datasets, without relying on any large-scale text-to-image priors. The source code of our method can be found at https://github.com/Faze-Hsw/LPNSR.

Multimodal electron microscopy of halide perovskite interfacial dynamics

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10238-8

A multimodal in situ electron microscopy approach enables direct visualization of structural and chemical evolution in a working halide perovskite light-emitting diode with nanometre precision.

Geodesic Gradient Descent: A Generic and Learning-rate-free Optimizer on Objective Function-induced Manifolds

arXiv:2603.06651v1 Announce Type: cross Abstract: Euclidean gradient descent algorithms barely capture the geometry of objective function-induced hypersurfaces and risk driving update trajectories off the hypersurfaces. Riemannian gradient descent algorithms address these issues but fail to represent complex hypersurfaces via a single classic manifold. We propose geodesic gradient descent (GGD), a generic and learning-rate-free Riemannian gradient descent algorithm. At each iteration, GGD uses an n-dimensional sphere to approximate a local neighborhood on the objective function-induced hypersurface, adapting to arbitrarily complex geometries. A tangent vector derived from the Euclidean gradient is projected onto the sphere to form a geodesic, ensuring the update trajectory stays on the hypersurface. Parameter updates are performed using the endpoint of the geodesic. The maximum step size of the gradient in GGD is equal to a quarter of the arc length on the n-dimensional sphere, thus eliminating the need for a learning rate. Experimental results show that compared with the classic Adam algorithm, GGD achieves test MSE reductions ranging from 35.79% to 48.76% for fully connected networks on the Burgers' dataset, and cross-entropy loss reductions ranging from 3.14% to 11.59% for convolutional neural networks on the MNIST dataset.
❌