❌

Normal view

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

arXiv:2609.10142v1 Announce Type: cross Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time, yet the mechanism behind its lasting protection remains unclear. Analyzing its temporal optimization dynamics, we find that the defense emerges from an early compensatory adaptation phase followed by a steady-state phase where the corrective signal decays; in parameter space, attention output projections emerge as the dominant residual-write route for defensive updates. Through Intervention Delta Preservation (IDP) and IDP Continuation experiments, we further show that preserving or reinjecting the weight offset fails to maintain protection, indicating that preventative steering relies on active adaptation rather than a static defense. Motivated by this finding, we propose Progressive Intensity Scheduling (PIS), which starts with a moderate injection strength and increases it after static-strength alignment begins to decay. Across the evaluated Qwen2.5 and Gemma-3 models, PIS improves safety robustness over static-strength steering while reducing harmful trait expression.

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

arXiv:2608.30935v2 Announce Type: replace-cross Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Pulmonary nodule prediction in the multi-omics era: Integrating radiomics, AI, liquid biopsy, and airway classifiers

10 July 2026 at 18:00

Crit Rev Oncol Hematol. 2026 Sep;225:105483. doi: 10.1016/j.critrevonc.2026.105483. Epub 2026 Jul 10.

ABSTRACT

Low-dose CT (LDCT) lung cancer screening significantly reduces mortality but has dramatically increased the detection of pulmonary nodules. Most of these nodules are benign, leading to a high false-positive rate that triggers unnecessary invasive procedures and patient anxiety, underscoring the need for more precise noninvasive diagnostic tools. Critically, single-modality liquid biopsy biomarkers, including circulating tumor cells, cell-free DNA mutations, or individual microRNAs, have demonstrated insufficient sensitivity or specificity for independent clinical deployment when used in isolation. This necessitates a paradigm shift toward multimodal molecular integration, wherein complementary biomarker classes are combined to overcome the inherent limitations of any single analyte. Traditional clinical prediction models (Mayo, VA, Brock, Herder) assist in estimating malignancy risk, yet their accuracy remains modest. Emerging approaches harness radiomics and artificial intelligence (AI) to extract high-dimensional imaging features from chest CT scans, improving risk stratification beyond human assessment alone. In parallel, minimally invasive liquid biopsy biomarkers offer complementary avenues to detect occult malignancy signals. Additionally, bronchial airway gene expression classifiers leverage the "field-of-injury" effect in normal respiratory epithelium to help identify lung cancer even when the nodule itself cannot be directly sampled via biopsy. Integrating these radiologic and molecular data streams into a multi-omics framework has the potential to enhance diagnostic precision for indeterminate pulmonary nodules, enabling more confident discrimination between benign and malignant lesions. However, most of these emerging tools have not yet been validated in large prospective trials and face technological barriers as well as challenges in real-world implementation. This review focuses primarily on LDCT screening detected pulmonary nodules, while incorporating evidence from incidentally detected and other indeterminate nodule cohorts when relevant to broader CT based management. By synthesizing advances in radiomics, AI, liquid biopsy, airway classifiers, and multi-omics integration, we highlight the need for prospective validation and multidisciplinary collaboration to translate these approaches into clinically useful pathways that improve early lung cancer detection, reduce unnecessary interventions, and enhance patient outcomes.

PMID:42431477 | DOI:10.1016/j.critrevonc.2026.105483

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

arXiv:2605.24993v1 Announce Type: new Abstract: Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically faithful surface-based models. We argue this is partly driven by inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present NeurIPS, a framework that improves surface-based decoding by reframing anatomical variation from a nuisance to a powerful inductive prior. NeurIPS unites two innovations: a Selective ROI Spherical Tokenizer (SRST) for efficient geometric encoding, and a Structure-Guided Mixture of Experts (SG-MoE) that explicitly models individual anatomy using cortical features. On the Natural Scenes Dataset, NeurIPS establishes a new state-of-the-art for surface decoders and achieves performance comparable to strong 1D baselines. This is achieved with unprecedented efficiency, as the model converges dramatically faster (10 vs. 600 epochs). This efficiency enables rapid adaptation to new subjects using only 20% of data and ensures robust scalability as the training cohort is expanded. Ablations provide causal evidence that these gains are driven by the model's use of cortical features, not by memorizing subject IDs. By leveraging anatomical priors, NeurIPS provides a principled and scalable path toward robust, generalizable brain decoding.

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows. However, real-world execution environments are far from ideal: pop-ups, resolution changes, and competing applications frequently interfere with agent perception and control. We introduce AgentHijack, a benchmark designed to evaluate the robustness of computer-use agents under common corruptions, where the uncertainties in dynamic environment disrupt the execution flow without direct adversarial intent. Specifically, AgentHijack introduces 9 configurable common corruptions to replicate realistic imperfect scenarios. We evaluate a variety of desktop tasks that utilize MLLM-based agents and discover that even minor instances of corruption can result in substantial performance degradation, which emphasizes the fragility of agents and underscores the necessity of robustness evaluation. Afterward, we propose AgentHijack-Agent, a framework that integrates an action generator with enhanced grounding capabilities and an onlooker responsible for behavior summarization and environment checking. Extensive experiments validate its effectiveness. Our code, environment, baseline models and data are publicly available at: https://AgentHijack.github.io.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

arXiv:2605.23954v1 Announce Type: cross Abstract: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18\%$\uparrow$ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02\%$\uparrow$ in Acc, 3.89\%$\uparrow$ in Noisy, and 4.53\%$\uparrow$ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented autonomous execution capabilities for LLM Agents by integrating external open-domain knowledge and tools. However, this interoperability introduces a covert attack surface targeting the agent's cognitive planning layer. This paper systematically investigates Tool Description Poisoning (TDP), a novel semantic attack. In TDP, malicious instructions are not embedded in a tool's executable code, but rather covertly injected into its descriptive metadata, the very "manual" an agent relies on for secure planning and decision-making. To rigorously and systematically evaluate this emerging threat, we introduce the MCP-TDP Security Benchmark. This high-fidelity sandbox environment comprises 32 realistic, real-world test cases spanning 6 distinct risk categories. Our evaluation of 8 mainstream LLMs reveals severe vulnerabilities, with leading models like GPT-4o exhibiting a nearly 100% Attack Success Rate (ASR) in six high-risk scenarios. Furthermore, our findings demonstrate that common prompt-guardrail defenses are largely ineffective and can, counterintuitively, even be counterproductive (a phenomenon which we term the "Firewall Fallacy"). Crucially, we also propose a defense mechanism: "Reactive Self-Correction," where an agent autonomously detects and reverts its own malicious actions post-execution. This work provides the first specialized security benchmark tailored for TDP, offering essential insights for securing the cognitive and planning layers of advanced agentic systems.

Multi-omics analysis of glutamine and fish collagen peptides in alleviating post-antibiotic Streptococcus pneumoniae injury in feline lung cells

Exp Ther Med. 2026 Mar 30;31(6):148. doi: 10.3892/etm.2026.13143. eCollection 2026 Jun.

ABSTRACT

Streptococcus pneumoniae (SP) infection often leads to persistent lung injury even after antibiotic treatment. Despite this phenomenon, the mechanisms underlying host cell recovery remain poorly understood. Upon breaching the epithelial barrier, SP primarily targets the pulmonary interstitial cells, which constitute the major mesenchymal component of the lung. These cells serve as essential effectors of tissue repair, extracellular matrix remodeling and epithelial restoration. Therefore, a feline pulmonary interstitial cell (FCA-L2) model of SP infection was established to investigate the protective effects of glutamine (GLU) and fish collagen peptides (FCP) through integrated transcriptomic and metabolomic analyses. Cells were infected with SP (0.05 McFarland units for 4 h) and then treated with doxycycline (7.5 µg/ml for 18 h) followed by GLU (40 mM) or FCP (500 µg/ml). Notably, SP infection increased lactate dehydrogenase (LDH) release by 3.5-fold, induced secretion of IL-1β, TNF-α and IL-8, disrupted tight-junction proteins (claudin, ZO-1 and occludin) and caused oxidative imbalance and apoptosis despite antibiotic (doxycycline) treatment. However, treatment with GLU or FCP significantly reduced LDH release by ~40%, restored junctional proteins, suppressed inflammatory cytokines and enhanced antioxidant enzyme activities. Multi-omics analysis revealed that GLU promoted amino acid biosynthesis and energy metabolism and suppressed aminoacyl-tRNA synthetases and cell-cycle regulators, thereby enhancing metabolic adaptability. By contrast, FCP activated amino and nucleotide sugar metabolism, increased polyunsaturated fatty-acid synthesis and supported glycocalyx repair and membrane reconstruction. GLU and FCP provided complementary metabolic and structural protection, which mitigated post-infectious stress and promoted cellular recovery. The findings of the present study underscore the potential of bioactive food-derived compounds as adjunctive therapies that may accelerate lung tissue repair and enhance the efficacy of conventional antibiotics.

PMID:41988354 | PMC:PMC13077270 | DOI:10.3892/etm.2026.13143

Correction: Steroid receptor coactivator-1 facilitates METTL3-mediated m6A modification by coactivating NF-κB and promotes the malignant progression of glioblastoma

Oncogene, Published online: 15 April 2026; doi:10.1038/s41388-026-03788-8

Correction: Steroid receptor coactivator-1 facilitates METTL3-mediated m6A modification by coactivating NF-κB and promotes the malignant progression of glioblastoma

Effect of the Maxing Huoqiao granule on nonsevere community-acquired pneumonia: A multicenter, double-blind, placebo-controlled randomized trial

Pharmacol Res. 2026 Apr 9:108186. doi: 10.1016/j.phrs.2026.108186. Online ahead of print.

ABSTRACT

Community-acquired pneumonia (CAP) remains a major global public health challenge with substantial morbidity and mortality. Although preclinical studies suggest that Maxing Huoqiao (MXHQ) granule may have therapeutic potential for pneumonia, high-quality clinical evidence is still limited. We conducted a multicenter, double-blind, randomized, placebo-controlled trial at two tertiary hospitals in China to evaluate the clinical efficacy of MXHQ as adjunctive therapy and to explore its potential mechanisms in adults with nonsevere CAP receiving standard moxifloxacin treatment. A total of 96 patients were enrolled and randomized (1:1:1) to receive standard-dose MXHQ, low-dose MXHQ, or placebo in addition to moxifloxacin for 7 days, with a 14-day follow-up. The primary endpoint was clinical cure, defined as composite recovery of major respiratory symptoms, lung rales, and fever; secondary endpoints included symptom relief, radiographic improvement, and safety. Compared with placebo, standard-dose MXHQ was associated with a higher day-14 clinical cure rate (30.78% vs. 68.97%; RR = 0.45, 95% CI = 0.24-0.83; P < 0.01). Furthermore, the standard-dose intervention was correlated with a shorter time to relief and recovery of cough and sputum (P < 0.05), as well as improvements in symptom scores (P < 0.05) and promoting lesion absorption on chest CT (P < 0.05). Low-dose MXHQ showed no significant clinical benefit, whereas safety profiles were comparable across all groups. Transcriptomic analyses of peripheral blood mononuclear cells, complemented by a Streptococcus pneumonia animal model, indicated that the clinical benefits of MXHQ are linked to the modulation of inflammation and innate immunity. These omics and in vivo observations suggest a potential mechanism underlying the protective effects of MXHQ against inflammatory injury and promotion of tissue repair, involving the regulation of anti-inflammatory mediators and tissue repair-related factors. (Chictr.org.cn, ID Number: ChiCTR2400082095).

PMID:41966499 | DOI:10.1016/j.phrs.2026.108186

Graphicalized vision-language modeling for comprehensive lung nodule analysis and risk stratification

npj Digital Medicine, Published online: 11 April 2026; doi:10.1038/s41746-026-02602-9

Graphicalized vision-language modeling for comprehensive lung nodule analysis and risk stratification

Clinical application of base editing for treating β-thalassaemia

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10342-9

A clinical phase 1 trial of a single infusion of CS-101, CD34+ cells modified using a transformer base editor to reactivate fetal haemoglobin production, led to early and enduring transfusion independence in patients with β-thalassaemia.

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

arXiv:2604.04247v1 Announce Type: new Abstract: Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based on previous agent runs. However, these methods primarily focus on single-agent or low-parallelism settings. This fundamentally limits their ability to efficiently learn from a large set of collected agentic traces. It would be efficient and beneficial to run prompt learning in parallel to accommodate the growing trend of learning from many agentic traces or parallel agent executions. Yet without a principled strategy for scaling, current methods suffer from quality degradation with high parallelism. To improve both the efficiency and quality of prompt learning, we propose Combee, a novel framework to scale parallel prompt learning for self-improving agents. Combee speeds up learning and enables running many agents in parallel while learning from their aggregate traces without quality degradation. To achieve this, Combee leverages parallel scans and employs an augmented shuffle mechanism; Combee also introduces a dynamic batch size controller to balance quality and delay. Evaluations on AppWorld, Terminal-Bench, Formula, and FiNER demonstrate that Combee achieves up to 17x speedup over previous methods with comparable or better accuracy and equivalent cost.

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics

arXiv:2604.03315v1 Announce Type: cross Abstract: Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter-shot consistency and explicit editability. While 2D diffusion-based generators produce vivid imagery, they often suffer from identity drift along with limited geometric control; conversely, traditional 3D animation workflows are consistent and editable but require expert-heavy, labor-intensive authoring. We present StoryBlender, a grounded 3D storyboard generation framework governed by a Story-centric Reflection Scheme. At its core, we propose the StoryBlender system, which is built on a three-stage pipeline: (1) Semantic-Spatial Grounding, to construct a continuity memory graph to decouple global assets from shot-specific variables for long-horizon consistency; (2) Canonical Asset Materialization, to instantiate entities in a unified coordinate space to maintain visual identity; and (3) Spatial-Temporal Dynamics, to achieve layout design and cinematic evolution through visual metrics. By orchestrating multiple agents in a hierarchical manner within a verification loop, StoryBlender iteratively self-corrects spatial hallucinations via engine-verified feedback. The resulting native 3D scenes support direct, precise editing of cameras and visual assets while preserving unwavering multi-shot continuity. Experiments demonstrate that StoryBlender significantly improves consistency and editability over both diffusion-based and 3D-grounded baselines. Code, data, and demonstration video will be available on https://engineeringai-lab.github.io/StoryBlender/

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing

arXiv:2604.04875v1 Announce Type: cross Abstract: Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio-visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross-level multimodal orchestration to achieve professional-grade fluidity, resulting in disjointed sequences with abrupt visual transitions and musical misalignment. To address this, we formulate video mashup creation as a Multimodal Coherency Satisfaction Problem (MMCSP) and propose the DIRECT framework. Simulating a professional production pipeline, our hierarchical multi-agent framework decomposes the challenge into three cascade levels: the Screenwriter for source-aware global structural anchoring, the Director for instantiating adaptive editing intent and guidance, and the Editor for intent-guided shot sequence editing with fine-grained optimization. We further introduce Mashup-Bench, a comprehensive benchmark with tailored metrics for visual continuity and auditory alignment. Extensive experiments demonstrate that DIRECT significantly outperforms state-of-the-art baselines in both objective metrics and human subjective evaluation. Project page and code: https://github.com/AK-DREAM/DIRECT

TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization

arXiv:2601.22776v2 Announce Type: replace Abstract: Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (RL) frameworks for search-augmented reasoning predominantly rely on sparse outcome-level rewards, leading to a "Double Homogenization Dilemma." This manifests as (1) Process homogenization, where the thinking, reasoning, and tooling involved in generation are ignored. (2) Intra-group homogenization, coarse-grained outcome rewards often lead to inefficiencies in intra-group advantage estimation with methods like Group Relative Policy Optimization (GRPO) during sampling. To address this, we propose Turn-level Stage-aware Policy Optimization (TSPO). TSPO introduces the First-Occurrence Latent Reward (FOLR) mechanism, allocating partial rewards to the step where the ground-truth answer first appears, thereby preserving process-level signals and increasing reward variance within groups without requiring external reward models or any annotations. Extensive experiments demonstrate that TSPO significantly outperforms state-of-the-art baselines, achieving average performance gains of 24% and 13.6% on Qwen2.5-3B and 7B models, respectively. Code is available at https://github.com/Flipped-May/TSPO.

Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering

arXiv:2501.07813v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agents are increasingly deployed to answer questions over local knowledge bases that cannot be centralized due to knowledge-sovereignty constraints. This results in two recurring failures in production: users do not know which agent to consult, and complex questions require evidence distributed across multiple agents. To overcome these challenges, we propose RIRS, a training-free orchestration framework to enable a multi-agent system for question answering. In detail, RIRS summarizes each agent's local corpus in an embedding space, enabling a user-facing server to route queries only to the most relevant agents, reducing latency and avoiding noisy "broadcast-to-all" contexts. For complicated questions, the server can iteratively aggregate responses to derive intermediate results and refine the question to bridge the gap toward a comprehensive answer. Extensive experiments demonstrate the effectiveness of RIRS, including its ability to precisely select agents and provide accurate responses to single-hop queries, and its use of an iterative strategy to achieve accurate, multi-step resolutions for complex queries.

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

arXiv:2512.02902v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose a one-shot adaptation framework that recalibrates visual representations through lightweight, learnable updates. Our first method, Feature Token Modulation (FTM), applies a global affine transformation to visual tokens and improves Libero viewpoint accuracy from 48.5% to 87.1% with only 4K parameters. Building on this, Feature Linear Adaptation (FLA) introduces low-rank updates to the ViT encoder, achieving 90.8% success with 4.7M parameters -- matching LoRA-scale finetuning at far lower cost. Together, these results reveal substantial untapped robustness in pretrained VLA models and demonstrate that targeted, minimal visual adaptation is sufficient to restore viewpoint generalization.
❌