❌

Normal view

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.

TM184C is a GPCR-like regulator of intercellular exchange and autophagy

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10993-8

TM184C—an ancient G-protein-coupled receptor-like superdark protein involved in regulation of autophagy, intercellular connectivity and material exchange—underscores the promise of exploring the understudied human proteome and beyond.

Itaconate and its derivatives in human health and diseases

Signal Transduct Target Ther. 2026 Sep 4;11(1):363. doi: 10.1038/s41392-026-02936-6.

ABSTRACT

Metabolic reprogramming forms the foundation of immune effector functions and the regulation of inflammation. As a pivotal node connecting the tricarboxylic acid cycle to immune signaling, the IRG1/ACOD1 and itaconate axes play a central role in coordinating inflammatory tone and redox balance. Itaconate, generated through the decarboxylation of cis aconitate, acts as an immunometabolic brake that engages multiple regulatory pathways to sustain the dynamic equilibrium between inflammation and tissue homeostasis. Across a broad spectrum of pathological conditions, including infectious diseases, metabolic disorders, ischemia‒reperfusion injury, neurodegenerative diseases, autoimmune disorders, and cancers, itaconate and its derivatives generally exert anti-inflammatory and cytoprotective effects. However, within specific microenvironments, these molecules may also be exploited by pathogens to evade immune clearance or promote immunosuppressive and protumorigenic responses. Future studies should further elucidate tissue- and lineage-specific functions, define bidirectional regulatory mechanisms, and optimize the pharmacokinetic properties of itaconate derivatives. With the advancement of multiomics integration, systems immunology, rational drug design, and engineered itaconate delivery technologies, the IRG1/ACOD1-itaconate axis and derivative-based therapeutic strategies are poised to emerge as key metabolic checkpoints and therapeutic targets in inflammatory-, metabolic-, immune-, and cancer-related diseases.

PMID:42693110 | PMC:PMC13542262 | DOI:10.1038/s41392-026-02936-6

CDO1 as a prognostic biomarker and therapeutic target in gastric cancer: Mechanistic insights into the PI3K/AKT-THBS1 axis and epigenetic reactivation by decitabine

Clin Transl Med. 2026 Sep;16(9):e70784. doi: 10.1002/ctm2.70784.

ABSTRACT

BACKGROUND: As a pivotal metabolic enzyme, cysteine dioxygenase type 1 (CDO1) exerts tumour-suppressive effects across diverse tumour types, and its expression is strongly correlated with clinical prognosis. However, the molecular mechanisms underlying CDO1-mediated tumour suppression in gastric cancer (GC), its relationship with the tumour-associated immune microenvironment, and pharmacological strategies to restore its expression remain poorly understood.

METHODS: CDO1 expression and prognosis were evaluated by multi-omics and tissue microarray analyses. Tumour microenvironment and immune infiltration were analyzed using ESTIMATE and ssGSEA. Downstream pathways and interacting proteins were identified by transcriptomics, co-immunoprecipitation, and GST pull-down. CDO1 function was assessed by proliferation, apoptosis, and migration assays in gain- and loss-of-function models. In vivo tumorigenesis and CDO1-dependent decitabine efficacy were evaluated by subcutaneous xenografts. Patient-derived organoids were used to assess decitabine sensitivity and 5-FU synergy.

RESULTS: Compared with normal controls, CDO1 expression was notably decreased in GC tissues, and its low expression was strongly linked to unfavourable prognosis, supporting its utility as a biomarker for prognosis. Elevated CDO1 levels correlated with an immune-active tumour microenvironment and reduced metastatic signatures. Mechanistically, CDO1 directly bound to PI3K p85α, disrupting p85α-p110α dimerization, thereby attenuating PI3K/AKT phosphorylation and downregulating THBS1 expression. CDO1 overexpression led to reduced proliferation, invasiveness, and EMT, accompanied by increased apoptosis. These effects were reversed by PI3K activation or THBS1 co-overexpression. Decitabine was identified as an agent that epigenetically restores CDO1 expression. Critically, CDO1 knockdown significantly attenuated the anti-tumour efficacy of decitabine in vivo, confirming that decitabine acts primarily through CDO1 reactivation. Decitabine synergized with 5-FU in both organoids and xenografts.

CONCLUSIONS: Our data identify CDO1 as both a biomarker for prognosis and a tumour suppressor in gastric cancer. They reveal a CDO1-PI3K/AKT-THBS1 signalling axis and support the epigenetic reactivation of CDO1 by decitabine as a translatable therapeutic strategy.

KEY POINTS: CDO1 is frequently downregulated in gastric cancer and serves as an independent favourable prognostic biomarker. CDO1 directly binds PI3K p85α, disrupting p85α-p110α dimerization to suppress the PI3K/AKT-THBS1 signalling axis. Decitabine epigenetically restores CDO1 expression, and its anti-tumour activity is critically CDO1-dependent in vivo. Combining decitabine with 5-FU synergistically overcomes gastric cancer growth in patient-derived organoids and subcutaneous xenograft models.

PMID:42670236 | PMC:PMC13527532 | DOI:10.1002/ctm2.70784

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

arXiv:2605.24154v1 Announce Type: new Abstract: Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textsc{Palette}, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. \textsc{Palette} further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that \textsc{Palette} delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor bit allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. To deliver the on-device benefits of our aggressive quantization, we further introduce OmniModel.cpp, an agentic conversion pipeline that ports architectures into a native C/C++ runtime with efficient low-bit kernels. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm, with all models deployed through OmniModel.cpp. On the LIBERO benchmark, ActQuant is the only method that operates at or below 3 bits-per-weight, retaining 95.0% on OpenVLA-OFT and 94.8% on $\pi_{0.5}$. Pushed further, ActQuant reaches 2.5 bpw at 90.1% on OpenVLA-OFT, compressing the backbone from 14.3 GB to 2.7 GB (5.3$\times$). On the physical UR3 arm, $\pi_{0.5}$ quantized with ActQuant retains the baseline's success rate while reducing the memory footprint by 2.5$\times$.

De novo design of quasisymmetric two-component protein cages

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10464-0

Researchers designed two-component proteins forming quasisymmetric cages via geometric frustration, enabling tunable virus-like assemblies for cargo delivery, cellular uptake and studying intracellular diffusion and protein localization.

Design of one-component quasisymmetric protein nanocages

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10554-z

Quasisymmetry could arise from spontaneous symmetry breaking in a system of strongly interacting building blocks with programmed curvatures, and this principle, coupled with a design approach, can generate a rich array of quasisymmetric assemblies.

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

arXiv:2603.16673v3 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the environment. However, invoking LLM reasoning introduces substantial computational latency and resource overhead, which can interrupt action execution and reduce system reliability. Excessive reasoning may delay actions, while insufficient reasoning often leads to incorrect decisions and task failures. This raises a fundamental question for embodied agents: when should the agent reason, and when should it act? In this work, we propose RARRL (Resource-Aware Reasoning via Reinforcement Learning), a hierarchical framework for resource-aware orchestration of embodied agents. Rather than learning low-level control policies, RARRL learns a high-level orchestration policy that operates at the agent's decision-making layer. This policy enables the agent to adaptively determine whether to invoke reasoning, which reasoning role to employ, and how much computational budget to allocate based on current observations, execution history, and remaining resources. Extensive experiments, including evaluations with empirical latency profiles derived from the ALFRED benchmark, show that RARRL consistently improves task success rates while reducing execution latency and enhancing robustness compared with fixed or heuristic reasoning strategies. These results demonstrate that adaptive reasoning control is essential for building reliable and efficient embodied robotic agents.

Perspectives from machine learning and multi-omics to decoding the effects of VDAC2 malignant subsets on tumor evolution

NPJ Precis Oncol. 2026 Mar 31. doi: 10.1038/s41698-026-01394-1. Online ahead of print.

ABSTRACT

VDAC2's known role in cancer and immune regulation via enhancing the CD8+ T cell-mediated killing, and it is worth systematically digging out the role of VDAC2 in pan-cancer based on this research. Bulk RNA sequencing, single-cell RNA sequencing, and spatial transcriptomic analyses were utilized to explore the role of VDAC2 from multiple perspectives in pan-cancers. RT-PCR, cell co-culture, CCK-8 assay, Transwell invasion assays, and ELISA were performed to validate the expression level and biological function. VDAC2 was upregulated in the majority of pan-cancers, and functional enrichment analyses displayed that VDAC2 may take part in the biological progress of energy metabolism, mitochondrial damage and cell proliferation. The landscape of VDAC2 expression and immune infiltration was constructed, and the VDAC2-BAK1-IFNγ pathway was identified in digestive cancer. VDAC2 had the potential to serve as a novel prognostic, screening cancer indicator and immune therapeutic target sensitive to various drugs. Overexpression of VDAC2 significantly promoted gastric cancer cell proliferation, invasion and immune invasion, as validated in vitro experiments. In short, our pan-cancer analysis constructed a comprehensive landscape of VDAC2's oncogenic role, establishing VDAC2 + -BAK1-IFNγ as an important pathway in tumor progression and immune evasion. VDAC2 emerges not only as a valuable prognostic biomarker but also as a promising novel therapeutic target.

PMID:41917254 | DOI:10.1038/s41698-026-01394-1

Perspectives from machine learning and multi-omics to decoding the effects of VDAC2 malignant subsets on tumor evolution

NPJ Precis Oncol. 2026 Mar 31. doi: 10.1038/s41698-026-01394-1. Online ahead of print.

ABSTRACT

VDAC2's known role in cancer and immune regulation via enhancing the CD8+ T cell-mediated killing, and it is worth systematically digging out the role of VDAC2 in pan-cancer based on this research. Bulk RNA sequencing, single-cell RNA sequencing, and spatial transcriptomic analyses were utilized to explore the role of VDAC2 from multiple perspectives in pan-cancers. RT-PCR, cell co-culture, CCK-8 assay, Transwell invasion assays, and ELISA were performed to validate the expression level and biological function. VDAC2 was upregulated in the majority of pan-cancers, and functional enrichment analyses displayed that VDAC2 may take part in the biological progress of energy metabolism, mitochondrial damage and cell proliferation. The landscape of VDAC2 expression and immune infiltration was constructed, and the VDAC2-BAK1-IFNγ pathway was identified in digestive cancer. VDAC2 had the potential to serve as a novel prognostic, screening cancer indicator and immune therapeutic target sensitive to various drugs. Overexpression of VDAC2 significantly promoted gastric cancer cell proliferation, invasion and immune invasion, as validated in vitro experiments. In short, our pan-cancer analysis constructed a comprehensive landscape of VDAC2's oncogenic role, establishing VDAC2 + -BAK1-IFNγ as an important pathway in tumor progression and immune evasion. VDAC2 emerges not only as a valuable prognostic biomarker but also as a promising novel therapeutic target.

PMID:41917254 | DOI:10.1038/s41698-026-01394-1

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

arXiv:2603.22435v1 Announce Type: cross Abstract: "Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.

From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG

arXiv:2603.03292v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields. While Retrieval-Augmented Generation (RAG) mitigates these issues, existing methods rely on noisy token-level signals and lack the multi-round refinement required for complex reasoning. In the paper, we propose MA-RAG (Multi-Round Agentic RAG), a framework that facilitates test-time scaling for complex medical reasoning by iteratively evolving both external evidence and internal reasoning history within an agentic refinement loop. At each round, the agent transforms semantic conflict among candidate responses into actionable queries to retrieve external evidence, while optimizing history reasoning traces to mitigate long-context degradation. MA-RAG extends the self-consistency principle by leveraging the lack of consistency as a proactive signal for multi-round agentic reasoning and retrieval, and mirrors a boosting mechanism that iteratively minimizes the residual error toward a stable, high-fidelity medical consensus. Extensive evaluations across 7 medical Q&A benchmarks show that MA-RAG consistently surpasses competitive inference-time scaling and RAG baselines, delivering substantial +6.8 points on average accuracy over the backbone model. Our code is available at https://github.com/NJU-RL/MA-RAG.

From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG

arXiv:2603.03292v1 Announce Type: cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields. While Retrieval-Augmented Generation (RAG) mitigates these issues, existing methods rely on noisy token-level signals and lack the multi-round refinement required for complex reasoning. In the paper, we propose **MA-RAG** (**M**ulti-Round **A**gentic RAG), a framework that facilitates test-time scaling for complex medical reasoning by iteratively evolving both external evidence and internal reasoning history within an agentic refinement loop. At each round, the agent transforms semantic **conflict** among candidate responses into actionable queries to retrieve external evidence, while optimizing history reasoning traces to mitigate long-context degradation. MA-RAG extends the *self-consistency* principle by leveraging the lack of consistency as a proactive signal for multi-round agentic reasoning and retrieval, and mirrors a *boosting* mechanism that iteratively minimizes the residual error toward a stable, high-fidelity medical **consensus**. Extensive evaluations across 7 medical Q&A benchmarks show that MA-RAG consistently surpasses competitive inference-time scaling and RAG baselines, delivering **substantial +6.8 points** on average accuracy over the backbone model. Our code is available at [this url](https://github.com/NJU-RL/MA-RAG).

CubeComposer: Spatio-Temporal Autoregressive 4K 360{\deg} Video Generation from Perspective Video

arXiv:2603.04291v1 Announce Type: cross Abstract: Generating high-quality 360{\deg} panoramic videos from perspective input is one of the crucial applications for virtual reality (VR), whereby high-resolution videos are especially important for immersive experience. Existing methods are constrained by computational limitations of vanilla diffusion models, only supporting $\leq$ 1K resolution native generation and relying on suboptimal post super-resolution to increase resolution. We introduce CubeComposer, a novel spatio-temporal autoregressive diffusion model that natively generates 4K-resolution 360{\deg} videos. By decomposing videos into cubemap representations with six faces, CubeComposer autoregressively synthesizes content in a well-planned spatio-temporal order, reducing memory demands while enabling high-resolution output. Specifically, to address challenges in multi-dimensional autoregression, we propose: (1) a spatio-temporal autoregressive strategy that orchestrates 360{\deg} video generation across cube faces and time windows for coherent synthesis; (2) a cube face context management mechanism, equipped with a sparse context attention design to improve efficiency; and (3) continuity-aware techniques, including cube-aware positional encoding, padding, and blending to eliminate boundary seams. Extensive experiments on benchmark datasets demonstrate that CubeComposer outperforms state-of-the-art methods in native resolution and visual quality, supporting practical VR application scenarios. Project page: https://lg-li.github.io/project/cubecomposer

GLEAN: Grounded Lightweight Evaluation Anchors for Contamination-Aware Tabular Reasoning

arXiv:2603.02212v1 Announce Type: cross Abstract: Tabular reasoning benchmarks mix semantic inference, numerical computation, and brittle table formatting, yet evaluations for small models remain vulnerable to contamination, dataset artifacts, and retrieval failures. We propose GLEAN, a lightweight evaluation protocol that integrates contamination-aware probes, weak-supervision governance, retrieval-reasoning diagnostics, and structured error attribution under tight hardware constraints. We evaluate across TabFact, WTQ via Squall, TableBench, RobuT, and SciTab under a 16GB GPU budget. Using Squall gold SQL as an executable anchor (95.2% execution), GLEAN assigns a deterministic error taxonomy (L0-L4 plus L0.5 context miss) and reveals a stable error-mode separation: TAPEX errors skew toward grounding (L3) while TAPAS errors skew toward hallucination/abstention (L2/L0). We validate evidence-row heuristics against SQL-derived rows on simple queries (0.62 precision / 0.71 recall; hybrid recall 0.81) and show that retrieval Recall@K can saturate even when end-to-end EM/F1 remains limited, motivating attribution beyond raw recall. We release a modular framework with audits and sensitivity checks to make small-model tabular evaluation more contamination-aware and diagnostic.

Proactive Guiding Strategy for Item-side Fairness in Interactive Recommendation

arXiv:2603.03094v1 Announce Type: cross Abstract: Item-side fairness is crucial for ensuring the fair exposure of long-tail items in interactive recommender systems. Existing approaches promote the exposure of long-tail items by directly incorporating them into recommended results. This causes misalignment between user preferences and the recommended long-tail items, which hinders long-term user engagement and reduces the effectiveness of recommendations. We aim for a proactive fairness-guiding strategy, which actively guides user preferences toward long-tail items while preserving user satisfaction during the interactive recommendation process. To this end, we propose HRL4PFG, an interactive recommendation framework that leverages hierarchical reinforcement learning to guide user preferences toward long-tail items progressively. HRL4PFG operates through a macro-level process that generates fairness-guided targets based on multi-step feedback, and a micro-level process that fine-tunes recommendations in real time according to both these targets and evolving user preferences. Extensive experiments show that HRL4PFG improves cumulative interaction rewards and maximum user interaction length by a larger margin when compared with state-of-the-art methods in interactive recommendation environments.

A Very Big Video Reasoning Suite

arXiv:2602.20159v1 Announce Type: cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/ .

Diversity-Incentivized Exploration for Versatile Reasoning

arXiv:2509.26209v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a crucial paradigm for incentivizing reasoning capabilities in Large Language Models (LLMs). Due to vast state-action spaces and reward sparsity in reasoning tasks, existing methods often struggle with deficient exploration and poor sample efficiency. In the paper, we propose \textbf{DIVER} (\textbf{D}iversity-\textbf{I}ncentivized Exploration for \textbf{V}ersatil\textbf{E} \textbf{R}easoning), an innovative framework that highlights the pivotal role of global sequence-level diversity to incentivize deep exploration for versatile reasoning. We first conduct a primary empirical study to reveal a strong positive correlation between global diversity and reasoning capacity. Building on this insight, we introduce global diversity incentives as an intrinsic reward to promote deep exploration in a semantically structured space. Incorporating the intrinsic reward, we develop a potential-based reward shaping mechanism to preserve optimal policy invariance and design simple heuristics to mitigate possible reward hacking. Experimental results show that DIVER outperforms competitive RLVR baselines with various exploration strategies on both in-domain and out-of-domain tasks, excelling in both Pass@1 and Pass@k evaluations. Our code is available at https://github.com/NJU-RL/DIVER.

WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity

arXiv:2602.14452v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer strong capabilities but incur high inference costs due to dense computation and memory access. Training-free activation sparsity is a promising approach for efficient LLM inference, yet existing methods often rely solely on activation information and uniform sparsity ratios. This overlooks the critical interplay with weights and inter-block sensitivity variation, leading to suboptimal performance. We identify two key phenomena in modern LLMs: 1) less significant activations may align with highly important weights, and 2) sparsity sensitivity varies non-monotonically across model blocks. We propose Weight-aware Mixed-Granularity Training-free Activation Sparsity (WiSparse), which leverages both activation and weight information for adaptive sparsity allocation. Specifically, we introduce a weight-aware mechanism integrating activation magnitudes with precomputed weight norms to accurately identify salient channels. This is combined with a mixed-granularity allocation scheme: a global budget is distributed across blocks via evolutionary search to protect sensitive regions, then refined within blocks to minimize reconstruction error. We improve sparse kernels and demonstrate effectiveness on three representative models. Notably, at 50% sparsity, WiSparse preserves 97% of Llama3.1's dense performance, surpassing the strongest baseline by 2.23 percentage points while achieving a 21.4% acceleration in end-to-end inference speed. Our research advances the limits of training-free approaches for efficient LLM inference, pushing the boundaries of achievable speedup without training.
❌