❌

Normal view

Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks

arXiv:2609.09698v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have achieved remarkable success across diverse applications, yet they remain highly vulnerable to adversarial attacks that maliciously perturb graph structure. Existing defenses often lack rigorous theoretical grounding, rely on attack-specific heuristics, or require costly retraining procedures such as adversarial training. To address these limitations, we propose Kernel-Complexity Edge Sanitization (KCES), a training-free and model-agnostic framework for defending against structural attacks. KCES is built upon Graph Kernel Complexity (GKC), a principled metric derived from the graph Gram matrix that appears in a generalization upper bound on the GNN test error. From this bound, we define an edge-specific KC score that quantifies each edge's structural influence via its induced change in GKC. KCES then identifies and prunes high-KC edges, which are empirically enriched with adversarial perturbations under structural attacks, to mitigate their harmful impact. Computationally efficient and scalable, KCES operates as a lightweight preprocessing step without retraining and can be seamlessly integrated with existing defenses. Extensive experiments demonstrate that KCES consistently outperforms representative robust baselines across diverse attack settings and scales effectively to large graphs. Supported by theoretical analysis and extensive empirical validation, KCES provides a principled and efficient framework for securing GNNs. Our code is available at https://github.com/karpning/KCScore.

Skin-innervating glutamatergic neurons modulate aging

Within the skin, glutamatergic neurons expressing neurofilament heavy chain (Nefh) play a role in aging. Loss of Nefh during aging drives skin fibroblast senescence and collagen loss, whereas glutamate supplementation improves skin aging phenotypes.

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm.

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

arXiv:2605.24041v2 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure that struggles to resolve high-frequency details, a limitation known as spectral bias. We introduce the Iterative Refinement Neural Operator (IRNO), which augments pre-trained operators with a learned refinement module iteratively applied via fixed-point iteration. IRNO decomposes the prediction into a coarse initialization followed by successive residual corrections, paralleling classical numerical solvers. Under local assumptions, we establish contraction of the induced operator, ensuring convergence to a unique fixed point. To explicitly target high-frequency errors, we propose a progressive spectral loss that adaptively increases penalty on high-frequency components over refinement steps during training. Across physical systems, IRNO consistently lowers error, with up to 56.05% improvement on turbulent flow. On Active Matter, spectral analysis reveals that, relative to base operator, the normalized error ratios decrease to 27.72-36.10% in low-, 5.07-6.68% in mid-, and 1.48-2.04% in high-frequencies, remaining stable beyond the trained iteration count. Code is available at https://github.com/xiaotianliu-dartmouth/Iterative_Refinement_Neural_Operator

NSR-Boost: A Neuro-Symbolic Residual Boosting Framework for Industrial Legacy Models

arXiv:2601.10457v3 Announce Type: replace Abstract: Although the Gradient Boosted Decision Trees (GBDTs) dominate industrial tabular applications, upgrading legacy models in high-concurrency production environments still faces prohibitive retraining costs and systemic risks. To address this problem, we present NSR-Boost, a neuro-symbolic residual boosting framework designed specifically for industrial scenarios. Its core advantage lies in being ``non-intrusive''. It treats the legacy model as a frozen model and performs targeted repairs on "hard regions" where predictions fail. The framework comprises three key stages: First, finding hard regions through residuals, then generating interpretable experts by generating symbolic code structures using Large Language Model (LLM) and fine-tuning parameters using Bayesian optimization, and finally dynamically integrating experts with legacy model output through a lightweight aggregator. Experimental results demonstrate that the framework significantly outperforms state-of-the-art (SOTA) baselines across six public datasets and one private dataset. More importantly, we report the successful deployment of NSR-Boost within the core financial risk control system of Qfin Holdings, where empirical results on real-world online traffic exhibit superior performance improvements and a significant reduction in the bad rate. In conclusion, it effectively captures long-tail risks missed by traditional models and offers a safe, low-cost evolutionary paradigm for industry.

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

arXiv:2605.23904v2 Announce Type: replace Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization. Code: https://aka.ms/skillopt

Forest carbon protocols underestimate climate-driven carbon loss risks

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10571-y

The buffer pool designed to compensate for unintended carbon losses from the largest forest climate mitigation programme in the United States is too small when considering the impact of future climate change scenarios.

Stromal ACTA2 Counteracts TCDD-Induced Hepatocarcinogenesis via Suppression of the PI3K-AKT-mTOR Pathway

J Hepatocell Carcinoma. 2026 May 10;13:586916. doi: 10.2147/JHC.S586916. eCollection 2026.

ABSTRACT

PURPOSE: 2,3,7,8-Tetrachlorodibenzo-p-dioxin (TCDD) is a persistent environmental pollutant that promotes hepatocellular carcinoma (HCC) through non-genotoxic mechanisms. However, stromal regulatory factors that counteract its tumor-promoting effects remain poorly defined. This study aimed to elucidate the role of actin alpha-2 (ACTA2) in TCDD-associated hepatocarcinogenesis.

METHODS: An integrative strategy combining network toxicology, Mendelian randomization, multi-omics and single-cell analyses, molecular docking and molecular dynamics simulations, along with in vitro experiments, was employed to investigate the functional role of ACTA2.

RESULTS: ACTA2 was identified as a stromal-associated factor linked to reduced HCC risk and improved patient survival. Single-cell and multi-omics analyses revealed that ACTA2 is predominantly expressed in hepatic stellate cells and fibroblast-like populations, reflecting tumor microenvironment composition rather than tumor cell-intrinsic expression. Functional enrichment analyses indicated that ACTA2 is associated with extracellular matrix remodeling and PI3K-AKT signaling. Molecular simulations demonstrated stable binding of TCDD to ACTA2 (ΔG_bind ≈ -7.05 kcal/mol), suggesting potential structural perturbation. In vitro experiments showed that TCDD downregulated ACTA2 expression, promoted proliferation of LX-2 and cancer-associated fibroblasts (CAFs), and activated PI3K-AKT-mTOR signaling, whereas ACTA2 overexpression attenuated these effects.

CONCLUSION: ACTA2 acts as a context-dependent stromal regulator that modulates PI3K-AKT-mTOR signaling in TCDD-induced hepatocarcinogenesis. These findings highlight the importance of stromal remodeling in environmental carcinogenesis and suggest ACTA2 as a potential biomarker and therapeutic target in dioxin-associated HCC.

PMID:42148320 | PMC:PMC13175077 | DOI:10.2147/JHC.S586916

Refined immune-based molecular subtypes of gastric cancer: Integrating mismatch repair status and tumor microenvironment for enhanced immunotherapy prediction

Chin J Cancer Res. 2026 Apr 30;38(2):234-251. doi: 10.21147/j.issn.1000-9604.2026.02.09.

ABSTRACT

OBJECTIVE: Gastric cancer (GC) is heterogeneous, and current mismatch repair (MMR)-based classifications incompletely predict response to immune checkpoint inhibitors (ICIs).

METHODS: RNA sequencing (RNA-seq) and immune infiltration profiles from 189 resected GC were used to derive four refined immune-MMR subtypes (R1-R4) by integrating MMR status, survival, and tumor microenvironment (TME) features. Multi-omics profiling and pathway analysis defined subtype biology. External transcriptomic cohorts and an ICI-treated cohort were classified with Nearest Template Prediction (NTP). Immune response-associated genes were identified from responder vs. non-responder comparisons within the ICI-sensitive subtype and validated by multiplex immunohistochemistry (mIHC).

RESULTS: R1 showed the best prognosis and highest immunotherapy response with objective response rate (ORR) 54.5%, while R4 had the worst prognosis. R2 represented an immune-unresponsive deficient mismatch repair (dMMR) subset, and R3 captured an immune-active proficient mismatch repair (pMMR) subgroup with moderate therapy sensitivity. Multi-omics integration revealed subtype-specific pathways (e.g., ECM remodeling in R1, metabolic reprogramming in R2). Reclassification of pMMR tumors based on transcriptional similarity to R1 identified a New R3 subset with enhanced immune features and higher ICI response. Eight immune response-associated genes (e.g., CXCL10, CXCL11, ELN, GAD1, IL32, MT1E, OR2I1P, SLC3A1) were identified and validated by mIHC for predictive relevance.

CONCLUSIONS: This immune-based molecular framework refines risk stratification beyond conventional MMR categories, identifies ICI-sensitive subsets among both dMMR and pMMR tumors, and proposes candidate biomarkers for patient selection.

PMID:42147371 | PMC:PMC13171420 | DOI:10.21147/j.issn.1000-9604.2026.02.09

SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling

arXiv:2603.23414v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) has shown strong promise for enhancing the reasoning abilities of large language models (LLMs), particularly in tasks requiring long chain-of-thought generation. However, RL training efficiency is often bottlenecked by the rollout phase, which can account for up to 70% of total training time when generating long trajectories (e.g., 16k tokens), due to slow autoregressive generation and synchronization overhead between rollout and policy updates. We propose SortedRL, an online length-aware scheduling strategy designed to address this bottleneck by improving rollout efficiency and maintaining training stability. SortedRL reorders rollout samples based on output lengths, prioritizing short samples forming groups for early updates. This enables large rollout batches, flexible update batches, and near on-policy micro-curriculum construction simultaneously. To further accelerate the pipeline, SortedRL incorporates a mechanism to control the degree of off-policy training through a cache-based mechanism, and is supported by a dedicated RL infrastructure that manages rollout and update via a stateful controller and rollout buffer. Experiments using LLaMA-3.1-8B and Qwen-2.5-32B on diverse tasks, including logical puzzles, and math challenges like AIME 24, Math 500, and Minerval, show that SortedRL reduces RL training bubble ratios by over 50%, while attaining 3.9% to 18.4% superior performance over baseline given same amount of data.

FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications

arXiv:2601.00150v3 Announce Type: replace-cross Abstract: FCMBench is the first large-scale and privacy-compliant multimodal benchmark for real-world financial credit applications, covering tasks and robustness challenges from domain specific workflows and constraints. The current version of FCMBench covers 26 certificate types, with 5198 privacy-compliant images and 13806 paired VQA samples. It evaluates models on Perception and Reasoning tasks under real-world Robustness interferences, including 3 foundational perception tasks, 4 credit-specific reasoning tasks demanding decision-oriented visual evidence interpretation, and 10 real-world challenges for rigorous robustness stress testing. Moreover, FCMBench offers privacy-compliant realism with minimal leakage risk through in-house scenario-aware captures of manually synthesized templates, without any publicly released images. We conduct extensive evaluations of 28 state-of-the-art vision-language models spanning 14 AI companies and research institutes. Among them, Gemini 3 Pro achieves the best F1 score as a commercial model (65.16), Kimi-K2.5 achieves the best score as an open-source baseline (60.58). The mean and the std. of all tested models is 44.8 and 10.3 respectively, indicating that FCMBench is non-trivial and provides strong resolution for separating modern vision-language model capabilities. Robustness evaluations reveal that even top-performing models experience notable performance degradation under the designed challenges. We have open-sourced this benchmark to advance AI research in the credit domain and provide a domain-specific task for real-world AI applications.

When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning

arXiv:2603.02266v1 Announce Type: cross Abstract: Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-training for direct answering. To investigate it, we introduce CAFE, an evaluation framework designed to precisely quantify audio reasoning errors. Evaluation results reveal LALMs struggle with perception during reasoning and encounter a critical bottleneck: reasoning performance suffers from audio perception decay as reasoning length extends. To address it, we propose MPAR$^2$, a paradigm that encourages dynamic perceptual reasoning and decomposes complex questions into perception-rich sub-problems. Leveraging reinforcement learning, MPAR$^2$ improves perception performance on CAFE from 31.74% to 63.51% and effectively mitigates perception decay, concurrently enhancing reasoning capabilities to achieve a significant 74.59% accuracy on the MMAU benchmark. Further analysis demonstrates that MPAR$^2$ reinforces LALMs to attend to audio input and dynamically adapts reasoning budget to match task complexity.

Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment

arXiv:2503.15937v5 Announce Type: replace Abstract: We propose V-Droid, a mobile GUI task automation agent. Unlike previous mobile agents that utilize Large Language Models (LLMs) as generators to directly generate actions at each step, V-Droid employs LLMs as verifiers to evaluate candidate actions before making final decisions. To realize this novel paradigm, we introduce a comprehensive framework for constructing verifier-driven mobile agents: the discretized action space construction coupled with the prefilling-only workflow to accelerate the verification process, the pair-wise progress preference training to significantly enhance the verifier's decision-making capabilities, and the scalable human-agent joint annotation scheme to efficiently collect the necessary data at scale. V-Droid obtains a substantial task success rate across several public mobile task automation benchmarks: 59.5% on AndroidWorld, 38.3% on AndroidLab, and 49% on MobileAgentBench, surpassing existing agents by 5.2%, 2.1%, and 9%, respectively. Furthermore, V-Droid achieves a remarkably low latency of 4.3s per step, which is 6.1x faster compared with existing mobile agents. The source code is available at https://github.com/V-Droid-Agent/V-Droid.

Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis

arXiv:2602.07135v3 Announce Type: replace-cross Abstract: Loss landscapes are a powerful tool for understanding neural network optimization and generalization, yet traditional low-dimensional analyses often miss complex topological features. We present Landscaper, an open-source Python package for arbitrary-dimensional loss landscape analysis. Landscaper combines Hessian-based subspace construction with topological data analysis to reveal geometric structures such as basin hierarchy and connectivity. A key component is the Saddle-Minimum Average Distance (SMAD) for quantifying landscape smoothness. We demonstrate Landscaper's effectiveness across various architectures and tasks, including those involving pre-trained language models, showing that SMAD captures training transitions, such as landscape simplification, that conventional metrics miss. We also illustrate Landscaper's performance in challenging chemical property prediction tasks, where SMAD can serve as a metric for out-of-distribution generalization, offering valuable insights for model diagnostics and architecture design in data-scarce scientific machine learning scenarios.
❌