❌

Normal view

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

arXiv:2605.10764v2 Announce Type: replace-cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.

Kaempferol functionally reprograms CD47 signaling to promote cytoprotection and attenuate oxeiptosis in severe acute pancreatitis

Phytomedicine. 2026 May 15;157:158305. doi: 10.1016/j.phymed.2026.158305. Online ahead of print.

ABSTRACT

BACKGROUND: Severe acute pancreatitis (SAP) lacks targeted therapies, and massive loss of functional pancreatic acinar cells (PAC) drives mortality. Kaempferol (KA) possesses well-established anti-inflammatory and cytoprotective activities and is derived from herbal medicinal plants, but its direct molecular targets and mechanism of action in SAP remain undefined.

PURPOSE: To evaluate the protective effects of KA against SAP and to elucidate its molecular mechanism of specific action, with a focus on identifying the direct cellular target through which KA exerts its cytoprotective effects.

STUDY DESIGN: Gain‑/loss‑of‑function in vitro and PAC‑specific CD47 SAP mouse models, combined with multi‑omics screening and biophysical assays.

METHODS: CD47 manipulation (siRNA/overexpression) was performed in primary PACs and cell lines, combined with WT/CD47-/-/Mist1‑CD47‑iOE (PAC‑specific) mouse models. Network pharmacology, transcriptomics and proteomics were integrated to screen and validate KA's protective effects. Computational‑experimental approaches (molecular docking/dynamics, CETSA, SPR, co‑IP, pharmacological epistasis) characterized KA's allosteric modulation of CD47 signaling.

RESULTS: CD47 was upregulated in SAP; its knockout reduced PAC death via KEAP1/PGAM5/AIFM1-driven oxeiptosis. KA reduced PAC death across genotypes, afforded no extra benefit in CD47-KO, and was not overridden by CD47‑OE. Mechanistically, KA allosterically binds CD47 ectodomain, stabilizes the CD47‑ UBQLN1 complex, and redirects signaling from Gαi‑mediated death to Gβγ/ ERK/NRF2‑mediated survival. ERK inhibition attenuated KA's protection. KA's action was CD47‑dependent.

CONCLUSION: This study identifies anti-oxeiptosis as a novel pharmacological activity of KA in SAP. This is achieved through allosteric modulation of CD47, redirecting its signaling from death‑promoting to a protective axis via activating Gβγ/ERK/NRF2 to suppress oxeiptosis. These findings reveal the CD47‑oxeiptosis axis as a therapeutic target and position KA as a promising candidate for SAP therapy, adding a new mechanistic dimension to KA's known pharmacological profile.

PMID:42184499 | DOI:10.1016/j.phymed.2026.158305

Tex3D: Objects as Attack Surfaces via Adversarial 3D Textures for Vision-Language-Action Models

arXiv:2604.01618v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have shown strong performance in robotic manipulation, yet their robustness to physically realizable adversarial attacks remains underexplored. Existing studies reveal vulnerabilities through language perturbations and 2D visual attacks, but these attack surfaces are either less representative of real deployment or limited in physical realism. In contrast, adversarial 3D textures pose a more physically plausible and damaging threat, as they are naturally attached to manipulated objects and are easier to deploy in physical environments. Bringing adversarial 3D textures to VLA systems is nevertheless nontrivial. A central obstacle is that standard 3D simulators do not provide a differentiable optimization path from the VLA objective function back to object appearance, making it difficult to optimize through an end-to-end manner. To address this, we introduce Foreground-Background Decoupling (FBD), which enables differentiable texture optimization through dual-renderer alignment while preserving the original simulation environment. To further ensure that the attack remains effective across long-horizon and diverse viewpoints in the physical world, we propose Trajectory-Aware Adversarial Optimization (TAAO), which prioritizes behaviorally critical frames and stabilizes optimization with a vertex-based parameterization. Built on these designs, we present Tex3D, the first framework for end-to-end optimization of 3D adversarial textures directly within the VLA simulation environment. Experiments in both simulation and real-robot settings show that Tex3D significantly degrades VLA performance across multiple manipulation tasks, achieving task failure rates of up to 96.7\%. Our empirical results expose critical vulnerabilities of VLA systems to physically grounded 3D adversarial attacks and highlight the need for robustness-aware training.

From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models

arXiv:2511.10788v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have made reasoning a central benchmark for evaluating intelligence. While prior surveys focus on efficiency by examining how to shorten reasoning chains or reduce computation, this view overlooks a fundamental challenge: current LLMs apply uniform reasoning strategies regardless of task complexity, generating long traces for trivial problems while failing to extend reasoning for difficult tasks. This survey reframes reasoning through the lens of {adaptivity}: the capability to allocate reasoning effort based on input characteristics such as difficulty and uncertainty. We make three contributions. First, we formalize deductive, inductive, and abductive reasoning within the LLM context, connecting these classical cognitive paradigms with their algorithmic realizations. Second, we formalize adaptive reasoning as a control-augmented policy optimization problem balancing task performance with computational cost, distinguishing learned policies from inference-time control mechanisms. Third, we propose a systematic taxonomy organizing existing methods into training-based approaches that internalize adaptivity through reinforcement learning, supervised fine-tuning, and learned controllers, and training-free approaches that achieve adaptivity through prompt conditioning, feedback-driven halting, and modular composition. This framework clarifies how different mechanisms realize adaptive reasoning in practice and enables systematic comparison across diverse strategies. We conclude by identifying open challenges in self-evaluation, meta-reasoning, and human-aligned reasoning control.

AI-guided multi-omics analysis identifies NPC1-modulated susceptibility to SARS-CoV-2 infection under PM(2.5) exposure

Nat Commun. 2026 Mar 30. doi: 10.1038/s41467-026-71196-3. Online ahead of print.

ABSTRACT

Exposure to airborne fine particulate matter (PM2.5) has been linked to increased risk of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, yet the underlying mechanisms remain unclear. Here, by leveraging a fine-tuned foundation model of single-cell transcriptomics, we uncover shared transcriptional signatures between PM2.5 exposure and SARS-CoV-2 infection. We further validate this association using population-level epidemiological analyses and perform genome-wide association studies (GWAS) to identify genetic variants that modulate infection risk under PM2.5 exposure. In addition, we identify NPC1 as a key modulator involved in SARS-CoV-2 infection efficiency under virus-laden PM2.5 exposure through integrative functional genomic analyses and in vitro experiments. Our findings suggest that PM2.5 facilitates viral entry through an NPC1-modulated endo-lysosomal pathway, providing a mechanistic explanation for observed pollution-related susceptibility. By integrating artificial intelligence (AI)-guided transcriptomics, epidemiology, GWAS, functional genomics, and in vitro verification, our study elucidates how environmental and genetic factors jointly influence SARS-CoV-2 susceptibility. This work highlights how AI-assisted multi-omics integration systematically decodes the health impacts of environmental exposures from molecular to population levels and informs air quality policy and infectious disease preparedness.

PMID:41912520 | DOI:10.1038/s41467-026-71196-3

Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases

arXiv:2603.22767v1 Announce Type: new Abstract: Observational studies can yield clinically actionable evidence at scale, but executing them on real-world databases is open-ended and requires coherent decisions across cohort construction, analysis, and reporting. Prior evaluations of LLM agents emphasize isolated steps or single answers, missing the integrity and internal structure of the resulting evidence bundle. To address this gap, we introduce RWE-bench, a benchmark grounded in MIMIC-IV and derived from peer-reviewed observational studies. Each task provides the corresponding study protocol as the reference standard, requiring agents to execute experiments in a real database and iteratively generate tree-structured evidence bundles. We evaluate six LLMs (three open-source, three closed-source) under three agent scaffolds using both question-level correctness and end-to-end task metrics. Across 162 tasks, task success is low: the best agent reaches 39.9%, and the best open-source model reaches 30.4%. Agent scaffolds also matter substantially, causing over 30% variation in performance metrics. Furthermore, we implement an automated cohort evaluation method to rapidly localize errors and identify agent failure modes. Overall, the results highlight persistent limitations in agents' ability to produce end-to-end evidence bundles, and efficient validation remains an important direction for future work. Code and data are available at https://github.com/somewordstoolate/RWE-bench.

ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models

arXiv:2505.13180v2 Announce Type: replace Abstract: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent work extending this idea to visual domains using Vision-Language Models (VLMs). However, a rigorous comparison with methods that plan directly with VLMs is missing, due to a lack of visual benchmarks that support symbolic planning. We present ViPlan, the first open-source benchmark for comparing VLM-grounded symbolic approaches (VLM-as-grounder) with direct VLM planning methods (VLM-as-planner). ViPlan introduces a series of increasingly challenging tasks in two visual domains: a visual variant of the classic Blocksworld planning problem and a simulated household robotics environment. We find VLM-as-grounder methods to outperform direct VLM planning in Blocksworld (solving 46% of the tasks against 9%), where image grounding is both crucial and accurate. However, in the household robotics tasks, where linguistic knowledge helps, VLM-as-planner methods are greatly superior to VLM-as-grounder approaches (solving 34% of the tasks against 5%), which are hindered by partial observability. Thus, ViPlan domains capture fundamental shortcomings of both planning approaches, which we further diagnose with a qualitative failure analysis. Finally, across methods, we observe no consistent benefit from Chain-of-Thought prompting, suggesting persistent limitations in current VLMs' visual reasoning abilities.
❌