❌

Reading view

SelfGrader: Stable Jailbreak Detection for Large Language Models using Token-Level Logits

arXiv:2604.01473v1 Announce Type: cross Abstract: Large Language Models (LLMs) are powerful tools for answering user queries, yet they remain highly vulnerable to jailbreak attacks. Existing guardrail methods typically rely on internal features or textual responses to detect malicious queries, which either introduce substantial latency or suffer from the randomness in text generation. To overcome these limitations, we propose SelfGrader, a lightweight guardrail method that formulates jailbreak detection as a numerical grading problem using token-level logits. Specifically, SelfGrader evaluates the safety of a user query within a compact set of numerical tokens (NTs) (e.g., 0-9) and interprets their logit distribution as an internal safety signal. To align these signals with human intuition of maliciousness, SelfGrader introduces a dual-perspective scoring rule that considers both the maliciousness and benignness of the query, yielding a stable and interpretable score that reflects harmfulness and reduces the false positive rate simultaneously. Extensive experiments across diverse jailbreak benchmarks, multiple LLMs, and state-of-the-art guardrail baselines demonstrate that SelfGrader achieves up to a 22.66% reduction in ASR on LLaMA-3-8B, while maintaining significantly lower memory overhead (up to 173x) and latency (up to 26x).
  •  

Editing strigolactone hormone receptor for robust antiviral silencing in rice

Precise genome editing of the rice strigolactone receptor DWARF14 confers robust, transgene-free antiviral resistance by blocking viral suppression of endogenous RNA silencing, offering a promising strategy for durable disease protection without a yield penalty.
  •  

Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries

arXiv:2603.29500v1 Announce Type: new Abstract: Large language models (LLMs) have recently demonstrated impressive performance on complex, multi-step reasoning tasks, especially when post-trained with outcome-rewarded reinforcement learning Guo et al. 2025. However, it has been observed that outcome rewards often overlook flawed intermediate steps, leading to unreliable reasoning steps even when final answers are correct. To address this unreliable reasoning, we propose PRoSFI (Process Reward over Structured Formal Intermediates), a novel reward method that enhances reasoning reliability without compromising accuracy. Instead of generating formal proofs directly, which is rarely accomplishable for a modest-sized (7B) model, the model outputs structured intermediate steps aligned with its natural language reasoning. Each step is then verified by a formal prover. Only fully validated reasoning chains receive high rewards. The integration of formal verification guides the model towards generating step-by-step machine-checkable proofs, thereby yielding more credible final answers. PRoSFI offers a simple and effective approach to training trustworthy reasoning models.
  •  

$V_0$: A Generalist Value Model for Any Policy at State Zero

arXiv:2602.03584v2 Announce Type: replace-cross Abstract: Policy gradient methods rely on a baseline to measure the relative advantage of an action, ensuring the model reinforces behaviors that outperform its current average capability. In the training of Large Language Models (LLMs) using Actor-Critic methods (e.g., PPO), this baseline is typically estimated by a Value Model (Critic) often as large as the policy model itself. However, as the policy continuously evolves, the value model requires expensive, synchronous incremental training to accurately track the shifting capabilities of the policy. To avoid this overhead, Group Relative Policy Optimization (GRPO) eliminates the coupled value model by using the average reward of a group of rollouts as the baseline; yet, this approach necessitates extensive sampling to maintain estimation stability. In this paper, we propose $V_0$, a Generalist Value Model capable of estimating the expected performance of any model on unseen prompts without requiring parameter updates. We reframe value estimation by treating the policy's dynamic capability as an explicit context input; specifically, we leverage a history of instruction-performance pairs to dynamically profile the model, departing from the traditional paradigm that relies on parameter fitting to perceive capability shifts. Focusing on value estimation at State Zero (i.e., the initial prompt, hence $V_0$), our model serves as a critical resource scheduler. During GRPO training, $V_0$ predicts success rates prior to rollout, allowing for efficient sampling budget allocation; during deployment, it functions as a router, dispatching instructions to the most cost-effective and suitable model. Empirical results demonstrate that $V_0$ significantly outperforms heuristic budget allocation and achieves a Pareto-optimal trade-off between performance and cost in LLM routing tasks.
  •  

In vivo generation of anti-BCMA CAR-T cells in relapsed or refractory multiple myeloma: a phase 1 study

Nature Medicine, Published online: 25 March 2026; doi:10.1038/s41591-026-04244-6

In a phase 1 trial, the in vivo generation of anti-BCMA CAR-T cells by lentiviral delivery was feasible and did not lead to dose-limiting toxicities in five patients with relapsed or refractory multiple myeloma.
  •  

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

  •  

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

  •  

Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices

arXiv:2510.05109v5 Announce Type: replace-cross Abstract: Large Multimodal Models (LMMs) are inherently modular, consisting of vision and audio encoders, projectors, and large language models. Yet, they are almost always executed monolithically, which underutilizes the heterogeneous accelerators (NPUs, GPUs, DSPs) in modern SoCs and leads to high end-to-end latency. In this paper, we present NANOMIND, a hardware--software co-design inference framework for Large Multimodal Models (LMMs) that breaks large models into modular ``bricks'' (vision, language, audio, etc.) and maps each to its ideal accelerator. The key insight is that large models can be broken into modular components and scheduled to run on the most appropriate compute units. It performs module-level dynamic offloading across accelerators on unified-memory SoCs. By combining customized hardware design, system-level scheduling, and optimized low-bit computation kernels, we demonstrate our framework with a compact, battery-powered device capable of running LMMs entirely on device. This prototype functions as a self-contained intelligent assistant that requires no network connectivity, while achieving higher throughput and superior power efficiency under strict resource constraints. The design further bypasses CPU bottlenecks and reduces redundant memory usage through token-aware buffer management and module-level coordination. Our system outperforms existing implementations in resource efficiency, cutting energy consumption by 42.3\% and GPU memory usage by 11.2\%. This enables a battery-powered device to run LLaVA-OneVision with a camera for nearly 20.8 hours.
  •  

Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing

arXiv:2511.01743v2 Announce Type: replace-cross Abstract: Recent advancements in large artificial intelligence models (LAMs) are driving significant innovations in mobile edge computing within next-generation wireless networks. However, the substantial demands for computational resources and larges-cale training data required to train LAMs conflict with the limited storage and computational capacity of edge devices, posing significant challenges to training and deploying LAMs at the edge. In this work, we introduce the Networked Mixture-of-Experts (NMoE) system, in which clients perform inference collaboratively by distributing tasks to suitable neighbors based on their expertise and aggregate the returned results. For training the NMoE, we propose a federated learning framework that integrates both supervised and self-supervised learning to balance personalization and generalization, while preserving communication efficiency and data privacy. We conduct extensive experiments to demonstrate the efficacy of the proposed NMoE system, providing insights for the NMoE training algorithms.
  •  

Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?

arXiv:2601.12349v2 Announce Type: replace-cross Abstract: Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted with perceiving screen content and injecting inputs. However, their design operates under the implicit assumption of Visual Atomicity: that the UI state remains invariant between observation and action. We demonstrate that this assumption is fundamentally invalid in Android, creating a critical attack surface. We present Action Rebinding, a novel attack that allows a seemingly-benign app with zero dangerous permissions to rebind an agent's execution. By exploiting the inevitable observation-to-action gap inherent in the agent's reasoning pipeline, the attacker triggers foreground transitions to rebind the agent's planned action toward the target app. We weaponize the agent's task-recovery logic and Android's UI state preservation to orchestrate programmable, multi-step attack chains. Furthermore, we introduce an Intent Alignment Strategy (IAS) that manipulates the agent's reasoning process to rationalize UI states, enabling it to bypass verification gates (e.g., confirmation dialogs) that would otherwise be rejected. We evaluate Action Rebinding Attacks on six widely-used Android GUI agents across 15 tasks. Our results demonstrate a 100% success rate for atomic action rebinding and the ability to reliably orchestrate multi-step attack chains. With IAS, the success rate in bypassing verification gates increases (from 0% to up to 100%). Notably, the attacker application requires no sensitive permissions and contains no privileged API calls, achieving a 0% detection rate across malware scanners (e.g., VirusTotal). Our findings reveal a fundamental architectural flaw in current agent-OS integration and provide critical insights for the secure design of future agent systems. To access experimental logs and demonstration videos, please contact yi_qian@smail.nju.edu.cn.
  •  

SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning

arXiv:2602.19455v1 Announce Type: cross Abstract: Time-series diagnostic reasoning is essential for many applications, yet existing solutions face a persistent gap: general reasoning large language models (GRLMs) possess strong reasoning skills but lack the domain-specific knowledge to understand complex time-series patterns. Conversely, fine-tuned time-series LLMs (TSLMs) understand these patterns but lack the capacity to generalize reasoning for more complicated questions. To bridge this gap, we propose a hybrid knowledge-injection framework that injects TSLM-generated insights directly into GRLM's reasoning trace, thereby achieving strong time-series reasoning with in-domain knowledge. As collecting data for knowledge injection fine-tuning is costly, we further leverage a reinforcement learning-based approach with verifiable rewards (RLVR) to elicit knowledge-rich traces without human supervision, then transfer such an in-domain thinking trace into GRLM for efficient knowledge injection. We further release SenTSR-Bench, a multivariate time-series-based diagnostic reasoning benchmark collected from real-world industrial operations. Across SenTSR-Bench and other public datasets, our method consistently surpasses TSLMs by 9.1%-26.1% and GRLMs by 7.9%-22.4%, delivering robust, context-aware time-series diagnostic insights.
  •  

PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling

arXiv:2602.19685v1 Announce Type: cross Abstract: Building Virtual Cells that can accurately simulate cellular responses to perturbations is a long-standing goal in systems biology. A fundamental challenge is that high-throughput single-cell sequencing is destructive: the same cell cannot be observed both before and after a perturbation. Thus, perturbation prediction requires mapping unpaired control and perturbed populations. Existing models address this by learning maps between distributions, but typically assume a single fixed response distribution when conditioned on observed cellular context (e.g., cell type) and the perturbation type. In reality, responses vary systematically due to unobservable latent factors such as microenvironmental fluctuations and complex batch effects, forming a manifold of possible distributions for the same observed conditions. To account for this variability, we introduce PerturbDiff, which shifts modeling from individual cells to entire distributions. By embedding distributions as points in a Hilbert space, we define a diffusion-based generative process operating directly over probability distributions. This allows PerturbDiff to capture population-level response shifts across hidden factors. Benchmarks on established datasets show that PerturbDiff achieves state-of-the-art performance in single-cell response prediction and generalizes substantially better to unseen perturbations. See our project page (https://katarinayuan.github.io/PerturbDiff-ProjectPage/), where code and data will be made publicly available (https://github.com/DeepGraphLearning/PerturbDiff).
  •  

FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation

arXiv:2602.13444v1 Announce Type: cross Abstract: Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An embodiment-agnostic interaction representation that captures this structure would make manipulation behaviors easier to validate and transfer across robots. We propose FlowHOI, a two-stage flow-matching framework that generates semantically grounded, temporally coherent HOI sequences, comprising hand poses, object poses, and hand-object contact states, conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry-centric grasping from semantics-centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion-text alignment loss to semantically ground the generated interactions in both the physical scene layout and the language instruction. To address the scarcity of high-fidelity HOI supervision, we introduce a reconstruction pipeline that recovers aligned hand-object trajectories and meshes from large-scale egocentric videos, yielding an HOI prior for robust generation. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7$\times$ higher physics simulation success rate than the strongest diffusion-based baseline, while delivering a 40$\times$ inference speedup. We further demonstrate real-robot execution on four dexterous manipulation tasks, illustrating the feasibility of retargeting generated HOI representations to real-robot execution pipelines.
  •  

Reinforcement Learning with Promising Tokens for Large Language Models

arXiv:2602.03195v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a key paradigm for aligning and optimizing large language models (LLMs). Standard approaches treat the LLM as the policy and apply RL directly over the full vocabulary space. However, this formulation includes the massive tail of contextually irrelevant tokens in the action space, which could distract the policy from focusing on decision-making among the truly reasonable tokens. In this work, we verify that valid reasoning paths could inherently concentrate within a low-rank subspace. Based on this insight, we introduce Reinforcement Learning with Promising Tokens (RLPT), a framework that mitigates the action space issue by decoupling strategic decision-making from token generation. Specifically, RLPT leverages the semantic priors of the base model to identify a dynamic set of promising tokens and constrains policy optimization exclusively to this refined subset via masking. Theoretical analysis and empirical results demonstrate that RLPT effectively reduces gradient variance, stabilizes the training process, and improves sample efficiency. Experiment results on math, coding, and telecom reasoning show that RLPT outperforms standard RL baselines and integrates effectively across various model sizes (4B and 8B) and RL algorithms (GRPO and DAPO).
  •  
❌