❌

Normal view

Safe Learning Under Irreversible Dynamics via Asking for Help

arXiv:2502.14043v4 Announce Type: replace-cross Abstract: Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.

Segmented poly(A) tails with microRNA target sites confer tissue-specific regulation for mRNA therapeutics

Zhang and colleagues engineered the poly(A) tail as a programmable regulatory element, showing that embedding cell-type-specific microRNA target sites directly within it confers robust, position-dependent silencing in off-target tissues while preserving activity in target cells. This strategy offers a new modular tool to enhance mRNA therapeutic safety and precision.

Safe Learning Under Irreversible Dynamics via Asking for Help

arXiv:2502.14043v3 Announce Type: replace-cross Abstract: Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parameter updates, and reusable components opens entry points for attackers. Threats have evolved from data poisoning and weight tampering to agent manipulation and interface exploitation, yet existing reviews lack a unified framework spanning the full fine-tuning lifecycle. Objective: This paper presents a systematic survey of LLM fine-tuning security and establishes a lifecycle-based framework for comparing attacks and defenses, complemented by unified empirical evaluation. Methods: We divide attack and defense mechanisms into three phases by intervention timing: pre-tuning, during-tuning, and post-tuning. Within each phase, strategies are reviewed and contrasted to expose their evolution and limitations. Representative methods are then evaluated under a unified model, hardware, and protocol setup, with cross-phase experiments pairing attacks and defenses from different phases. Results: Attack effectiveness is highly model-dependent and non-monotonic with scale: weight-editing attacks effective on earlier models lose impact on modern open-source LLMs; cross-lingual backdoor transfer, reported as near-perfect at larger scales, fails entirely on tested 1B-4B models; and purely benign samples can compromise safety alignment in instruction-tuned models. Single-phase defenses rarely generalize across phases, and defense effectiveness depends jointly on model architecture and alignment state. Conclusion: We identify key open problems (configuration-robust defense, cross-phase defense composition, and embedding-space attacks beyond behavioral assumptions) and propose concrete future research directions.

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

arXiv:2604.11557v2 Announce Type: replace Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the structural distribution of tool-use trajectories, and relies on incompatible evaluation benchmarks. We present UniToolCall, a unified framework for tool learning that standardizes the entire pipeline from toolset construction and dataset generation to evaluation. The framework curates a large tool pool of 22k+ tools and constructs a hybrid training corpus of 390k+ instances by combining 10 standardized public datasets with structurally controlled synthetic trajectories. It explicitly models diverse interaction patterns, including single-hop vs. multi-hop and single-turn vs. multi-turn, while capturing both serial and parallel execution structures. To support coherent multi-turn reasoning, we further introduce an Anchor Linkage mechanism that enforces cross-turn dependencies. Furthermore, we convert 7 public benchmarks into a unified Query--Action--Observation--Answer (QAOA) representation with fine-grained evaluation at the function-call, turn, and conversation levels. Experiments show that fine-tuning Qwen3-8B on our dataset substantially improves tool-use performance. Under the distractor-heavy Hybrid-20 setting, achieves 93.0% single-turn Strict Precision, outperforming commercial models including GPT, Gemini, and Claude.

UBTF-HSP90A-MIF stress circuit drives lenvatinib resistance and immune exclusion in hepatocellular carcinoma

J Adv Res. 2026 Apr 5:S2090-1232(26)00280-8. doi: 10.1016/j.jare.2026.04.002. Online ahead of print.

ABSTRACT

INTRODUCTION: The clinical benefit of combining lenvatinib with PD-1 blockade in HCC is frequently constrained by adaptive resistance and the development of an immune-cold tumor microenvironment.

OBJECTIVES: This study aimed to elucidate the molecular mechanisms underlying adaptive resistance and immune exclusion during lenvatinib-PD-1 therapy in HCC, with a particular focus on a UBTF/HSP90A/MIF regulatory circuit. We examined whether genetic or pharmacologic targeting of macrophage migration inhibitory factor (MIF) could restore lenvatinib sensitivity, remodel the tumor immune microenvironment, and serve as a predictive biomarker in clinical cohorts.

METHODS: Paired lenvatinib-sensitive and -resistant HCC models were interrogated using integrated multi-omic and functional approaches, including RNA sequencing, promoter pull-down assays, ChIP, luciferase reporter assays, PLA, and flow cytometry. Key findings were validated in patient-derived organoids and xenografts, as well as in an immunocompetent hydrodynamic HCC mouse model. Clinical relevance was evaluated in independent cohorts treated with lenvatinib plus anti-PD-1 therapy.

RESULTS: UBTF directly bound to and transcriptionally activated the HSP90A promoter, resulting in increased HSP90A expression and stabilization of MIF. MIF signaling through CD74 co-activated the PI3K-AKT and MAPK pathways, sustaining tumor cell proliferation under lenvatinib pressure. Single-cell RNA sequencing and multiplex immunohistochemistry revealed macrophage enrichment and CD8+ T-cell exclusion in resistant tumors. Genetic ablation of Mif (Alb-Cre; Mifflox/flox) or pharmacologic inhibition with 4-IPP (4-Iodo-6-phenylpyrimidine) restored lenvatinib sensitivity, reprogrammed the tumor immune microenvironment, and, when combined with PD-1 blockade, achieved superior tumor control and prolonged survival. In clinical datasets, low pretreatment MIF expression was associated with improved responses to lenvatinib plus PD-1 therapy.

CONCLUSIONS: These findings define a UBTF/HSP90A/MIF axis linking proteostasis and cytokine signaling to immune-metabolic dysfunction and lenvatinib resistance in HCC. MIF emerges as both a mechanistic driver and a predictive biomarker, supporting prospective evaluation of therapeutic strategies combining lenvatinib-PD-1 with MIF- or HSP90A-targeted interventions to personalize TKI-ICI therapy.

PMID:41946392 | DOI:10.1016/j.jare.2026.04.002

Suicidal Thoughts and Behaviors Among Chinese Adolescents in Relation to Negative Life Events, Internet Addiction, and Sexual Abuse: Cross-Sectional Study

Background: Increasing suicidal thoughts and behaviors (STB) among adolescents raise social concerns and have a well-recognized association with sexual abuse (SA). However, research regarding the mechanisms explaining the association between SA and STB remains limited. Objective: This study aims to examine the chained mediating effects of negative life events (NLE) and internet addiction (IA) between SA and STB among adolescents in China. Methods: This cross-sectional study used data from the Science Database of the People Mental Health survey conducted between March 2013 and December 2022 by the National Population Health Data Center of the National Research Institute for Family Planning. Through stratified sampling, 20,893 adolescents were recruited from 16 Chinese provinces. After excluding samples with missing relevant variables, 10,664 (55.89%; aged 16-17.9 y; n=5826, 54.63% women) adolescents were included in the final analysis. STB was the outcome variable, with NLE and IA as mediators, all assessed via a questionnaire that was uniformly administered by trained investigators in school settings. The Pearson χ test was used to analyze the association between SA and STB. Using a combination of multiple linear regression and bootstrap testing, the study constructed a chain mediation model to explore how SA influences STB in adolescents through NLE and IA. Results: The scores for SA, NLE, IA, and STB were 1.330 (SD 1.714), 51.960 (SD 23.822), 34.88 (SD 13.852), and 0.690 (SD 1.396), respectively. Multiple linear regression analysis indicated SA was associated with NLE (β=2.382, 95% CI 2.112‐2.653;

Effect of a Digital-Driven Physician-Pharmacist Collaborative Model for Diabetes in Primary Health Care: Cluster Randomized Trial

Background: Evidence-based physician-pharmacist collaborative clinics have demonstrated significant short-term benefits for patients with type 2 diabetes (T2D), but their long-term effectiveness remains unclear, especially in primary health care settings. Objective: This study aimed to explore the long-term effectiveness and cost-effectiveness of a novel, digital-driven, multifaceted physician-pharmacist collaborative model for managing patients with T2D in underresourced settings. Methods: We conducted a 12-month cluster randomized controlled trial from May 2021 to December 2022 across 6 primary health care settings in China. Guided by the theory of planned behavior, the intervention involved routine therapy from physicians along with pharmaceutical interventions from pharmacists. These were delivered through a combination of face-to-face visits and mobile health care. The intervention group received 4 face-to-face visits and biweekly remote education sessions over the 12 months. We conducted intention-to-treat analyses to estimate differences in clinical and behavior indicators between the intervention and control groups. Primary outcomes included glycosylated hemoglobin and 10-year atherosclerotic cardiovascular risk. Data were analyzed using adjusted generalized estimation equations. Results: This study included 574 patients (291 in the intervention group and 283 in the control group). Over 12 months, patients in the intervention group had significant reductions in hemoglobin A1c (–2.57 vs –1.96, respectively; P<.001; 95% CI –1.027 to –0.238) and 10-year atherosclerotic cardiovascular risk (–1.35 vs 0.01, respectively; P<.001; 95% CI –1.690 to –0.630) compared with the control group. Substantial improvements were also observed in several secondary outcomes, including fasting blood glucose, 2-hour postprandial blood glucose, waist circumference, waist-to-hip ratio, blood pressure, triglyceride, and total cholesterol. Total diabetes-related costs decreased, and patient satisfaction improved significantly in the intervention group. There were no significant differences in BMI, high-density lipoprotein, or low-density lipoprotein. Conclusions: These findings suggest that the physician-pharmacist collaborative model could improve the long-term quality and efficiency of T2D management and reduce medical costs in underresourced areas globally. Patients with T2D, especially those with central obesity or high cardiovascular risk, may benefit more from collaborative clinics. Trial Registration: Chinese Clinical Trial Registry ChiCTR2000031839; https://www.chictr.org.cn/showproj.html?proj=51910

Multimodal electron microscopy of halide perovskite interfacial dynamics

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10238-8

A multimodal in situ electron microscopy approach enables direct visualization of structural and chemical evolution in a working halide perovskite light-emitting diode with nanometre precision.

Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis

arXiv:2409.03597v4 Announce Type: replace-cross Abstract: This paper presents the Multimodal Laryngoscopic Video Analyzing System (MLVAS), a novel system that leverages both audio and video data to automatically extract key video segments and metrics from raw laryngeal videostroboscopic videos for assisted clinical assessment. The system integrates video-based glottis detection with an audio keyword spotting method to analyze both video and audio data, identifying patient vocalizations and refining video highlights to ensure optimal inspection of vocal fold movements. Beyond key video segment extraction from the raw laryngeal videos, MLVAS is able to generate effective audio and visual features for Vocal Fold Paralysis (VFP) detection. Pre-trained audio encoders are utilized to encode the patient voice to get the audio features. Visual features are generated by measuring the angle deviation of both the left and right vocal folds to the estimated glottal midline on the segmented glottis masks. To get better masks, we introduce a diffusion-based refinement that follows traditional U-Net segmentation to reduce false positives. We conducted several ablation studies to demonstrate the effectiveness of each module and modalities in the proposed MLVAS. The experimental results on a public segmentation dataset show the effectiveness of our proposed segmentation module. In addition, unilateral VFP classification results on a real-world clinic dataset demonstrate MLVAS's ability of providing reliable and objective metrics as well as visualization for assisted clinical diagnosis.

ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signals

arXiv:2508.14689v4 Announce Type: replace-cross Abstract: Pre-trained foundation models have demonstrated remarkable success in audio, vision and language, yet their potential for general machine signal modeling with arbitrary sampling rates-covering acoustic, vibration, and other industrial sensor data-remains under-explored. In this work, we propose a novel foundation model ECHO that integrates an advanced band-split architecture with frequency positional embeddings, enabling spectral localization across arbitrary sampling configurations. Moreover, the model incorporates sliding patches to support inputs of variable length without padding or cropping, producing a concise embedding that retains both temporal and spectral fidelity and naturally extends to streaming scenarios. We evaluate our method on various kinds of machine signal datasets, including previous DCASE task 2 challenges (2020-2025), and widely-used industrial signal corpora. Experimental results demonstrate consistent state-of-the-art performance in machine signal anomaly detection and fault classification, confirming the effectiveness and generalization capability of the proposed model. We open-sourced ECHO on https://github.com/yucongzh/ECHO.

High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning

arXiv:2506.04051v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) currently respond to every prompt. However, they can produce incorrect answers when they lack knowledge or capability -- a problem known as hallucination. We instead propose post-training an LLM to generate content only when confident in its correctness and to otherwise (partially) abstain. Specifically, our method, HALT, produces capability-aligned post-training data that encodes what the model can and cannot reliably generate. We generate this data by splitting responses of the pretrained LLM into factual fragments (atomic statements or reasoning steps), and use ground truth information to identify incorrect fragments. We achieve capability-aligned finetuning responses by either removing incorrect fragments or replacing them with "Unsure from Here" -- according to a tunable threshold that allows practitioners to trade off response completeness and mean correctness of the response's fragments. We finetune four open-source models for biography writing, mathematics, coding, and medicine with HALT for three different trade-off thresholds. HALT effectively trades off response completeness for correctness, increasing the mean correctness of response fragments by 15% on average, while resulting in a 4% improvement in the F1 score (mean of completeness and correctness of the response) compared to the relevant baselines. By tuning HALT for highest correctness, we train a single reliable Llama3-70B model with correctness increased from 51% to 87% across all four domains while maintaining 53% of the response completeness achieved with standard finetuning.
❌