❌

Normal view

Well-Being and Cognitive Factors Influencing Health Care Workers’ Adherence to Internet-Based Stress Management: Mixed Methods Analysis of a Nonrandomized Controlled Study

Background: High stress levels are common among health care workers (HCWs), threatening their health and workforce stability. Internet-based mobile stress management (MSM) is a promising intervention for reducing work-related stress; however, poor adherence limits effectiveness. Exploring factors influencing HCWs’ adherence may thus aid in developing optimal interventions. Objective: The research aimed to investigate (1) how HCWs’ well-being and cognitive factors influenced MSM treatment adherence and (2) what HCWs’ specific needs for MSM were. Methods: This study was a convergent mixed methods secondary analysis of a nonrandomized controlled trial. HCWs who were currently employed, had internet access, had no serious medical problems, and were willing to participate were recruited by convenience sampling through an MSM project in a large Chinese general hospital from August 11, 2021, to January 31, 2022. Those intending to leave the hospital or with insufficient medical condition for follow-up were excluded. Quantitative data were collected from 157 HCWs (n=135, 86% female participants; mean age of 33.7, SD 4.9 y) electronically via Research Electronic Data Capture (REDCap). Measures included sociodemographic characteristics, the Fatigue Assessment Scale, the 14-item Perceived Stress Scale, a user experience questionnaire, an attitudes scale (perceived usefulness, feasibility, and enjoyment), and self-reported practice frequency. Qualitative data were collected via an open-ended question answered by 96 participants. Quantitative data were analyzed using hierarchical regression and structural equation modeling. Qualitative data were analyzed using reflexive-thematic analysis in NVivo (QSR International). Results: In the quantitative study (n=157), hierarchical regression analyses showed that fatigue was a significant negative predictor of adherence (b=−0.050, 95% CI −0.086 to −0.015; =−2.859; 2-tailed =.005), while user experience (b=0.074, 95% CI 0.042-0.106; =4.569; 2-tailed

Unified modeling of 3D molecular generation via atomic interactions with PocketXMol

A versatile, atom-level generative AI model enables unified pocket-interacting tasks, from docking to de novo design, and demonstrates robust experimental validation for both small-molecule and peptide therapeutics.

SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model

arXiv:2603.26738v2 Announce Type: replace-cross Abstract: While automated sleep staging has achieved expert-level accuracy, its clinical adoption is hindered by a lack of auditable reasoning. We introduce SleepVLM, a rule-grounded vision-language model (VLM) designed to stage sleep from multi-channel polysomnography (PSG) waveform images while generating clinician-readable rationales based on American Academy of Sleep Medicine (AASM) scoring criteria. Utilizing waveform-perceptual pre-training and rule-grounded supervised fine-tuning, SleepVLM achieved Cohen's kappa scores of 0.767 on an held out test set (MASS-SS1) and 0.743 on an external cohort (ZUAMHCS), matching state-of-the-art performance. Expert evaluations further validated the quality of the model's reasoning, with mean scores exceeding 4.0/5.0 for factual accuracy, evidence comprehensiveness, and logical coherence. By coupling competitive performance with transparent, rule-based explanations, SleepVLM may improve the trustworthiness and auditability of automated sleep staging in clinical workflows. To facilitate further research in interpretable sleep medicine, we release MASS-EX, a novel expert-annotated dataset.

R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning

arXiv:2505.21668v3 Announce Type: replace Abstract: Practical guidance on training Large Language Models (LLMs) to leverage Code Interpreter across diverse tasks remains lacking. We present R1-Code-Interpreter, an extension of a text-only LLM trained via multi-turn supervised fine-tuning (SFT) and reinforcement learning (RL) to autonomously generate multiple code queries during step-by-step reasoning. Unlike prior RL + tool-use efforts focused on narrow domains such as math or retrieval, we curate 144 diverse reasoning and planning tasks and show that training a general-purpose Code Interpreter across them presents significant challenges due to task heterogeneity and scarcity of effective samples. To address this, we introduce a multi-stage curriculum learning approach that partitions training samples by measured improvement potential. The RL training prioritizes samples with higher potential and gradually shifts to lower-potential ones, increasing the average RL gains from merely +3.4% to +9.3% across Qwen-2.5 models (3/7/14B). Our final model, R1-CI-14B, improves average accuracy on the 37 test tasks from 44.1% to 72.4%, outperforming text-only GPT-4o (58.6%) and GPT-4o with Code Interpreter (70.9%). Notably, R1-CI-14B also exhibits emergent self-checking behavior through code generation. Datasets, Codes, and Models are available at https://github.com/yongchao98/R1-Code-Interpreter and https://huggingface.co/yongchao98.

FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations

arXiv:2602.18437v1 Announce Type: cross Abstract: Generating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-optimize citation fidelity while overlooking relevance to the user query, which degrades answer quality and robustness in real-world settings with noisy or irrelevant retrieved content. Moreover, the prevailing single-pass paradigm struggles to deliver optimal answers in long-form generation that requiring multiple citations. To address these limitations, we propose FineRef, a framework based on Fine-grained error Reflection, which explicitly teaches the model to self-identify and correct two key citation errors, mismatch and irrelevance, on a per-citation basis. FineRef follows a two-stage training strategy. The first stage instills an "attempt-reflect-correct" behavioral pattern via supervised fine-tuning, using fine-grained and controllable reflection data constructed by specialized lightweight models. An online self-reflective bootstrapping strategy is designed to improve generalization by iteratively enriching training data with verified, self-improving examples. To further enhance the self-reflection and correction capability, the second stage applies process-level reinforcement learning with a multi-dimensional reward scheme that promotes reflection accuracy, answer quality, and correction gain. Experiments on the ALCE benchmark demonstrate that FineRef significantly improves both citation performance and answer accuracy. Our 7B model outperforms GPT-4 by up to 18% in Citation F1 and 4% in EM Recall, while also surpassing the state-of-the-art model across key evaluation metrics. FineRef also exhibits strong generalization and robustness in domain transfer settings and noisy retrieval scenarios.
❌