Normal view
-
Journal of Medical Internet Research
-
Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study
Background: Gastroesophageal reflux disease (GERD) symptoms significantly affect patients’ quality of life. Patient-reported outcome (PRO) instruments for symptoms measurement in GERD patients is advocated by regulatory authority. Current tools for GERD symptoms evaluation are limited and the results can be biased by the recall bias. To better characterize the GERD symptoms, an e-diary was developed for daily GERD symptom monitoring. Objective: To build up and optimize a PRO-based e-diary, and t
-
cs.AI, q-bio.NC updates on arXiv.org
-
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
-
cs.AI, q-bio.NC updates on arXiv.org
-
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
arXiv:2601.01627v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety o
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
arXiv:2506.05623v2 Announce Type: replace-cross Abstract: Infrastructure-as-Code (IaC) generation holds significant promise for automating cloud infrastructure provisioning. Recent advances in Large Language Models (LLMs) present a promising opportunity to democratize IaC development by generating deployable infrastructure templates from natural language descriptions. However, current evaluation focuses on syntactic correctness while ignoring deployability, the critical measure of the utility o
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Wearable-informed generative digital avatars predict task-conditioned post-stroke locomotion
arXiv:2512.14329v2 Announce Type: replace-cross Abstract: Dynamic prediction of locomotor capacity after stroke could enable more individualized rehabilitation, yet current assessments largely provide static impairment scores and do not indicate whether patients can perform specific tasks such as slope walking or stair climbing. Here, we present a wearable-informed data-physics hybrid generative framework that reconstructs a stroke survivor's locomotor control from wearable inertial sensing and
Wearable-informed generative digital avatars predict task-conditioned post-stroke locomotion
-
Nature Medicine
-
A minimally invasive dried blood spot biomarker test for the detection of Alzheimer’s disease pathology
Nature Medicine, Published online: 05 January 2026; doi:10.1038/s41591-025-04080-0This multicenter study demonstrates use of dried and capillary blood as a minimally invasive, scalable approach for Alzheimer’s biomarker testing in research, with potential as a widely scalable population-based research approach, especially in resource-limited settings.
A minimally invasive dried blood spot biomarker test for the detection of Alzheimer’s disease pathology
Nature Medicine, Published online: 05 January 2026; doi:10.1038/s41591-025-04080-0
This multicenter study demonstrates use of dried and capillary blood as a minimally invasive, scalable approach for Alzheimer’s biomarker testing in research, with potential as a widely scalable population-based research approach, especially in resource-limited settings.-
npj Digital Medicine
-
A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer
npj Digital Medicine, Published online: 31 December 2025; doi:10.1038/s41746-025-02260-3A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer
A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer
npj Digital Medicine, Published online: 31 December 2025; doi:10.1038/s41746-025-02260-3
A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer-
cs.AI, q-bio.NC updates on arXiv.org
-
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Gene
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
arXiv:2512.23184v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choic
From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
arXiv:2512.23545v1 Announce Type: cross Abstract: Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to produce predictions, without reassessment or targeted evidence acquisition under ambiguous diagnoses. This contrasts with clinical diagnostic workflows that refine hypotheses through repeated slide observations
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
-
Omics in Hepatocellular
-
Immunotherapy for virus-related hepatocellular carcinoma: recent progress and future directions
Ann Med. 2026 Dec;58(1):2607229. doi: 10.1080/07853890.2025.2607229. Epub 2025 Dec 26.ABSTRACTBACKGROUND: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with hepatitis B virus (HBV) and hepatitis C virus (HCV) infections remaining the predominant etiological factors. Chronic viral infection not only drives carcinogenesis but also reshapes the hepatic immune microenvironment, profoundly influencing the efficacy and safety of immunotherapy.RECENT ADVANCES:
Immunotherapy for virus-related hepatocellular carcinoma: recent progress and future directions
Ann Med. 2026 Dec;58(1):2607229. doi: 10.1080/07853890.2025.2607229. Epub 2025 Dec 26.
ABSTRACT
BACKGROUND: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with hepatitis B virus (HBV) and hepatitis C virus (HCV) infections remaining the predominant etiological factors. Chronic viral infection not only drives carcinogenesis but also reshapes the hepatic immune microenvironment, profoundly influencing the efficacy and safety of immunotherapy.
RECENT ADVANCES: Immune checkpoint inhibitors (ICIs) have revolutionized systemic therapy for advanced HCC, with agents targeting PD-1/PD-L1 demonstrating clinical benefit. Combination strategies - such as ICIs with anti-angiogenic therapies, multikinase inhibitors, or locoregional treatments - have shown synergistic efficacy and are now standard of care in certain settings. For virus-related HCC, antiviral therapy improves immune responsiveness and reduces risks such as HBV reactivation, underscoring the need for integrated management.
FUTURE PERSPECTIVES: Emerging therapeutic approaches include next-generation immune checkpoints (e.g. TIM-3, LAG-3, TIGIT), bispecific antibodies, cellular therapies (CAR-T, TCR-T, TILs), and tumor vaccines targeting viral or tumor-associated antigens. Advances in biomarker discovery, including circulating tumor DNA, immune signatures, and microbiome modulation, are expected to guide personalized treatment. Integration of multi-omics and clinical data will further refine patient selection and optimize treatment sequencing.
CONCLUSION: Immunotherapy offers new hope for patients with virus-related HCC, but challenges remain in response heterogeneity, resistance, and toxicity. Individualized strategies that combine immunotherapy with effective antiviral management and biomarker-|guided patient selection are essential. Continued translational and clinical research into virus-immune-tumor interactions will enable safer, more effective, and more durable treatment outcomes, ultimately transforming HCC into a more manageable disease.
PMID:41454610 | PMC:PMC12777805 | DOI:10.1080/07853890.2025.2607229
-
Journal of Medical Internet Research
-
Developing and Evaluating Guidelines to Prevent Overdependence on Digital Therapeutics in Children and Adolescents: Randomized Controlled Trial
Background: Digital therapeutics (DTx) for children and adolescents with mental health problems have been developed in the health care industry. Despite reports of side effects from DTx for children and adolescents, there have been no guidelines to address the prevention of DTx overdependence among young users. Objective: This study aimed to identify the requirements for guidelines to prevent DTx overdependence in children and adolescents and to develop and evaluate these guidelines. Methods: We
Developing and Evaluating Guidelines to Prevent Overdependence on Digital Therapeutics in Children and Adolescents: Randomized Controlled Trial
-
npj Digital Medicine
-
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
npj Digital Medicine, Published online: 26 December 2025; doi:10.1038/s41746-025-02277-8A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
npj Digital Medicine, Published online: 26 December 2025; doi:10.1038/s41746-025-02277-8
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains-
cs.AI, q-bio.NC updates on arXiv.org
-
Distributional AGI Safety
arXiv:2512.16856v1 Announce Type: new Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI emergence hypothesis, where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills and affordances, has received far less attention.
Distributional AGI Safety
-
cs.AI, q-bio.NC updates on arXiv.org
-
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
arXiv:2512.15769v1 Announce Type: cross Abstract: The increasing use of generative models such as diffusion models for synthetic data augmentation has greatly reduced the cost of data collection and labeling in downstream perception tasks. However, this new data source paradigm may introduce important security concerns. This work investigates backdoor propagation in such emerging generative data supply chains, namely Data-Chain Backdoor (DCB). Specifically, we find that open-source diffusion mo
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
arXiv:2512.16455v1 Announce Type: cross Abstract: In this paper, we describe a federated compute platform dedicated to support Artificial Intelligence in scientific workloads. Putting the effort into reproducible deployments, it delivers consistent, transparent access to a federation of physically distributed e-Infrastructures. Through a comprehensive service catalogue, the platform is able to offer an integrated user experience covering the full Machine Learning lifecycle, including model deve
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
arXiv:2512.09566v2 Announce Type: replace Abstract: Drug discovery is a time-consuming and expensive process, with traditional high-throughput and docking-based virtual screening hampered by low success rates and limited scalability. Recent advances in generative modelling, including autoregressive, diffusion, and flow-based approaches, have enabled de novo ligand design beyond the limits of enumerative screening. Yet these models often suffer from inadequate generalization, limited interpretab
Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
-
cs.AI, q-bio.NC updates on arXiv.org
-
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
arXiv:2511.07392v3 Announce Type: replace-cross Abstract: In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, we propose a Voice-Interactive Surgical Agent (VISA) built on a hierarchical multi-agent framework consisting of an orchestration agent and three task-specific agents driven by Large Language Models (LLMs). These LLM-base
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
-
cs.AI, q-bio.NC updates on arXiv.org
-
First, do NOHARM: towards clinically safe large language models
arXiv:2512.01241v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark using 100 real primary care-to-specialist consultation cases to measure frequency and severity of harm from LLM-generated medical recommendations. NOHARM covers 10 specialties, with 12,747 expert
First, do NOHARM: towards clinically safe large language models
-
cs.AI, q-bio.NC updates on arXiv.org
-
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
arXiv:2512.13716v1 Announce Type: new Abstract: Personalized decision-making is essential for human-AI interaction, enabling AI agents to act in alignment with individual users' value preferences. As AI systems expand into real-world applications, adapting to personalized values beyond task completion or collective alignment has become a critical challenge. We address this by proposing a value-driven approach to personalized decision-making. Human values serve as stable, transferable signals th