Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Gene
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
arXiv:2512.23184v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choic
From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
arXiv:2512.23545v1 Announce Type: cross Abstract: Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to produce predictions, without reassessment or targeted evidence acquisition under ambiguous diagnoses. This contrasts with clinical diagnostic workflows that refine hypotheses through repeated slide observations
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
-
Omics in Hepatocellular
-
Immunotherapy for virus-related hepatocellular carcinoma: recent progress and future directions
Ann Med. 2026 Dec;58(1):2607229. doi: 10.1080/07853890.2025.2607229. Epub 2025 Dec 26.ABSTRACTBACKGROUND: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with hepatitis B virus (HBV) and hepatitis C virus (HCV) infections remaining the predominant etiological factors. Chronic viral infection not only drives carcinogenesis but also reshapes the hepatic immune microenvironment, profoundly influencing the efficacy and safety of immunotherapy.RECENT ADVANCES:
Immunotherapy for virus-related hepatocellular carcinoma: recent progress and future directions
Ann Med. 2026 Dec;58(1):2607229. doi: 10.1080/07853890.2025.2607229. Epub 2025 Dec 26.
ABSTRACT
BACKGROUND: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with hepatitis B virus (HBV) and hepatitis C virus (HCV) infections remaining the predominant etiological factors. Chronic viral infection not only drives carcinogenesis but also reshapes the hepatic immune microenvironment, profoundly influencing the efficacy and safety of immunotherapy.
RECENT ADVANCES: Immune checkpoint inhibitors (ICIs) have revolutionized systemic therapy for advanced HCC, with agents targeting PD-1/PD-L1 demonstrating clinical benefit. Combination strategies - such as ICIs with anti-angiogenic therapies, multikinase inhibitors, or locoregional treatments - have shown synergistic efficacy and are now standard of care in certain settings. For virus-related HCC, antiviral therapy improves immune responsiveness and reduces risks such as HBV reactivation, underscoring the need for integrated management.
FUTURE PERSPECTIVES: Emerging therapeutic approaches include next-generation immune checkpoints (e.g. TIM-3, LAG-3, TIGIT), bispecific antibodies, cellular therapies (CAR-T, TCR-T, TILs), and tumor vaccines targeting viral or tumor-associated antigens. Advances in biomarker discovery, including circulating tumor DNA, immune signatures, and microbiome modulation, are expected to guide personalized treatment. Integration of multi-omics and clinical data will further refine patient selection and optimize treatment sequencing.
CONCLUSION: Immunotherapy offers new hope for patients with virus-related HCC, but challenges remain in response heterogeneity, resistance, and toxicity. Individualized strategies that combine immunotherapy with effective antiviral management and biomarker-|guided patient selection are essential. Continued translational and clinical research into virus-immune-tumor interactions will enable safer, more effective, and more durable treatment outcomes, ultimately transforming HCC into a more manageable disease.
PMID:41454610 | PMC:PMC12777805 | DOI:10.1080/07853890.2025.2607229
-
Journal of Medical Internet Research
-
Developing and Evaluating Guidelines to Prevent Overdependence on Digital Therapeutics in Children and Adolescents: Randomized Controlled Trial
Background: Digital therapeutics (DTx) for children and adolescents with mental health problems have been developed in the health care industry. Despite reports of side effects from DTx for children and adolescents, there have been no guidelines to address the prevention of DTx overdependence among young users. Objective: This study aimed to identify the requirements for guidelines to prevent DTx overdependence in children and adolescents and to develop and evaluate these guidelines. Methods: We
Developing and Evaluating Guidelines to Prevent Overdependence on Digital Therapeutics in Children and Adolescents: Randomized Controlled Trial
-
npj Digital Medicine
-
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
npj Digital Medicine, Published online: 26 December 2025; doi:10.1038/s41746-025-02277-8A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
npj Digital Medicine, Published online: 26 December 2025; doi:10.1038/s41746-025-02277-8
A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains-
cs.AI, q-bio.NC updates on arXiv.org
-
Distributional AGI Safety
arXiv:2512.16856v1 Announce Type: new Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI emergence hypothesis, where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills and affordances, has received far less attention.
Distributional AGI Safety
-
cs.AI, q-bio.NC updates on arXiv.org
-
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
arXiv:2512.15769v1 Announce Type: cross Abstract: The increasing use of generative models such as diffusion models for synthetic data augmentation has greatly reduced the cost of data collection and labeling in downstream perception tasks. However, this new data source paradigm may introduce important security concerns. This work investigates backdoor propagation in such emerging generative data supply chains, namely Data-Chain Backdoor (DCB). Specifically, we find that open-source diffusion mo
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
arXiv:2512.16455v1 Announce Type: cross Abstract: In this paper, we describe a federated compute platform dedicated to support Artificial Intelligence in scientific workloads. Putting the effort into reproducible deployments, it delivers consistent, transparent access to a federation of physically distributed e-Infrastructures. Through a comprehensive service catalogue, the platform is able to offer an integrated user experience covering the full Machine Learning lifecycle, including model deve
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
arXiv:2512.09566v2 Announce Type: replace Abstract: Drug discovery is a time-consuming and expensive process, with traditional high-throughput and docking-based virtual screening hampered by low success rates and limited scalability. Recent advances in generative modelling, including autoregressive, diffusion, and flow-based approaches, have enabled de novo ligand design beyond the limits of enumerative screening. Yet these models often suffer from inadequate generalization, limited interpretab
Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
-
cs.AI, q-bio.NC updates on arXiv.org
-
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
arXiv:2511.07392v3 Announce Type: replace-cross Abstract: In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, we propose a Voice-Interactive Surgical Agent (VISA) built on a hierarchical multi-agent framework consisting of an orchestration agent and three task-specific agents driven by Large Language Models (LLMs). These LLM-base
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
-
cs.AI, q-bio.NC updates on arXiv.org
-
First, do NOHARM: towards clinically safe large language models
arXiv:2512.01241v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark using 100 real primary care-to-specialist consultation cases to measure frequency and severity of harm from LLM-generated medical recommendations. NOHARM covers 10 specialties, with 12,747 expert
First, do NOHARM: towards clinically safe large language models
-
cs.AI, q-bio.NC updates on arXiv.org
-
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
arXiv:2512.13716v1 Announce Type: new Abstract: Personalized decision-making is essential for human-AI interaction, enabling AI agents to act in alignment with individual users' value preferences. As AI systems expand into real-world applications, adapting to personalized values beyond task completion or collective alignment has become a critical challenge. We address this by proposing a value-driven approach to personalized decision-making. Human values serve as stable, transferable signals th
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
-
cs.AI, q-bio.NC updates on arXiv.org
-
A data-physics hybrid generative model for patient-specific post-stroke motor rehabilitation using wearable sensor data
arXiv:2512.14329v1 Announce Type: cross Abstract: Dynamic prediction of locomotor capacity after stroke is crucial for tailoring rehabilitation, yet current assessments provide only static impairment scores and do not indicate whether patients can safely perform specific tasks such as slope walking or stair climbing. Here, we develop a data-physics hybrid generative framework that reconstructs an individual stroke survivor's neuromuscular control from a single 20 m level-ground walking trial an
A data-physics hybrid generative model for patient-specific post-stroke motor rehabilitation using wearable sensor data
-
cs.AI, q-bio.NC updates on arXiv.org
-
COMMA: A Communicative Multimodal Multi-Agent Benchmark
arXiv:2410.07553v5 Announce Type: replace Abstract: The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative tasks. This oversight presents a critical gap in understanding their effectiveness in real-world deployments, particularly when communicating with humans. Existing agentic benchmarks fail to address key aspects of inter-agent communication and collaboration, particular
COMMA: A Communicative Multimodal Multi-Agent Benchmark
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
arXiv:2508.15811v2 Announce Type: replace-cross Abstract: Generative query suggestion using large language models offers a powerful way to enhance conversational systems, but aligning outputs with nuanced user preferences remains a critical challenge. To address this, we introduce a multi-stage framework designed for progressive alignment between the generation policy and user intent. Our pipeline begins with prompt engineering as a cold-start strategy, followed by the Supervised Fine-Tuning st
From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
-
cs.AI, q-bio.NC updates on arXiv.org
-
Grounding Large Language Models in Clinical Evidence: A Retrieval-Augmented Generation System for Querying UK NICE Clinical Guidelines
arXiv:2510.02967v3 Announce Type: replace-cross Abstract: This paper presents the development and evaluation of a Retrieval-Augmented Generation (RAG) system for querying the United Kingdom's National Institute for Health and Care Excellence (NICE) clinical guidelines using Large Language Models (LLMs). The extensive length and volume of these guidelines can impede their utilisation within a time-constrained healthcare system, a challenge this project addresses through the creation of a system
Grounding Large Language Models in Clinical Evidence: A Retrieval-Augmented Generation System for Querying UK NICE Clinical Guidelines
-
npj Digital Medicine
-
H&E-based MSI/MMR testing with AI in colorectal cancer: a multi-centred blinded evaluation
npj Digital Medicine, Published online: 15 December 2025; doi:10.1038/s41746-025-02218-5H&E-based MSI/MMR testing with AI in colorectal cancer: a multi-centred blinded evaluation
H&E-based MSI/MMR testing with AI in colorectal cancer: a multi-centred blinded evaluation
npj Digital Medicine, Published online: 15 December 2025; doi:10.1038/s41746-025-02218-5
H&E-based MSI/MMR testing with AI in colorectal cancer: a multi-centred blinded evaluation-
cs.AI, q-bio.NC updates on arXiv.org
-
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
arXiv:2512.11661v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly embedded in academic writing practices. Although numerous studies have explored how researchers employ these tools for scientific writing, their concrete implementation, limitations, and design challenges within the literature review process remain underexplored. In this paper, we report a user study with researchers across multiple disciplines to characterize current practices, benefits, and \textit
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
-
Journal of Medical Internet Research
-
Stakeholder Criteria for Trust in Artificial Intelligence–Based Computer Perception Tools in Health Care: Qualitative Interview Study
Background: Computer perception (CP) technologies hold significant promise for advancing precision mental health care systems, given their ability to leverage algorithmic analysis of continuous, passive sensing data from wearables and smartphones (eg, behavioral activity, geolocation, vocal features, and ambient environmental data) to infer clinically meaningful behavioral and physiological states. However, successful implementation critically depends on cultivating well-founded stakeholder trus