❌

Normal view

Surface-based Molecular Design with Multi-modal Flow Matching

arXiv:2601.04506v1 Announce Type: cross Abstract: Therapeutic peptides show promise in targeting previously undruggable binding sites, with recent advancements in deep generative models enabling full-atom peptide co-design for specific protein receptors. However, the critical role of molecular surfaces in protein-protein interactions (PPIs) has been underexplored. To bridge this gap, we propose an omni-design peptides generation paradigm, called SurfFlow, a novel surface-based generative algorithm that enables comprehensive co-design of sequence, structure, and surface for peptides. SurfFlow employs a multi-modality conditional flow matching (CFM) architecture to learn distributions of surface geometries and biochemical properties, enhancing peptide binding accuracy. Evaluated on the comprehensive PepMerge benchmark, SurfFlow consistently outperforms full-atom baselines across all metrics. These results highlight the advantages of considering molecular surfaces in de novo peptide discovery and demonstrate the potential of integrating multiple protein modalities for more effective therapeutic peptide discovery.

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

arXiv:2601.04790v1 Announce Type: cross Abstract: Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexplored. We present the first systematic analysis of role-based authority bias in free-form multi-agent evaluation using ChatEval. Applying French and Raven's power-based theory, we classify authoritative roles into legitimate, referent, and expert types and analyze their influence across 12-turn conversations. Experiments with GPT-4o and DeepSeek R1 reveal that Expert and Referent power roles exert stronger influence than Legitimate power roles. Crucially, authority bias emerges not through active conformity by general agents, but through authoritative roles consistently maintaining their positions while general agents demonstrate flexibility. Furthermore, authority influence requires clear position statements, as neutral responses fail to generate bias. These findings provide key insights for designing multi-agent frameworks with asymmetric interaction patterns.

Atlas 2 -- Foundation models for clinical deployment

arXiv:2601.05148v1 Announce Type: cross Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology -- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this report, we present Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models which bridge these shortcomings by showing state-of-the-art performance in prediction performance, robustness, and resource efficiency in a comprehensive evaluation across eighty public benchmarks. Our models were trained on the largest pathology foundation model dataset to date comprising 5.5 million histopathology whole slide images, collected from three medical institutions Charit\'e - Universt\"atsmedizin Berlin, LMU Munich, and Mayo Clinic.

PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor

arXiv:2601.01802v3 Announce Type: replace Abstract: To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key challenges: \textbf{1) Can we train a highly realistic AI counselor?} Realistic counseling is a longitudinal task requiring sustained memory and dynamic goal tracking. We propose a multi-session benchmark (spanning 6-10 sessions across three distinct stages) that demands critical capabilities such as memory continuity, adaptive reasoning, and longitudinal planning. The dataset is annotated with extensive professional skills, comprising over 677 meta-skills and 4577 atomic skills. \textbf{2) How to train a multi-therapy AI counselor?} While existing models often focus on a single therapy, complex cases frequently require flexible strategies among various therapies. We construct a diverse dataset covering five therapeutic modalities (Psychodynamic, Behaviorism, CBT, Humanistic Existentialist, and Postmodernist) alongside an integrative therapy with a unified three-stage clinical framework across six core psychological topics. \textbf{3) How to systematically evaluate an AI counselor?} We establish a holistic evaluation framework with 18 therapy-specific and therapy-shared metrics across Client-Level and Counselor-Level dimensions. To support this, we also construct over 2,000 diverse client profiles. Extensive experimental analysis fully validates the superior quality and clinical fidelity of our dataset. Crucially, \texttt{PsychEval} transcends static benchmarking to serve as a high-fidelity reinforcement learning environment that enables the self-evolutionary training of clinically responsible and adaptive AI counselors.

A Web-Based Cancer Prevention Intervention for Rural Emerging Adults: Mixed Methods Development and Pilot-Testing Study

Background: The rapid growth of user-generated web-based health information increases the complexity of cancer information seeking. One promising strategy for promoting high-quality cancer information consumption is through targeted interventions that are intentionally designed to reach individuals in the web-based spaces they occupy. However, there is a paucity of evidence-based information on the best strategies for designing and implementing web-based health behavior change interventions to improve individuals’ cancer-related knowledge and prevent cancer. Objective: This study aimed to develop and pilot test a theory-based intervention via the web to reduce 6 cancer risk factors among rural emerging adults (EAs) through community-engaged research. Methods: This mixed methods evaluation describes the development of a web-based cancer prevention intervention aimed at rural EAs aged 18-26 years in the United States and delivered in Facebook private groups. The intervention was guided by behavior change theory and cocreated with EA and Stakeholder Organization Advisory Boards to ensure relevance, accessibility, and appropriateness. We report on 3 formative surveys, a pilot intervention, protocol development, and the community-engaged process for intervention development. Descriptive statistics were applied to the surveys and pilot intervention baseline results to produce means and SDs using R. Results: We developed posts (n=400) for a Facebook feed aimed at reducing 6 cancer risk behaviors (unhealthy diet, lack of physical activity, tobacco use, alcohol use, sun exposure, and human papillomavirus infection) with iterative input from the EA and stakeholder advisory boards. Formative surveys with rural EAs (n=297) and a pilot study of the intervention with this population (n=26) were conducted. In the pilot study, the intervention reached participants across rural counties, with sustained engagement (post views=1060, reactions=346, comments=72) over a one-month period. Key modifications to the intervention content and design emerged from both advisory boards, the formative surveys, and the pilot intervention, focusing on using perceived reliable sources and direct links to source material. Conclusions: This web-based cancer prevention intervention is scalable and delivers engaging, evidence-informed health information to rural EAs. We offer key insights into the design and implementation of web-based cancer prevention interventions for EAs by describing the resources, timelines, and expertise needed to design and implement the intervention. Considerations for fully engaging EA and community stakeholder partners are presented, and we discuss how their involvement resulted in modifications that strengthened the intervention. Finally, we highlight the importance of theory-based health-behavior messaging, digital messaging skillsets, and platform-tailored dissemination strategies for maximizing web-based intervention acceptability. Trial Registration: ClinicalTrials.gov NCT05618158; https://classic.clinicaltrials.gov/ct2/show/NCT05618158

Adaptive therapy for perioperative non-small cell lung cancer: strategies guided by dynamic minimal residual disease adjustment

Transl Oncol. 2026 Jan 6;64:102660. doi: 10.1016/j.tranon.2025.102660. Online ahead of print.

ABSTRACT

Lung cancer remains the leading cause of cancer incidence and mortality worldwide, with non-small cell lung cancer (NSCLC) accounting for about 85% of cases. The low rate of early diagnosis and the high rate of occult metastases limit the survival benefits of conventional treatments. The current TNM staging system fails to fully reflect tumor heterogeneity or the dynamic molecular evolution of the disease, thus affecting the prediction of recurrence and the prognostic stratification. Some recent advances in minimal residual disease (MRD) detection, such as ultra-sensitive liquid biopsy technologies, have largely overcome the limitations of traditional imaging and offered a transformative approach for continuous, precision-based management of lung cancer. This review systematically summarized the technological evolution of MRD detection and highlighted its clinical significance in guiding adaptive therapy for NSCLC, including treatment escalation, de-escalation, and the emerging concept of precision-guided drug holidays. Moreover, the authors comprehensively discussed the "Four-Dimensional TNMB Staging System," which incorporates continuous molecular monitoring to address the static limitations of conventional staging and enhance the accuracy of prognostic stratification. Although ongoing challenges, such as the lack of standardized interpretation criteria and limited detection sensitivity, the combinations with the third-generation liquid biopsy platforms, multi-omics analyses, and multi-center prospective validation studies are expected to advance the clinical implementation of MRD-guided strategies. The paradigm change will enable the transition of NSCLC management from conventional standardized models to a precision-guided, closed-loop system of "monitoring-intervention-remonitoring," establishing a solid theoretical and practical foundation for comprehensive, molecularly driven management strategies.

PMID:41496417 | DOI:10.1016/j.tranon.2025.102660

Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study

Background: Gastroesophageal reflux disease (GERD) symptoms significantly affect patients’ quality of life. Patient-reported outcome (PRO) instruments for symptoms measurement in GERD patients is advocated by regulatory authority. Current tools for GERD symptoms evaluation are limited and the results can be biased by the recall bias. To better characterize the GERD symptoms, an e-diary was developed for daily GERD symptom monitoring. Objective: To build up and optimize a PRO-based e-diary, and to investigate the effect of symptom frequency on adherence. Methods: The GERD e-diary evaluated 8 daytime (acid regurgitation, cough, heartburn, sour taste in the mouth, hiccups, hoarseness, dysphagia, and chest pain) and 2 nighttime symptoms (acid regurgitation and cough) for consecutive 8 weeks. The adherence of e-diary, defined as daily completing rate of e-diary, was evaluated and optimized from First Stage to Third Stage with no reminder implemented in First Stage, sending reminding SMS (Short Message Service) text messaging upon detecting missing data in Second Stage (no reminder during the first 3 to 5 days after enrollment), and immediate installation of reminding system at enrollment in Third Stage. GERD symptom frequency was obtained by summation of the symptomatic days in each week. A multiple regression analysis was performed to examine the effects of system optimization and GERD symptom frequency on patient adherence, while controlling for potential confounding variables. Results: 138 GERD patients (M/F=70/68; age: mean 52.9, SD 12.3 years) were recruited. At First Stage, the adherence was 47.2%, 40% and 57.6% for nighttime, daytime and overall symptom. System optimization significantly improved adherence with increased adherence of nighttime symptoms by 12.5% (P=.005) and 10.9% (P=.01), daytime symptom by 21.7% (P

OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment

arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via semantic search engine; (3) constructing a hierarchical taxonomy of core-task-related work and performing contribution-level full-text comparisons against each contribution; and (4) synthesizing all analyses into a structured novelty report with explicit citations and evidence snippets. Unlike naive LLM-based approaches, \textsc{OpenNovelty} grounds all assessments in retrieved real papers, ensuring verifiable judgments. We deploy our system on 500+ ICLR 2026 submissions with all reports publicly available on our website, and preliminary analysis suggests it can identify relevant prior work, including closely related papers that authors may overlook. OpenNovelty aims to empower the research community with a scalable tool that promotes fair, consistent, and evidence-backed peer review.

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

arXiv:2601.01627v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety of LLMs for Japanese healthcare. Our benchmark is based on 67 guidelines from the Japan Medical Association and contains over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Using a dual-LLM scoring protocol, we evaluate 27 models and find that commercial models maintain robust safety while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.0, $p

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

arXiv:2506.05623v2 Announce Type: replace-cross Abstract: Infrastructure-as-Code (IaC) generation holds significant promise for automating cloud infrastructure provisioning. Recent advances in Large Language Models (LLMs) present a promising opportunity to democratize IaC development by generating deployable infrastructure templates from natural language descriptions. However, current evaluation focuses on syntactic correctness while ignoring deployability, the critical measure of the utility of IaC configuration files. Six state-of-the-art LLMs performed poorly on deployability, achieving only 20.8$\sim$30.2% deployment success rate on the first attempt. In this paper, we construct DPIaC-Eval, the first deployability-centric IaC template benchmark consisting of 153 real-world scenarios cross 58 unique services. Also, we propose an LLM-based deployability-centric framework, dubbed IaCGen, that uses iterative feedback mechanism encompassing format verification, syntax checking, and live deployment stages, thereby closely mirroring the real DevOps workflows. Results show that IaCGen can make 54.6$\sim$91.6% generated IaC templates from all evaluated models deployable in the first 10 iterations. Additionally, human-in-the-loop feedback that provide direct guidance for the deployability errors, can further boost the performance to over 90% passItr@25 on all evaluated LLMs. Furthermore, we explore the trustworthiness of the generated IaC templates on user intent alignment and security compliance. The poor performance (25.2% user requirement coverage and 8.4% security compliance rate) indicates a critical need for continued research in this domain.

Wearable-informed generative digital avatars predict task-conditioned post-stroke locomotion

arXiv:2512.14329v2 Announce Type: replace-cross Abstract: Dynamic prediction of locomotor capacity after stroke could enable more individualized rehabilitation, yet current assessments largely provide static impairment scores and do not indicate whether patients can perform specific tasks such as slope walking or stair climbing. Here, we present a wearable-informed data-physics hybrid generative framework that reconstructs a stroke survivor's locomotor control from wearable inertial sensing and predicts task-conditioned post-stroke locomotion in new environments. From a single 20 m level-ground walking trial recorded by five IMUs, the framework personalizes a physics-based digital avatar using a healthy-motion prior and hybrid imitation learning, generating dynamically feasible, patient-specific movements for inclined walking and stair negotiation. Across 11 stroke inpatients, predicted postures reached 82.2% similarity for slopes and 69.9% for stairs, substantially exceeding a physics-only baseline. In a multicentre pilot randomized study (n = 21; 28 days), access to scenario-specific locomotion predictions to support task selection and difficulty titration was associated with larger gains in Fugl-Meyer lower-extremity scores than standard care (mean change 6.0 vs 3.7 points; $p

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer

npj Digital Medicine, Published online: 31 December 2025; doi:10.1038/s41746-025-02260-3

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.

From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research

arXiv:2512.23184v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

arXiv:2512.23545v1 Announce Type: cross Abstract: Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to produce predictions, without reassessment or targeted evidence acquisition under ambiguous diagnoses. This contrasts with clinical diagnostic workflows that refine hypotheses through repeated slide observations and further examination requests. We propose PathFound, an agentic multimodal model designed to support evidence-seeking inference in pathological diagnosis. PathFound integrates the power of pathological visual foundation models, vision-language models, and reasoning models trained with reinforcement learning to perform proactive information acquisition and diagnosis refinement by progressing through the initial diagnosis, evidence-seeking, and final decision stages. Across several large multimodal models, adopting this strategy consistently improves diagnostic accuracy, indicating the effectiveness of evidence-seeking workflows in computational pathology. Among these models, PathFound achieves state-of-the-art diagnostic performance across diverse clinical scenarios and demonstrates strong potential to discover subtle details, such as nuclear features and local invasions.

Immunotherapy for virus-related hepatocellular carcinoma: recent progress and future directions

Ann Med. 2026 Dec;58(1):2607229. doi: 10.1080/07853890.2025.2607229. Epub 2025 Dec 26.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality worldwide, with hepatitis B virus (HBV) and hepatitis C virus (HCV) infections remaining the predominant etiological factors. Chronic viral infection not only drives carcinogenesis but also reshapes the hepatic immune microenvironment, profoundly influencing the efficacy and safety of immunotherapy.

RECENT ADVANCES: Immune checkpoint inhibitors (ICIs) have revolutionized systemic therapy for advanced HCC, with agents targeting PD-1/PD-L1 demonstrating clinical benefit. Combination strategies - such as ICIs with anti-angiogenic therapies, multikinase inhibitors, or locoregional treatments - have shown synergistic efficacy and are now standard of care in certain settings. For virus-related HCC, antiviral therapy improves immune responsiveness and reduces risks such as HBV reactivation, underscoring the need for integrated management.

FUTURE PERSPECTIVES: Emerging therapeutic approaches include next-generation immune checkpoints (e.g. TIM-3, LAG-3, TIGIT), bispecific antibodies, cellular therapies (CAR-T, TCR-T, TILs), and tumor vaccines targeting viral or tumor-associated antigens. Advances in biomarker discovery, including circulating tumor DNA, immune signatures, and microbiome modulation, are expected to guide personalized treatment. Integration of multi-omics and clinical data will further refine patient selection and optimize treatment sequencing.

CONCLUSION: Immunotherapy offers new hope for patients with virus-related HCC, but challenges remain in response heterogeneity, resistance, and toxicity. Individualized strategies that combine immunotherapy with effective antiviral management and biomarker-|guided patient selection are essential. Continued translational and clinical research into virus-immune-tumor interactions will enable safer, more effective, and more durable treatment outcomes, ultimately transforming HCC into a more manageable disease.

PMID:41454610 | PMC:PMC12777805 | DOI:10.1080/07853890.2025.2607229

Developing and Evaluating Guidelines to Prevent Overdependence on Digital Therapeutics in Children and Adolescents: Randomized Controlled Trial

Background: Digital therapeutics (DTx) for children and adolescents with mental health problems have been developed in the health care industry. Despite reports of side effects from DTx for children and adolescents, there have been no guidelines to address the prevention of DTx overdependence among young users. Objective: This study aimed to identify the requirements for guidelines to prevent DTx overdependence in children and adolescents and to develop and evaluate these guidelines. Methods: We conducted 2 phases. This study first involved a phase I survey to develop guidelines, including assessments of smartphone usage and mental health conditions. The second phase evaluated the guidelines’ effectiveness, reliability, necessity, and satisfaction using a visual analog scale through a randomized controlled trial. Participants—45 children and adolescents aged 9-16 years and 42 caregivers—were randomly assigned to the experimental and control groups. Results: Phase I revealed that blocking mobile applications and notifications (mean 8.5, SD 1.8) and parental monitoring (mean 8.5, SD 2.1) were effective preventive features. Caregivers, children, and adolescents expressed concerns about the side effects and overdependence of DTx and decreased effects due to nonindividualized guidelines in subjective responses to the phase I survey. Based on these insights, personalized guidelines for phase II were developed, in which overall mean visual analog scale scores for guideline evaluation were higher in the experimental group, except for necessity among caregivers (mean 8.5, SD 1.3 versus mean 8.7, SD 1.2). Conclusions: Both caregivers and children and adolescents demonstrated the need for guidelines to prevent overdependence on DTx distinct from smartphone usage. Tailored guidelines may be acceptable for use in real-world therapeutic protocols. Guidelines to prevent overdependence on DTx in children and adolescents and to achieve a balance between their benefits and risks need to be established. Trial Registration: Clinical Research Information Service (CRiS) of the Republic of Korea KCT0008893; https://cris.nih.go.kr/cris/search/detailSearch.do?seq=25609
❌