❌

Reading view

A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties

arXiv:2512.08185v1 Announce Type: cross Abstract: Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most researchers. Existing security benchmarks require GPU clusters, commercial API access, or protected health data -- barriers that limit community participation in this critical research area. We propose a practical, fully reproducible framework for evaluating medical AI security under realistic resource constraints. Our framework design covers multiple medical specialties stratified by clinical risk -- from high-risk domains such as emergency medicine and psychiatry to general practice -- addressing jailbreaking attacks (role-playing, authority impersonation, multi-turn manipulation) and privacy extraction attacks. All evaluation utilizes synthetic patient records requiring no IRB approval. The framework is designed to run entirely on consumer CPU hardware using freely available models, eliminating cost barriers. We present the framework specification including threat models, data generation methodology, evaluation protocols, and scoring rubrics. This proposal establishes a foundation for comparative security assessment of medical-specialist models and defense mechanisms, advancing the broader goal of ensuring safe and trustworthy medical AI systems.
  •  

ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Access

arXiv:2512.08193v1 Announce Type: cross Abstract: We present ClinicalTrialsHub, an interactive search-focused platform that consolidates all data from ClinicalTrials.gov and augments it by automatically extracting and structuring trial-relevant information from PubMed research articles. Our system effectively increases access to structured clinical trial data by 83.8% compared to relying on ClinicalTrials.gov alone, with potential to make access easier for patients, clinicians, researchers, and policymakers, advancing evidence-based medicine. ClinicalTrialsHub uses large language models such as GPT-5.1 and Gemini-3-Pro to enhance accessibility. The platform automatically parses full-text research articles to extract structured trial information, translates user queries into structured database searches, and provides an attributed question-answering system that generates evidence-grounded answers linked to specific source sentences. We demonstrate its utility through a user study involving clinicians, clinical researchers, and PhD students of pharmaceutical sciences and nursing, and a systematic automatic evaluation of its information extraction and question answering capabilities.
  •  

Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models

arXiv:2510.21084v2 Announce Type: replace-cross Abstract: Intelligent drug recommendation based on Electronic Health Records (EHRs) is critical for improving the quality and efficiency of clinical decision-making. By leveraging large-scale patient data, drug recommendation systems can assist physicians in selecting the most appropriate medications according to a patient's medical history, diagnoses, laboratory results, and comorbidities. Recent advances in large language models (LLMs) have shown remarkable capabilities in complex reasoning and medical text understanding, making them promising tools for drug recommendation tasks. However, the application of LLMs for Chinese clinical medication recommendation remains largely unexplored. In this work, we conduct a systematic investigation of LLM-based methodologies for Chinese discharge medication recommendation. We evaluate several representative LLM families (GLM, Llama, Qwen) under a unified methodological framework including zero-shot prompting, in-context learning, chain-of-thought prompting, and supervised fine-tuning using LoRA. We analyze model behavior across reasoning styles, error patterns, domain adaptation mechanisms, and robustness. Experimental results show that while supervised fine-tuning improves model performance, there remains substantial room for improvement, with the best model achieving the F1 score of 0.5648 and Jaccard score of 0.4477. Our findings highlight both the potential and limitations of LLMs for Chinese drug recommendation.
  •  

Systematic benchmarking of high-throughput subcellular spatial transcriptomics platforms across human tumors

Nat Commun. 2025 Oct 17;16(1):9232. doi: 10.1038/s41467-025-64292-3.

ABSTRACT

Recent advancements in spatial transcriptomics technologies have significantly enhanced resolution and throughput, underscoring an urgent need for systematic benchmarking. Here, we generate serial tissue sections from colon adenocarcinoma, hepatocellular carcinoma, and ovarian cancer samples for systematic evaluation. Using these uniformly processed samples, we generate spatial transcriptomics data across four high-throughput platforms with subcellular resolution: Stereo-seq v1.3, Visium HD FFPE, CosMx 6K, and Xenium 5K. To establish ground truth datasets, we profile proteins on tissue sections adjacent to all platforms using CODEX and perform single-cell RNA sequencing on the same samples. Leveraging manual nuclear segmentation and detailed annotations, we systematically assess each platform's performance across capture sensitivity, specificity, diffusion control, cell segmentation, cell annotation, spatial clustering, and concordance with adjacent CODEX. The uniformly generated and processed multi-omics dataset could advance computational method development and biological discoveries. The dataset is accessible via SPATCH, a user-friendly web server for visualization and download.

PMID:41107232 | PMC:PMC12534522 | DOI:10.1038/s41467-025-64292-3

  •  

Organoids in Genetic Disorders: from Disease Modeling to Translational Applications

Stem Cell Rev Rep. 2025 Sep 11. doi: 10.1007/s12015-025-10973-x. Online ahead of print.

ABSTRACT

The emergence of organoid models has significantly bridged the gap between traditional cell cultures/animal models and authentic human disease states, particularly for genetic disorders, where their inherent genetic fidelity enables more biologically relevant research directions and enhances translational validity. This review systematically analyzes established organoid models of genetic diseases across organs (e.g., brain, eye, kidney, lung, and heart), highlighting their pivotal roles in identifying novel pathogenic genes, elucidating disease mechanisms, and advancing therapeutic strategies such as drug screening platforms, gene-editing therapies, and organ transplantation strategies. Furthermore, we critically address current limitations-including challenges in recapitulating complex pathologies and scaling production-while underscoring their potential for personalized medicine through multi-omics integration and bioengineering innovations. Although the scope of "genetic diseases" is broad, this synthesis focuses on disorders with well-defined inheritance patterns, such as monogenic disorders, copy number variations (CNVs), and aneuploidies. Despite covering only a subset of these conditions, this review aims to provide researchers with a comprehensive overview of the field, emphasizing how organoid-based approaches could accelerate both mechanistic discoveries and clinical translation in genetic disease research.

PMID:40931310 | DOI:10.1007/s12015-025-10973-x

  •  

An eyecare foundation model for clinical assistance: a randomized controlled trial

Nature Medicine, Published online: 28 August 2025; doi:10.1038/s41591-025-03900-7

Trained and validated on multimodal data from 14.5 million images from multicountry datasets, a foundation model is shown to increase diagnostic and referral accuracy of clinicians when used as an assistant in a trial involving 16 ophthalmologists and 668 patients.
  •  
❌