❌

Reading view

A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains

npj Digital Medicine, Published online: 26 December 2025; doi:10.1038/s41746-025-02277-8

A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains
  •  

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

arXiv:2510.24161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.
  •  

Early Screening and Subtype Identification of High-Risk Lung Nodules via Breathprint by Graphene eNose Platform: A Large Cohort Study

ACS Sens. 2025 Apr 25;10(4):3101-3111. doi: 10.1021/acssensors.5c00314. Epub 2025 Apr 7.

ABSTRACT

Early screening of individuals with high-risk lung nodules can significantly improve the prognosis of lung cancer patients, and accurate identification of lung nodule subtypes can provide guidance for medical treatment. Exhaled breath (EB) analysis via eNoses offers a quick and noninvasive approach, but current eNose technology lacks quality control and solid validation in large population studies. Herein, an eNose platform integrated with a metal ion-decorated graphene sensor array and a breath sampling accessory was established. EB samples from 427 healthy subjects and 2586 subjects with lung nodules, including various benign and malignant subtypes, were collected through the breath sampling accessory for quality control. The large-cohort clinical EB samples were analyzed by the eNose platform to acquire the cross-reactive resistance response. Breathprint analysis for high-risk lung nodules using SVM and age-matched training sets yielded strong and robust performance. Combined with baseline data, the model achieved an AUC of 0.93 (95% CI, 0.89-0.96) on the external test set, with 97% sensitivity and 73% specificity. Moreover, dimensionality reduction analysis of breathprints demonstrated separability across different lung nodule subtypes. This study demonstrates the reliability of the graphene eNose platform to identify high-risk lung nodules and classify lung nodule subtypes in a noninvasive and rapid method.

PMID:40193324 | DOI:10.1021/acssensors.5c00314

  •  
❌