❌

Reading view

Longitudinal liquid biopsy identifies an early predictive biomarker of immune checkpoint blockade response in head and neck squamous cell carcinoma

Nat Commun. 2025 Sep 1;16(1):8161. doi: 10.1038/s41467-025-63538-4.

ABSTRACT

Immune checkpoint blockade (ICB) has improved outcomes for patients with head and neck squamous cell carcinoma (HNSCC), but predictive biomarkers remain limited. Here, we use a time-resolved, multi-omic approach in a murine HNSCC model to characterize peripheral immune responses to ICB. Single-cell transcriptomics and T/B cell receptor analyses reveal early on-treatment expansion of effector memory T and B cell repertoires in responders, preceding tumor regression. These dynamic immune features inform a composite transcriptional signature that accurately predicts ICB response in independent human HNSCC cohorts. LiBIO outperforms existing biomarkers and generalizes to melanoma, non-small cell lung cancer, and breast cancer without retraining. These findings suggest that early treatment-induced changes in circulating immune repertoires reflect the host's capacity to mount an effective antitumor response. This work provides a framework for leveraging transient peripheral immune dynamics to develop non-invasive, high-fidelity biomarkers for response to immunotherapy across cancer types.

PMID:40890155 | PMC:PMC12402333 | DOI:10.1038/s41467-025-63538-4

  •  

Token Probabilities to Mitigate Large Language Models Overconfidence in Answering Medical Questions: Quantitative Study

Background: Chatbots have demonstrated promising capabilities in Medicine, scoring passing grades for board examinations across various specialties. However, their tendency to express high levels of confidence in their responses, even when incorrect, poses a limitation to their utility in clinical settings. Objective: To examine whether token probabilities outperform chatbots' Expressed Confidence levels in predict-ing the accuracy of their responses to medical questions. Methods: Seven large languages models (LLMs), comprising both commercial (GPT-3.5, GPT-4 and GPT-4o) and open-source (Llama 3-8b, Llama 3-70b, Phi-3-Mini, and Phi-3-Medium), were prompted to respond to a set of 2,522 questions from the US Medical Licensing Examination (MedQA database). Addition-ally, the models rated their confidence from 0 to 100 and the token probability of each response was extracted. The models’ success rates were measured, and the predictive performances of both Ex-pressed Confidence and Response Token Probability in predicting response accuracy were evaluated using Area Under the Receiver Operating Characteristic Curve (AUROC), Adapted Calibration Error (ACE) and Brier score. Sensitivity analyses were conducted using additional questions sourced from other databases in English (MedMCQA, n=2,797), Chinese (MedQA Main-land China, n=3,413 and Taiwan, n=2,808), and French (FrMedMCQA, n=1,079). Results: Overall, mean accuracy ranged from 52.7%[50.8-54.7] for Phi-3-Mini to 87.6%[86.2-88.9] for GPT-4o. Across the US Medical Licensing Examination questions, all chatbots consistently expressed high levels of confidence in their responses (ranging from 90[90-90] for Llama 3-70B to 100[100–100] for GPT-3.5). However, Expressed Confidence failed to predict response accuracy (AUROC ranging from 0.52[0.50-0.53] for Phi 3 Mini to 0.68[0.65-0.71] for GPT-4o). In contrast, the Response Token Probability consistently outperformed Expressed Confidence for predicting response accuracy (AU-ROC ranging from 0.67[0.65-0.69] for Phi-3-Mini to 0.83[0.81-0.85] for Llama 3-70B, all p-values
  •  

An eyecare foundation model for clinical assistance: a randomized controlled trial

Nature Medicine, Published online: 28 August 2025; doi:10.1038/s41591-025-03900-7

Trained and validated on multimodal data from 14.5 million images from multicountry datasets, a foundation model is shown to increase diagnostic and referral accuracy of clinicians when used as an assistant in a trial involving 16 ophthalmologists and 668 patients.
  •  
❌