❌

Reading view

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

arXiv:2602.13332v1 Announce Type: cross Abstract: Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose MedScope, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. By interleaving intermediate reasoning with targeted tool calls and verification on retrieved observations, MedScope produces more accurate and trustworthy predictions that are explicitly grounded in temporally localized visual evidence. To address the lack of high-fidelity supervision, we build ClinVideoSuite, an evidence-centric, fine-grained clinical video suite. We then optimize MedScope with Grounding-Aware Group Relative Policy Optimization (GA-GRPO), which directly reinforces tool use with grounding-aligned rewards and evidence-weighted advantages. On full and fine-grained video understanding benchmarks, MedScope achieves state-of-the-art performance in both in-domain and out-of-domain evaluations. Our approach illuminates a path toward medical AI agents that can genuinely "think with videos" through tool-integrated reasoning. We will release our code, models, and data.
  •  

AI Deception: Risks, Dynamics, and Controls

arXiv:2511.22619v2 Announce Type: replace Abstract: As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.
  •  

Longitudinal liquid biopsy identifies an early predictive biomarker of immune checkpoint blockade response in head and neck squamous cell carcinoma

Nat Commun. 2025 Sep 1;16(1):8161. doi: 10.1038/s41467-025-63538-4.

ABSTRACT

Immune checkpoint blockade (ICB) has improved outcomes for patients with head and neck squamous cell carcinoma (HNSCC), but predictive biomarkers remain limited. Here, we use a time-resolved, multi-omic approach in a murine HNSCC model to characterize peripheral immune responses to ICB. Single-cell transcriptomics and T/B cell receptor analyses reveal early on-treatment expansion of effector memory T and B cell repertoires in responders, preceding tumor regression. These dynamic immune features inform a composite transcriptional signature that accurately predicts ICB response in independent human HNSCC cohorts. LiBIO outperforms existing biomarkers and generalizes to melanoma, non-small cell lung cancer, and breast cancer without retraining. These findings suggest that early treatment-induced changes in circulating immune repertoires reflect the host's capacity to mount an effective antitumor response. This work provides a framework for leveraging transient peripheral immune dynamics to develop non-invasive, high-fidelity biomarkers for response to immunotherapy across cancer types.

PMID:40890155 | PMC:PMC12402333 | DOI:10.1038/s41467-025-63538-4

  •  
❌