❌

Normal view

MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questions

arXiv:2602.17308v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for diagnostic tasks in medicine. In clinical practice, the correct diagnosis can rarely be immediately inferred from the initial patient presentation alone. Rather, reaching a diagnosis often involves systematic history taking, during which clinicians reason over multiple potential conditions through iterative questioning to resolve uncertainty. This process requires considering differential diagnoses and actively excluding emergencies that demand immediate intervention. Yet, the ability of medical LLMs to generate informative follow-up questions and thus reason over differential diagnoses remains underexplored. Here, we introduce MedClarify, an AI agent for information-seeking that can generate follow-up questions for iterative reasoning to support diagnostic decision-making. Specifically, MedClarify computes a list of candidate diagnoses analogous to a differential diagnosis, and then proactively generates follow-up questions aimed at reducing diagnostic uncertainty. By selecting the question with the highest expected information gain, MedClarify enables targeted, uncertainty-aware reasoning to improve diagnostic performance. In our experiments, we first demonstrate the limitations of current LLMs in medical reasoning, which often yield multiple, similarly likely diagnoses, especially when patient cases are incomplete or relevant information for diagnosis is missing. We then show that our information-theoretic reasoning approach can generate effective follow-up questioning and thereby reduces diagnostic errors by ~27 percentage points (p.p.) compared to a standard single-shot LLM baseline. Altogether, MedClarify offers a path to improve medical LLMs through agentic information-seeking and to thus promote effective dialogues with medical LLMs that reflect the iterative and uncertain nature of real-world clinical reasoning.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Intent Laundering: AI Safety Datasets Are Not What They Seem Shahriar Golchin · Marc Wetter
    arXiv:2602.16729v1 Announce Type: cross Abstract: We systematically evaluate the quality of widely used AI safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world attacks based on three key properties: driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger sa
     

Intent Laundering: AI Safety Datasets Are Not What They Seem

arXiv:2602.16729v1 Announce Type: cross Abstract: We systematically evaluate the quality of widely used AI safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world attacks based on three key properties: driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger safety mechanisms explicitly, which is unrealistic compared to real-world attacks. In practice, we evaluate whether these datasets genuinely measure safety risks or merely provoke refusals through triggering cues. To explore this, we introduce "intent laundering": a procedure that abstracts away triggering cues from attacks (data points) while strictly preserving their malicious intent and all relevant details. Our results indicate that current AI safety datasets fail to faithfully represent real-world attacks due to their overreliance on triggering cues. In fact, once these cues are removed, all previously evaluated "reasonably safe" models become unsafe, including Gemini 3 Pro and Claude Sonnet 3.7. Moreover, when intent laundering is adapted as a jailbreaking technique, it consistently achieves high attack success rates, ranging from 90% to over 98%, under fully black-box access. Overall, our findings expose a significant disconnect between how model safety is evaluated and how real-world adversaries behave.
  • ✇STAT
  • STAT+: Key study of Grail’s cancer detection test fails in setback for company Matthew Herper and Angus Chen
    A blood test for detecting cancer early being developed by the diagnostics firm Grail failed to meet its main goal in a giant study being conducted with England’s National Health Service, the company said Thursday. Grail’s test has been the standard bearer for new technologies that promise a blood test can be used to detect many different types of cancer early and eventually even to indicate to scientists where in the body to look for tumors. The company already sells its test, called Galleri
     

STAT+: Key study of Grail’s cancer detection test fails in setback for company

20 February 2026 at 06:38

A blood test for detecting cancer early being developed by the diagnostics firm Grail failed to meet its main goal in a giant study being conducted with England’s National Health Service, the company said Thursday.

Grail’s test has been the standard bearer for new technologies that promise a blood test can be used to detect many different types of cancer early and eventually even to indicate to scientists where in the body to look for tumors. The company already sells its test, called Galleri, for a list price of $1,000, although it is not yet approved by the Food and Drug Administration. Grail said Thursday it sold 185,000 tests in 2025, generating $136.8 million. 

The company’s shares were down 47% in after-hours trading.

Continue to STAT+ to read the full story…

© Adobe

  • ✇STAT
  • STAT+: AI-guided cancer treatments, telehealth usage, and other health tech news Mario Aguilar
    You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. Good morning health tech readers! Please join me in congratulating my colleagues as STAT, who have been honored with a fourth Polk Award for our coverage of the Trump administration’s impacts on the federal health department and American science. The award recognizes the whole newsroom and in pa
     

STAT+: AI-guided cancer treatments, telehealth usage, and other health tech news

19 February 2026 at 22:25

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Please join me in congratulating my colleagues as STAT, who have been honored with a fourth Polk Award for our coverage of the Trump administration’s impacts on the federal health department and American science. The award recognizes the whole newsroom and in particular the work of Lizzy Lawrence covering a dramatic year of changes at the Food and Drug Administration. 

Continue to STAT+ to read the full story…

© Adobe

Live biotherapeutics in cancer therapy

Prog Mol Biol Transl Sci. 2026;220:361-403. doi: 10.1016/bs.pmbts.2026.01.002. Epub 2026 Jan 23.

ABSTRACT

Cancer poses a global challenge in diagnostics and therapeutics. Treatments like chemotherapy, radiotherapy, surgery, and immunotherapy have significantly decreased the fatality rate, but drug resistance, therapy side effects, and relapse remain as major concerns. Live biotherapeutics are microorganisms that can be developed as therapeutic agents to modulate cancer pathophysiology and aid in disease management. Live biotherapeutic products (LBPs) have the potential to suppress tumour growth, enhance the effectiveness of conventional therapies, and reduce treatment-related side effects. Dysbiosis in the gut and cancer-specific tissues is linked to cancers of the colon, stomach, pancreas, and liver. Live biotherapeutics aim either to re-establish microbial balance or to employ microbes directly as anticancer tools. Both native and engineered LBPs (bacteria and viruses) represent promising interventions that may form part of next-generation cancer treatment strategies. Their clinical application draws on the integration of microbiology, immunology, synthetic biology, and oncology. LBPs can be used to target cancer cells by delivering antitumour payloads such as immune modulators, toxins, exposing cancer antigens, and molecules for targeted killing. LBPs offer advantages such as reduced systemic toxicity, overcoming drug resistance, and synergy with chemo-, radio-, and immunotherapies. Despite challenges in safety, manufacturing, regulation, and personalization, advances in synthetic biology and omics are enabling precision approaches. Future innovations such as bacteriobots, biocontainment systems, and patient-specific microbiome integration highlight their potential as next-generation cancer therapeutics.

PMID:41714084 | DOI:10.1016/bs.pmbts.2026.01.002

Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation

arXiv:2602.15019v1 Announce Type: new Abstract: Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, non-English channels. Recent data suggests >85% of patent filings originate outside the U.S., with China accounting for nearly half of the global total; a growing share of scholarly output is also non-U.S. Industry estimates put China at ~30% of global drug development, spanning 1,200+ novel candidates. In this high-stakes environment, failing to surface "under-the-radar" assets creates multi-billion-dollar risk for investors and business development teams, making asset scouting a coverage-critical competition where speed and completeness drive value. Yet today's Deep Research AI agents still lag human experts in achieving high-recall discovery across heterogeneous, multilingual sources without hallucinations. We propose a benchmarking methodology for drug asset scouting and a tuned, tree-based self-learning Bioptic Agent aimed at complete, non-hallucinated scouting. We construct a challenging completeness benchmark using a multilingual multi-agent pipeline: complex user queries paired with ground-truth assets that are largely outside U.S.-centric radar. To reflect real deal complexity, we collected screening queries from expert investors, BD, and VC professionals and used them as priors to conditionally generate benchmark queries. For grading, we use LLM-as-judge evaluation calibrated to expert opinions. We compare Bioptic Agent against Claude Opus 4.6, OpenAI GPT-5.2 Pro, Perplexity Deep Research, Gemini 3 Pro + Deep Research, and Exa Websets. Bioptic Agent achieves 79.7% F1 versus 56.2% (Claude Opus 4.6), 50.6% (Gemini 3 Pro + Deep Research), 46.6% (GPT-5.2 Pro), 44.2% (Perplexity Deep Research), and 26.9% (Exa Websets). Performance improves steeply with additional compute, supporting the view that more compute yields better results.

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

arXiv:2602.13332v1 Announce Type: cross Abstract: Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose MedScope, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. By interleaving intermediate reasoning with targeted tool calls and verification on retrieved observations, MedScope produces more accurate and trustworthy predictions that are explicitly grounded in temporally localized visual evidence. To address the lack of high-fidelity supervision, we build ClinVideoSuite, an evidence-centric, fine-grained clinical video suite. We then optimize MedScope with Grounding-Aware Group Relative Policy Optimization (GA-GRPO), which directly reinforces tool use with grounding-aligned rewards and evidence-weighted advantages. On full and fine-grained video understanding benchmarks, MedScope achieves state-of-the-art performance in both in-domain and out-of-domain evaluations. Our approach illuminates a path toward medical AI agents that can genuinely "think with videos" through tool-integrated reasoning. We will release our code, models, and data.

Rare, Yet Targetable: New Perspectives on Ampullary Carcinomas

Int J Mol Sci. 2026 Feb 6;27(3):1597. doi: 10.3390/ijms27031597.

ABSTRACT

Ampullary carcinoma (AC) is a rare gastrointestinal malignancy with dual intestinal and pancreatobiliary differentiation, complicating diagnosis, staging, and treatment. This review synthesizes current epidemiology, pathology, and multi-omic data to outline a pragmatic care pathway: lineage-first at presentation, mutation-fast at progression. Histology remains the primary classifier: the intestinal subtype generally aligns with colorectal regimens, whereas pancreatobiliary and mixed subtypes favor pancreaticobiliary therapy. In selected fit patients, modified FOLFIRINOX may address mixed phenotypes. Next-generation sequencing adds precision by identifying therapeutically relevant alterations, including ERBB2/HER2 amplifications, MSI-high/dMMR, BRAF V600E, and rare NTRK or RET fusions, while KRAS mutations are enriched in pancreatobiliary tumors. We recommend early application of a rapid-core panel (KRAS/BRAF, MSI/dMMR, ERBB2/HER2, RNA-based fusions) to capture high-impact targets, followed by comprehensive profiling at first progression. Liquid biopsy, plasma circulating tumor DNA (ctDNA), or bile-derived DNA may complement tissue and help identify the dominant lineage. Research priorities include ampulla-enriched umbrella trials, explicit AC subcohorts in tissue-agnostic studies, and ctDNA-informed endpoints. This lineage-first, mutation-fast paradigm supports precision care and evidence generation in AC.

PMID:41684016 | PMC:PMC12897727 | DOI:10.3390/ijms27031597

  • ✇STAT
  • STAT+: FDA’s rejection of Moderna threatens to stifle broader vaccine industry Jason Mast
    The Food and Drug Administration’s refusal to review Moderna’s flu vaccine this month has renewed fears that Trump administration policies could paralyze the vaccine industry, dissuading companies from developing new shots in the U.S. and leaving the country flat-footed in the event of future pandemics.  “I consider it an unprecedented action that really violates the basic principles of a data-driven regulatory agency and the fundamentals of public health, and it’s that simple,” said Gary Nab
     

STAT+: FDA’s rejection of Moderna threatens to stifle broader vaccine industry

13 February 2026 at 01:49

The Food and Drug Administration’s refusal to review Moderna’s flu vaccine this month has renewed fears that Trump administration policies could paralyze the vaccine industry, dissuading companies from developing new shots in the U.S. and leaving the country flat-footed in the event of future pandemics. 

“I consider it an unprecedented action that really violates the basic principles of a data-driven regulatory agency and the fundamentals of public health, and it’s that simple,” said Gary Nabel, former head of the National Institutes of Health’s Vaccine Research Center and chief scientist at Sanofi, who now runs a vaccine and cancer startup. “It’s a destructive precedent that will undermine the future of vaccine development and the preeminence of American research.”

Executives at large vaccine developers were already grappling with a litany of changes to vaccine policy. Under Robert F. Kennedy Jr., a longtime vaccine critic, the Department of Health and Human Services has unilaterally removed six shots from the childhood vaccination schedule, canceled hundreds of millions of dollars in grants for mRNA shots, and fired and replaced a key immunization advisory board. 

Continue to STAT+ to read the full story…

© John Tlumacki/Globe Staff

  • ✇STAT
  • STAT+: Researchers take another look at Apple’s hypertension feature Mario Aguilar
    You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. Good morning health tech readers! Today, we’ve got a ton of updates including news about venture capital funding, telehealth policy, the government’s progress on information blocking, and research into the accuracy of Apple’s new hypertension feature.Continue to STAT+ to read the full story…
     

STAT+: Researchers take another look at Apple’s hypertension feature

12 February 2026 at 21:42

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Today, we’ve got a ton of updates including news about venture capital funding, telehealth policy, the government’s progress on information blocking, and research into the accuracy of Apple’s new hypertension feature.

Continue to STAT+ to read the full story…

© Business Wire via AP

STAT+: Pharmalittle: We’re reading about FDA rejecting a Moderna vaccine, compounding in the crosshairs and more

11 February 2026 at 22:24

Hello, everyone, and welcome to the middle of the week. Congratulations on making it this far. It is an accomplishment, after all. The next step is to… keep going. And why not? Just consider the alternatives. On that optimistic note, please join us for a needed cup or three of stimulation. Our choice today is coconut rum. Meanwhile, here are some items of interest to get you going. Have a wonderful day and do drop us a line when you hear something juicy …

The U.S. Food and Drug Administration refused to review Moderna’s application for a new influenza vaccine, a surprise decision that could  raise concerns about the agency’s posture toward drug companies and the Trump administration’s policies on vaccines, STAT writes. Moderna, revealing the rejection, took the unusual step of releasing the letter it had received from Vinay Prasad, who heads the FDA’s biologics division. They also issued a strongly worded statement from its chief executive officer Stephane Bancel, who said the decision “does not further our shared goal of enhancing America’s leadership in developing innovative medicines.” At the heart of the dispute is what existing influenza vaccine Moderna should have used as a control when testing the efficacy of its new shot, which utilizes the same mRNA technology the company used in its Covid-19 vaccine.

The recent moves by the Trump administration against Hims & Hers might only be the start of a crackdown on compounding, STAT explains. In recent days, the Food and Drug Administration issued a warning, the Department of Health & Human Services asked the Department of Justice to open an investigation and, meanwhile, Novo Nordisk filed a patent infringement lawsuit against the company. But while compounded weight-loss drugs proliferated during recent shortages and continued to remain available, the flurry of developments underscores growing unease among regulators with mass-marketed compounded drugs sold by national, vertically integrated telehealth platforms. The FDA has so far focused publicly on misleading marketing, but signs that it may scrutinize compounding practices themselves have the industry on edge, given how many telehealth companies rely on compounded versions of everything from acne treatments to libido drugs.

Continue to STAT+ to read the full story…

© Alex Hogan/STAT

Advancing healthcare AI governance through a comprehensive maturity model based on systematic review

npj Digital Medicine, Published online: 11 February 2026; doi:10.1038/s41746-026-02418-7

Advancing healthcare AI governance through a comprehensive maturity model based on systematic review

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

arXiv:2602.06609v1 Announce Type: cross Abstract: Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

arXiv:2602.05183v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Nature Medicine, Published online: 09 February 2026; doi:10.1038/s41591-025-04074-y

In a randomized controlled study involving 1,298 participants from a general sample, performance of humans when assisted by a large language model (LLM) was sensibly inferior to that of the LLM alone when assessing ten medical scenarios leading to disease identification and recommendations for treatment.

The Feasibility of Smartwatch Micro–Ecological Momentary Assessment for Tracking Eating Patterns of Malaysian Children and Adolescents in the South-East Asian Community Observatory Child Health Update 2020: Cross-Sectional Study

Background: Mobile phone ecological momentary assessment (EMA) methods are a well-established measure of eating and drinking behaviors, but compliance can be poor. Micro-EMA (μEMA), which collects information with a single tap response to brief questions on smartwatches, offers a novel application that may improve response rates. To our knowledge, there is no data evaluating μEMA to measure eating habits in children or in low-to-middle-income countries. Objective: In this study, we investigated the feasibility of micro-EMA to measure eating patterns in Malaysian children and adolescents. Methods: We invited 100 children and adolescents aged 7-18 years in Segamat, Malaysia, to participate in 2021-2022. Smartwatches were distributed to 83 children and adolescents who agreed to participate. Participants were asked to wear the smartwatch for 8 days and respond to 12 prompts per day, hourly, from 9AM to 8PM, asking for information on their meals, snacks, and drinks consumed. A questionnaire captured their experiences using the smartwatch and μEMA interface. Response rate (proportion of prompts responded to) assessed participants’ adherence. We explored associations between response rate with time of day, across days, age, and sex using multilevel binomial logistic regression modeling. Results: Eighty-two participants provided usable smartwatch data. The median number (IQR) of meals, drinks, and snacks per day was 2 (2-4), 3 (1-5), and 1 (0-2), respectively, on the first day of the study. The median response rate across the study was 68% (IQR 50-83). The response rate decreased across study days from 74% (68-78) on Day 1 to 40% (30-50) on Day 7 (odds ratio [OR] per study day 0.73, 95% CI 0.64-0.83). Response rate was lowest at the start of the day and highest between the hours of 12 PM and 2 PM. Female participants responded to more prompts than male participants (OR 1.72, 95% CI 1.03-2.86). There was no evidence of differential response by age (OR 0.73, 95% CI 0.41-1.28). Most participants (65%) rated their experience using the smartwatch positively, with 33% saying they were happy to participate in future studies using the smartwatch. For children that did not wear the smartwatch for the full study duration (n=22), discomfort was the most common complaint (41%). Conclusions: In this study of the feasibility of μEMA on smartwatches to measure eating in Malaysian children, we found the method was acceptable. However, response rates declined across study days, resulting in substantial missingness. Future studies (eg, through focus groups) should explore approaches to improving response to event prompts, trial alternative devices to increase children’s comfort, and evaluate revised protocols for reporting of intake events.
❌