❌

Reading view

Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation

arXiv:2602.15019v1 Announce Type: new Abstract: Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, non-English channels. Recent data suggests >85% of patent filings originate outside the U.S., with China accounting for nearly half of the global total; a growing share of scholarly output is also non-U.S. Industry estimates put China at ~30% of global drug development, spanning 1,200+ novel candidates. In this high-stakes environment, failing to surface "under-the-radar" assets creates multi-billion-dollar risk for investors and business development teams, making asset scouting a coverage-critical competition where speed and completeness drive value. Yet today's Deep Research AI agents still lag human experts in achieving high-recall discovery across heterogeneous, multilingual sources without hallucinations. We propose a benchmarking methodology for drug asset scouting and a tuned, tree-based self-learning Bioptic Agent aimed at complete, non-hallucinated scouting. We construct a challenging completeness benchmark using a multilingual multi-agent pipeline: complex user queries paired with ground-truth assets that are largely outside U.S.-centric radar. To reflect real deal complexity, we collected screening queries from expert investors, BD, and VC professionals and used them as priors to conditionally generate benchmark queries. For grading, we use LLM-as-judge evaluation calibrated to expert opinions. We compare Bioptic Agent against Claude Opus 4.6, OpenAI GPT-5.2 Pro, Perplexity Deep Research, Gemini 3 Pro + Deep Research, and Exa Websets. Bioptic Agent achieves 79.7% F1 versus 56.2% (Claude Opus 4.6), 50.6% (Gemini 3 Pro + Deep Research), 46.6% (GPT-5.2 Pro), 44.2% (Perplexity Deep Research), and 26.9% (Exa Websets). Performance improves steeply with additional compute, supporting the view that more compute yields better results.
  •  

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

arXiv:2602.13332v1 Announce Type: cross Abstract: Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose MedScope, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. By interleaving intermediate reasoning with targeted tool calls and verification on retrieved observations, MedScope produces more accurate and trustworthy predictions that are explicitly grounded in temporally localized visual evidence. To address the lack of high-fidelity supervision, we build ClinVideoSuite, an evidence-centric, fine-grained clinical video suite. We then optimize MedScope with Grounding-Aware Group Relative Policy Optimization (GA-GRPO), which directly reinforces tool use with grounding-aligned rewards and evidence-weighted advantages. On full and fine-grained video understanding benchmarks, MedScope achieves state-of-the-art performance in both in-domain and out-of-domain evaluations. Our approach illuminates a path toward medical AI agents that can genuinely "think with videos" through tool-integrated reasoning. We will release our code, models, and data.
  •  
  •  

Rare, Yet Targetable: New Perspectives on Ampullary Carcinomas

Int J Mol Sci. 2026 Feb 6;27(3):1597. doi: 10.3390/ijms27031597.

ABSTRACT

Ampullary carcinoma (AC) is a rare gastrointestinal malignancy with dual intestinal and pancreatobiliary differentiation, complicating diagnosis, staging, and treatment. This review synthesizes current epidemiology, pathology, and multi-omic data to outline a pragmatic care pathway: lineage-first at presentation, mutation-fast at progression. Histology remains the primary classifier: the intestinal subtype generally aligns with colorectal regimens, whereas pancreatobiliary and mixed subtypes favor pancreaticobiliary therapy. In selected fit patients, modified FOLFIRINOX may address mixed phenotypes. Next-generation sequencing adds precision by identifying therapeutically relevant alterations, including ERBB2/HER2 amplifications, MSI-high/dMMR, BRAF V600E, and rare NTRK or RET fusions, while KRAS mutations are enriched in pancreatobiliary tumors. We recommend early application of a rapid-core panel (KRAS/BRAF, MSI/dMMR, ERBB2/HER2, RNA-based fusions) to capture high-impact targets, followed by comprehensive profiling at first progression. Liquid biopsy, plasma circulating tumor DNA (ctDNA), or bile-derived DNA may complement tissue and help identify the dominant lineage. Research priorities include ampulla-enriched umbrella trials, explicit AC subcohorts in tissue-agnostic studies, and ctDNA-informed endpoints. This lineage-first, mutation-fast paradigm supports precision care and evidence generation in AC.

PMID:41684016 | PMC:PMC12897727 | DOI:10.3390/ijms27031597

  •  

STAT+: FDA’s rejection of Moderna threatens to stifle broader vaccine industry

The Food and Drug Administration’s refusal to review Moderna’s flu vaccine this month has renewed fears that Trump administration policies could paralyze the vaccine industry, dissuading companies from developing new shots in the U.S. and leaving the country flat-footed in the event of future pandemics. 

“I consider it an unprecedented action that really violates the basic principles of a data-driven regulatory agency and the fundamentals of public health, and it’s that simple,” said Gary Nabel, former head of the National Institutes of Health’s Vaccine Research Center and chief scientist at Sanofi, who now runs a vaccine and cancer startup. “It’s a destructive precedent that will undermine the future of vaccine development and the preeminence of American research.”

Executives at large vaccine developers were already grappling with a litany of changes to vaccine policy. Under Robert F. Kennedy Jr., a longtime vaccine critic, the Department of Health and Human Services has unilaterally removed six shots from the childhood vaccination schedule, canceled hundreds of millions of dollars in grants for mRNA shots, and fired and replaced a key immunization advisory board. 

Continue to STAT+ to read the full story…

© John Tlumacki/Globe Staff

  •  

STAT+: Researchers take another look at Apple’s hypertension feature

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Today, we’ve got a ton of updates including news about venture capital funding, telehealth policy, the government’s progress on information blocking, and research into the accuracy of Apple’s new hypertension feature.

Continue to STAT+ to read the full story…

© Business Wire via AP

  •  

STAT+: Pharmalittle: We’re reading about FDA rejecting a Moderna vaccine, compounding in the crosshairs and more

Hello, everyone, and welcome to the middle of the week. Congratulations on making it this far. It is an accomplishment, after all. The next step is to… keep going. And why not? Just consider the alternatives. On that optimistic note, please join us for a needed cup or three of stimulation. Our choice today is coconut rum. Meanwhile, here are some items of interest to get you going. Have a wonderful day and do drop us a line when you hear something juicy …

The U.S. Food and Drug Administration refused to review Moderna’s application for a new influenza vaccine, a surprise decision that could  raise concerns about the agency’s posture toward drug companies and the Trump administration’s policies on vaccines, STAT writes. Moderna, revealing the rejection, took the unusual step of releasing the letter it had received from Vinay Prasad, who heads the FDA’s biologics division. They also issued a strongly worded statement from its chief executive officer Stephane Bancel, who said the decision “does not further our shared goal of enhancing America’s leadership in developing innovative medicines.” At the heart of the dispute is what existing influenza vaccine Moderna should have used as a control when testing the efficacy of its new shot, which utilizes the same mRNA technology the company used in its Covid-19 vaccine.

The recent moves by the Trump administration against Hims & Hers might only be the start of a crackdown on compounding, STAT explains. In recent days, the Food and Drug Administration issued a warning, the Department of Health & Human Services asked the Department of Justice to open an investigation and, meanwhile, Novo Nordisk filed a patent infringement lawsuit against the company. But while compounded weight-loss drugs proliferated during recent shortages and continued to remain available, the flurry of developments underscores growing unease among regulators with mass-marketed compounded drugs sold by national, vertically integrated telehealth platforms. The FDA has so far focused publicly on misleading marketing, but signs that it may scrutinize compounding practices themselves have the industry on edge, given how many telehealth companies rely on compounded versions of everything from acne treatments to libido drugs.

Continue to STAT+ to read the full story…

© Alex Hogan/STAT

  •  

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

arXiv:2602.06609v1 Announce Type: cross Abstract: Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.
  •  

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

arXiv:2602.05183v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.
  •  

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Nature Medicine, Published online: 09 February 2026; doi:10.1038/s41591-025-04074-y

In a randomized controlled study involving 1,298 participants from a general sample, performance of humans when assisted by a large language model (LLM) was sensibly inferior to that of the LLM alone when assessing ten medical scenarios leading to disease identification and recommendations for treatment.
  •  

The Feasibility of Smartwatch Micro–Ecological Momentary Assessment for Tracking Eating Patterns of Malaysian Children and Adolescents in the South-East Asian Community Observatory Child Health Update 2020: Cross-Sectional Study

Background: Mobile phone ecological momentary assessment (EMA) methods are a well-established measure of eating and drinking behaviors, but compliance can be poor. Micro-EMA (μEMA), which collects information with a single tap response to brief questions on smartwatches, offers a novel application that may improve response rates. To our knowledge, there is no data evaluating μEMA to measure eating habits in children or in low-to-middle-income countries. Objective: In this study, we investigated the feasibility of micro-EMA to measure eating patterns in Malaysian children and adolescents. Methods: We invited 100 children and adolescents aged 7-18 years in Segamat, Malaysia, to participate in 2021-2022. Smartwatches were distributed to 83 children and adolescents who agreed to participate. Participants were asked to wear the smartwatch for 8 days and respond to 12 prompts per day, hourly, from 9AM to 8PM, asking for information on their meals, snacks, and drinks consumed. A questionnaire captured their experiences using the smartwatch and μEMA interface. Response rate (proportion of prompts responded to) assessed participants’ adherence. We explored associations between response rate with time of day, across days, age, and sex using multilevel binomial logistic regression modeling. Results: Eighty-two participants provided usable smartwatch data. The median number (IQR) of meals, drinks, and snacks per day was 2 (2-4), 3 (1-5), and 1 (0-2), respectively, on the first day of the study. The median response rate across the study was 68% (IQR 50-83). The response rate decreased across study days from 74% (68-78) on Day 1 to 40% (30-50) on Day 7 (odds ratio [OR] per study day 0.73, 95% CI 0.64-0.83). Response rate was lowest at the start of the day and highest between the hours of 12 PM and 2 PM. Female participants responded to more prompts than male participants (OR 1.72, 95% CI 1.03-2.86). There was no evidence of differential response by age (OR 0.73, 95% CI 0.41-1.28). Most participants (65%) rated their experience using the smartwatch positively, with 33% saying they were happy to participate in future studies using the smartwatch. For children that did not wear the smartwatch for the full study duration (n=22), discomfort was the most common complaint (41%). Conclusions: In this study of the feasibility of μEMA on smartwatches to measure eating in Malaysian children, we found the method was acceptable. However, response rates declined across study days, resulting in substantial missingness. Future studies (eg, through focus groups) should explore approaches to improving response to event prompts, trial alternative devices to increase children’s comfort, and evaluate revised protocols for reporting of intake events.
  •  

Tumor microbiome differences in early-onset versus average-onset pancreatic adenocarcinoma

ESMO Gastrointest Oncol. 2025 Jul 7;9:100194. doi: 10.1016/j.esmogo.2025.100194. eCollection 2025 Sep.

ABSTRACT

BACKGROUND: Compelling evidence supports the biomarker potential of microbiome in pancreatic adenocarcinoma. Given the knowledge gap on the characteristics and significance of microbiome in early-onset pancreatic ductal adenocarcinoma (eoPDAC, age <50 years), we aimed to evaluate microbiome profiles in resected specimens from individuals with eoPDAC and average-onset PDAC (aoPDAC, age >50 years).

MATERIALS AND METHODS: We carried out shotgun metagenomic sequencing in resected specimens from individuals with eoPDAC (n = 24) and aoPDAC (n = 20). Statistical tests included Wilcoxon test, permutational analysis of variance, multiomic classifier modeling, differential abundance analysis, and linear regression. All P values were adjusted for multiple testing and P < 0.05 was considered statistically significant.

RESULTS: We successfully sequenced several bacteria and fungi in the tumor specimens from 44 individuals with resected PDAC (24 eoPDAC and 20 aoPDAC). The alpha diversity of the bacterial microbiome was higher in eoPDAC tumor tissue compared with aoPDAC (P = 0.04). In contrast, the fungal mycobiome's alpha diversity was higher for aoPDAC tumor tissue (P = 0.02). Key organisms with differential abundance between tumor tissue from individuals with eoPDAC and aoPDAC included Bacillus, Candida, Collimonas, Cupriavidus, Enterobacter, Escherichia, Klebsiella, Malasseiza, Mucilaginibacter, Neisseria, and Sphingomonas. Higher bacterial diversity in tumor tissue was associated with better overall survival for individuals with eoPDAC (R = 0.26, P = 0.02).

CONCLUSIONS: Shotgun metagenomic sequencing identified bacterial microbiome and fungal mycobiome in tumors from individuals with eoPDAC and aoPDAC. We observed significant differences in alpha and beta diversity and relative abundances of organisms suggesting distinct microbiome signatures. Microbiome associations with survival were observed in eoPDAC indicating unique potential as prognostic biomarker.

PMID:41647993 | PMC:PMC12836659 | DOI:10.1016/j.esmogo.2025.100194

  •  

EcDNA-borne structural variants drive oncogenic fusion transcript amplification

Extrachromosomal DNA (ecDNA) is a major source of oncogenic fusions across cancer types, generating tissue-specific fusion landscapes with diagnostic potential. EcDNA-borne PVT1 5′-end fusions stabilize partner RNAs and boost oncogene output.
  •  

DISCOVER: Identifying Patterns of Daily Living in Human Activities from Smart Home Data

arXiv:2503.01733v3 Announce Type: replace-cross Abstract: Smart homes equipped with ambient sensors offer a transformative approach to continuous health monitoring and assisted living. Traditional research in this domain primarily focuses on Human Activity Recognition (HAR), which relies on mapping sensor data to a closed set of predefined activity labels. However, the fixed granularity of these labels often constrains their practical utility, failing to capture the subtle, household-specific nuances essential, for example, for tracking individual health over time. To address this, we propose DISCOVER, a framework for discovering and annotating Patterns of Daily Living (PDL) - fine-grained, recurring sequences of sensor events that emerge directly from a resident's unique routines. DISCOVER utilizes a self-supervised feature extraction and representation-aware clustering pipeline, supported by a custom visualization interface that enables experts to interpret and label discovered patterns with minimal effort. Our evaluation across multiple smart-home environments demonstrates that DISCOVER identifies cohesive behavioral clusters with high inter-rater agreement while achieving classification performance comparable to fully-supervised baselines using only 0.01% of the labels. Beyond reducing annotation overhead, DISCOVER establishes a foundation for longitudinal analysis. By grounding behavior in a resident's specific environment rather than rigid semantic categories, our framework facilitates the observation of within-person habitual drift. This capability positions the system as a potential tool for identifying subtle behavioral indicators associated with early-stage cognitive decline in future longitudinal studies.
  •  

Phenome-wide analysis of copy number variants in 470,727 UK Biobank genomes

Nature, Published online: 04 February 2026; doi:10.1038/s41586-025-10087-x

A multiancestry phenome-wide analysis of copy number variants in the UK Biobank genomes increases power to detect genetic associations with complex traits across human populations.
  •  

STAT+: AI doctors are coming. Should FDA make sure they’re safe?

When is an AI doctor a medical device?

Call it a sign of things to come. A startup called Doctronic made a splash recently when it announced the use AI to renew prescriptions without clinician input in the state of Utah. Something didn’t sit right with me about the announcement. Sure it got approval from Utah, but why isn’t it a medical device subject to Food and Drug Administration review? The company claimed it was “the practice of medicine” and so exempt from FDA authority. That didn’t seem entirely right either. 

So I did some asking around and after talking to over a dozen executives, legal scholars, and policy experts, it turns out the question is not nearly as clear-cut as Doctronic would have us believe. Indeed, it appears the company may be planning to market a medical device without authorization. In my story, I explain the law and why it all matters.

Read more here

Continue to STAT+ to read the full story…

© Adobe

  •  
❌