❌

Reading view

STAT+: FDA’s rejection of Moderna threatens to stifle broader vaccine industry

The Food and Drug Administration’s refusal to review Moderna’s flu vaccine this month has renewed fears that Trump administration policies could paralyze the vaccine industry, dissuading companies from developing new shots in the U.S. and leaving the country flat-footed in the event of future pandemics. 

“I consider it an unprecedented action that really violates the basic principles of a data-driven regulatory agency and the fundamentals of public health, and it’s that simple,” said Gary Nabel, former head of the National Institutes of Health’s Vaccine Research Center and chief scientist at Sanofi, who now runs a vaccine and cancer startup. “It’s a destructive precedent that will undermine the future of vaccine development and the preeminence of American research.”

Executives at large vaccine developers were already grappling with a litany of changes to vaccine policy. Under Robert F. Kennedy Jr., a longtime vaccine critic, the Department of Health and Human Services has unilaterally removed six shots from the childhood vaccination schedule, canceled hundreds of millions of dollars in grants for mRNA shots, and fired and replaced a key immunization advisory board. 

Continue to STAT+ to read the full story…

© John Tlumacki/Globe Staff

  •  

OpenAI Scales Single Primary Postgresql to Millions of Queries per Second for ChatGPT

OpenAI described how it scaled PostgreSQL to support ChatGPT and its API platform, handling millions of queries per second for hundreds of millions of users. By running a single-primary PostgreSQL deployment on Azure with nearly 50 read replicas, optimizing query patterns, and offloading write-heavy workloads to sharded systems, OpenAI maintained low-latency reads while managing write pressure.

By Leela Kumili
  •  

STAT+: Researchers take another look at Apple’s hypertension feature

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Good morning health tech readers!

Today, we’ve got a ton of updates including news about venture capital funding, telehealth policy, the government’s progress on information blocking, and research into the accuracy of Apple’s new hypertension feature.

Continue to STAT+ to read the full story…

© Business Wire via AP

  •  

Clinical utility of OGN in pan-cancer: diagnostic biomarker and immune microenvironment regulator

Transl Cancer Res. 2026 Jan 31;15(1):43. doi: 10.21037/tcr-2025-1499. Epub 2026 Jan 27.

ABSTRACT

BACKGROUND: Osteoglycin (OGN), an extracellular matrix protein, has emerging but poorly characterized roles in cancer. This study presents the first pan-cancer investigation of OGN's expression patterns, clinical significance, immune interactions, and functional mechanisms.

METHODS: Multi-omics data from Genotype Tissue Expression (GTEx), Cancer Cell Line Encyclopedia (CCLE), The Cancer Genome Atlas (TCGA), and Human Protein Atlas (HPA) databases were integrated. Differential expression was analyzed in normal tissues and tumor samples. Diagnostic utility was evaluated using area under the curve (AUC) of receiver operating characteristic (ROC) curve. Prognostic value was assessed via Kaplan-Meier [overall survival (OS); disease-specific survival (DSS); disease free interval (DFI); progression-free interval (PFI)] and Cox regression analyses. Immune microenvironment correlations were quantified using ESTIMATE, CIBERSORT, and gene set enrichment. Functional pathways were explored through gene set enrichment analysis (GSEA) and correlation with hallmark cancer signatures.

RESULTS: OGN was broadly expressed in normal tissues (brain, liver, kidney) but significantly downregulated in most tumor types (P<0.05, TCGA; validated at protein level, HPA). OGN demonstrated high diagnostic accuracy in pan-cancer (AUC: 0.703-0.990), achieving near-perfect performance in colon adenocarcinoma (COAD) (AUC: 0.966) and thyroid cancer (THCA) (AUC: 0.920). High OGN expression correlated with improved survival outcomes in thymoma (THYM) (OS/DSS) and cholangiocarcinoma (CHOL) (PFI/DFI), but worse prognosis in lung adenocarcinoma​/liver hepatocellular carcinoma​ (LUAD/LIHC), indicating cancer-type specificity. OGN expression strongly associated with immune cell infiltration (macrophages, natural killer cells, T cells), chemokine signaling, programmed death-ligand 1 (PD-L1) levels, microsatellite instability (MSI), and tumor mutation burden (TMB). GSEA revealed enrichment of OGN-linked genes in epithelial-mesenchymal transition (EMT), angiogenesis, JAK-STAT, and PI3K pathways across cancers.

CONCLUSIONS: Our pan-cancer analysis highlights OGN as a context-dependent regulator linking extracellular matrix (ECM) remodeling with immune and angiogenic signaling. Its pan-cancer dysregulation, diagnostic/prognostic value, and crosstalk with immune evasion mechanisms nominate OGN as a promising multi-functional biomarker and therapeutic target.

PMID:41674945 | PMC:PMC12885879 | DOI:10.21037/tcr-2025-1499

  •  

Upconversion mesoporous silica nanoparticles co-delivering celecoxib and rose bengal enable multimodal immunogenic and anti-angiogenic therapy for spinal metastasis of non-small cell lung cancer

Oncogene. 2026 Feb 11. doi: 10.1038/s41388-026-03679-y. Online ahead of print.

ABSTRACT

Non-small cell lung cancer (NSCLC) with spinal metastasis represents a clinical challenge due to its aggressive nature, limited treatment options, and profound impact on patient quality of life. Here, we report the development of an innovative upconversion mesoporous silica nanoparticle (UCMS) platform co-loaded with celecoxib and rose bengal (UCMS@CXB/RB), engineered to synergistically combine photodynamic therapy (PDT) and cyclooxygenase-2 (COX-2) inhibition. Upon near-infrared (NIR) irradiation, UCMS@CXB/RB generated abundant reactive oxygen species, triggered immunogenic cell death, and significantly suppressed prostaglandin E2 signaling, leading to reduced angiogenesis and improved antitumor immunity. In vitro and in vivo studies confirmed that this nanoplatform effectively remodeled the tumor microenvironment, inhibited tumor growth, and alleviated cancer-induced spinal dysfunction. Single-cell multi-omics analysis further revealed dynamic crosstalk among immune cells, tumor cells, and endothelial populations, providing mechanistic insights into the multifaceted therapeutic effects of UCMS@CXB/RB. Our results underscore the clinical potential of integrating PDT with targeted COX-2 blockade to address the complex pathophysiology of NSCLC spinal metastasis. This study presents a promising minimally invasive therapeutic strategy with strong translational relevance for managing metastatic NSCLC and improving patient outcomes.

PMID:41673094 | DOI:10.1038/s41388-026-03679-y

  •  

STAT+: The unusual Prasad missive in the FDA’s rejection of the Moderna flu shot application

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

So: Moderna says it was blindsided by the FDA, and Vinay Prasad, on its mRNA flu vaccine. Meanwhile, regulators are moving aggressively against Hims & Hers, and midsized biotechs have come together in solidarity — saying they could be crushed by President Trump’s drug pricing policy. Needless to say, it is not business as usual in Washington.

Out west, STAT’s Jonathan Wosen spoke with Novartis’ chief of biomedical research, who was in San Diego for the groundbreaking on a $1.1 billion research hub. She explained why company is pruning its pipeline and how it’s harnessing AI.

Continue to STAT+ to read the full story…

© David L Ryan/Globe Staff

  •  

STAT+: Pharmalittle: We’re reading about FDA rejecting a Moderna vaccine, compounding in the crosshairs and more

Hello, everyone, and welcome to the middle of the week. Congratulations on making it this far. It is an accomplishment, after all. The next step is to… keep going. And why not? Just consider the alternatives. On that optimistic note, please join us for a needed cup or three of stimulation. Our choice today is coconut rum. Meanwhile, here are some items of interest to get you going. Have a wonderful day and do drop us a line when you hear something juicy …

The U.S. Food and Drug Administration refused to review Moderna’s application for a new influenza vaccine, a surprise decision that could  raise concerns about the agency’s posture toward drug companies and the Trump administration’s policies on vaccines, STAT writes. Moderna, revealing the rejection, took the unusual step of releasing the letter it had received from Vinay Prasad, who heads the FDA’s biologics division. They also issued a strongly worded statement from its chief executive officer Stephane Bancel, who said the decision “does not further our shared goal of enhancing America’s leadership in developing innovative medicines.” At the heart of the dispute is what existing influenza vaccine Moderna should have used as a control when testing the efficacy of its new shot, which utilizes the same mRNA technology the company used in its Covid-19 vaccine.

The recent moves by the Trump administration against Hims & Hers might only be the start of a crackdown on compounding, STAT explains. In recent days, the Food and Drug Administration issued a warning, the Department of Health & Human Services asked the Department of Justice to open an investigation and, meanwhile, Novo Nordisk filed a patent infringement lawsuit against the company. But while compounded weight-loss drugs proliferated during recent shortages and continued to remain available, the flurry of developments underscores growing unease among regulators with mass-marketed compounded drugs sold by national, vertically integrated telehealth platforms. The FDA has so far focused publicly on misleading marketing, but signs that it may scrutinize compounding practices themselves have the industry on edge, given how many telehealth companies rely on compounded versions of everything from acne treatments to libido drugs.

Continue to STAT+ to read the full story…

© Alex Hogan/STAT

  •  

Agentic AI in healthcare: How Life Sciences marketing could achieve $450B in value by 2028

Agentic AI in healthcare is graduating from answering prompts to autonomously executing complex marketing tasks – and life sciences companies are betting their commercial strategies on it.

According to a recent report cited by Capgemini Invent, AI agents could generate up to $450 billion in economic value through revenue uplift and cost savings globally by 2028, with 69% of executives planning to deploy agents in marketing processes by year’s end.

The stakes are particularly high in pharmaceutical marketing, where sales representatives have increasingly limited face-time with healthcare professionals (HCPs) – a trend accelerated by Covid-19. The challenge isn’t just access; it’s making those rare interactions count with intelligence that’s currently trapped in data silos.

The fragmented intelligence problem

Briggs Davidson, senior director of digital, data & marketing strategy for life Sciences at Capgemini Invent, outlines a scenario that will sound familiar to anyone in pharma marketing: An HCP attends a conference where a competitor showcases promising drug results, publishes research, and shifts their prescriptions to a rival product – in a single quarter.

“In most companies, legacy IT infrastructure and data silos keep this information in disparate systems in CRM, events databases and claims data,” Davidson writes. “Chances are, none of that information was accessible to sales reps before they met with the HCP.”

The solution, according to Davidson, isn’t to connect these systems, it’s deploying agentic AI in healthcare marketing to autonomously query, synthesising and acting on unified data. Unlike conversational AI that responds to queries, agentic systems can independently execute multi-step tasks.

Instead of a data engineer building a new pipeline, an AI agent could autonomously query the CRM and claims database to answer business questions like: “Identify oncologists in the Northwest who have a 20% lower prescription volume but attended our last medical congress.”

From orchestration to autonomous execution

Davidson frames the change as moving from an “omnichannel view” – coordinating experiences in channels – to true orchestration powered by agentic AI.

In practice, this means a sales representative could have an agent assist with call and visit planning by asking: “What messages has my HCP responded to most recently?” or “Can you create a detailed intelligence brief on my HCP?”

The agentic system would compile:

  • Their most recent conversation with the HCP,
  • The HCP’s prescribing behaviour,
  • Thought-leaders the HCP follows,
  • Relevant content to share,
  • The HCP’s preferred outreach channels (in-person visits, emails, webinars).

More significantly, the AI agent would then create a custom call plan for each HCP based on their unified profile and recommend follow-up steps based on engagement outcomes. “Agentic AI systems are about driving action, graduating from ‘answer my prompt,’ to ‘autonomously execute my task,'” Davidson explains.

“That means evolving the sales representative mindset from asking questions to coordinating small teams of specialised agents that work together: one plans, another retrieves and checks content, a third schedules and measures, and a fourth enforces compliance guardrails – all under human oversight.”

The AI-ready data prerequisite

The operational promise hinges on what Davidson calls “AI-ready data” – standardised, accessible, complete, and trustworthy information that enables three abilities:

Faster decision making: Predictive analytics that provide near real-time alerts on what’s about to happen, letting sales representatives act proactively.

Personalisation at scale: Delivering customised experiences to thousands of HCPs simultaneously with small human teams enabled by specialised agent networks.

True marketing ROI: Moving beyond monthly historical reports to understanding which marketing activities are actively driving prescriptions.

Davidson emphasises that successful deployment starts with marketing and IT alignment on initial use cases, with stakeholders identifying KPIs that demonstrate tangible outcomes – like specific percentage increases in HCP engagement or sales representative productivity.

Critical implementation questions

The article frames agentic AI in healthcare as “not simply another technology-led ability; it’s a new operating layer for commercial teams.” But it acknowledges that “agentic AI’s full value only materialises with AI-ready data, trustworthy deployment and workflow redesign.”

What remains unaddressed is the regulatory and compliance complexity of autonomous systems querying claims databases containing prescriber behaviour, particularly under HIPAA’s minimum necessary standard. The piece also doesn’t detail actual client implementations or metrics beyond the aspirational $450B economic value projection.

For global organisations, Davidson says use cases “can and should be tailored to fit each market’s maturity for maximum ROI,” suggesting that deployment will vary in regulatory environments. The fundamental value proposition, according to Davidson, centres on bidirectional benefit: “The HCP receives directly relevant content, and the marketing teams can drive increased HCP engagement and conversion.”

Whether that vision of autonomous marketing agents coordinating in CRM, events, and claims systems becomes standard practice by 2028 – or remains constrained by data governance realities – will likely determine if life sciences achieves anything close to that $450 billion opportunity.

See also: China’s hyperscalers bet billions on agentic AI as commerce becomes the new battleground

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Agentic AI in healthcare: How Life Sciences marketing could achieve $450B in value by 2028 appeared first on AI News.

  •  

Spatial and multi-omics transcriptomic dissects platinum resistance in lung adenocarcinoma: a five-gene predictive model with tumor microenvironment dynamics

Chem Biol Interact. 2026 Feb 7:111952. doi: 10.1016/j.cbi.2026.111952. Online ahead of print.

ABSTRACT

The scarcity of reliable biomarkers and predictive models for platinum resistance in lung adenocarcinoma (LUAD) poses a significant clinical challenge. This study endeavors to identify molecular subtypes related to platinum resistance and construct a robust predictive model through multi-omics techniques. We performed integrative analysis of public datasets using advanced bioinformatics strategies, including spatial transcriptome deconvolution and consensus clustering. Bulk RNA deconvolution analysis was conducted to characterize tumor microenvironment heterogeneity. Feature selection was performed using the Supervised Principal Component (SuperPC) algorithm, followed by diagnostic model construction validated through receiver operating characteristic (ROC) analysis. Functional validation was performed through cytological experiments measuring cisplatin IC50 alterations following gene manipulation in LUAD cell lines. Consensus clustering revealed distinct LUAD subtypes, with Cluster1 demonstrating significant platinum resistance. We first subtyped the patients in the bulk transcriptome data based on consistency clustering, and then analyzed the differences between different platinum-resistant subtypes (Cluster 1 and Cluster 2), so as to screen 333 isotype-specific differentially expressed genes and 15 platinum resistance-related (PRR) genes were selected through machine learning. A refined 5-gene signature (ANKRD29/CACNA2D2/DSP/HSD17B6/SPP1) achieved exceptional predictive performance (AUC=0.9639). Spatial transcriptomics demonstrated compartmentalized expression patterns: SPP1/DSP localized to tumor niches, HSD17B6/CACNA2D2 to epithelial regions, and ANKRD29 depletion in stromal areas. Cellular colocalization analysis revealed malignant epithelial PH proximity to myeloid and mast cells. Functional validation confirmed that ANKRD29/CACNA2D2 overexpression sensitized A549/DDP cells to cisplatin, while DSP/SPP1/HSD17B6 overexpression induced resistance. Experiments in nude mice have shown that these genes are closely related to cisplatin resistance in LUAD. This study identifies the Cluster1 subtype and malignant epithelial PH as crucial determinants of platinum resistance in LUAD. Our innovative 5-gene predictive model exhibits clinical-grade diagnostic accuracy, and spatial transcriptomic characterization offers mechanistic insights into the dynamics of the tumor microenvironment.

PMID:41662930 | DOI:10.1016/j.cbi.2026.111952

  •  

Problems and Barriers Regarding the Admission, Financing, and Service Provision of Digital Health Apps: Qualitative Stakeholder Survey

Background: Since their introduction with the Digital Care Act in 2019, DiGA are a part of the German statutory healthcare system. In order to become a DiGA, mHealth apps have to complete a certification process covering both technical and evidence related aspects. After completion, DiGA are added to the DiGA-directory, containing a list of all reimbursable DiGA within German statutory health insurance (SHI). The first apps were added at the end of 2020 with the number steadily increasing. The novelty of the introduction leads to problems and barriers to optimal use along the way, which is studied from different stakeholder perspectives in this research article. Objective: The aim of the survey was to identify problems and barriers in the context of certification, financing and use of DiGA in Germany. Methods: We used semi-structured expert interviews to evaluate the perspective of stakeholders of the German healthcare system on DiGA. The interview guide was developed according to Helfferich, the interviews were transcribed and analyzed using the qualitative content approach by Mayring and Kuckartz. Results: We identified problems from stakeholder perspectives regarding the certification/admission, financing and service distribution regarding DiGA. The interviewed stakeholders reported problems with authorization of DiGA and the corresponding process. DiGA prices and the different negotiation positions were criticized, as well as financial challenges for smaller DiGA-manufacturers. Within service provision, technical problems, e. g., with activation codes or software surrounding DiGA-prescription were mentioned. Problems were also seen in insufficient knowledge and skills on the side of the patients as well as the medical providers. Conclusions: mHealth applications provide potentially disruptive innovations within the healthcare sector. Nevertheless, since the evidence-based and regulated use of this technology is relatively new there are still problems and barriers limiting the optimized, patient-centered use. This study provides an overview of problems in the context of DiGA in Germany from the stakeholder perspective. Since other countries showed interest in potentially adopting the German system, valuable implications can be drawn from this survey.
  •  

Quantifying Individual Health Status from Multi-omics Data by Health State Manifold

Phenomics. 2025 Dec 15;5(5):469-486. doi: 10.1007/s43657-024-00188-4. eCollection 2025 Oct.

ABSTRACT

Quantifying individual health status from increasingly accumulated omics data is essential for both early prevention and intervention of diseases, which attracts great attention from communities of biology and medicine. Most of the existing approaches mainly classify individuals into different catalogues or classes based on phenotypes and biomarkers. However, an individual's health status from a dynamical systems viewpoint can be viewed as a non-equilibrium steady state, which can generally be characterized by two key features, i.e. (1) homeostatic potential that represents the ability of homeostatic resilience to withstand perturbations or maintain functions at the current state/phenotype of this individual and (2) phenotypic potential that represents the state/phenotype of the individual on the whole process from health to disease. Here, we proposed a health state manifold (HSM) method derived from dynamic network biomarker method and diffusion map theory to quantify individual health status with the characterization of such two features in a robust and accurate manner based on multi-omics data. To verify our method, HSM method was applied to the quantification of diabetes mellitus (rat subjects) and the Roux-en-Y Gastric Bypass (human subjects) for both disease progression process and recovery process, which demonstrated its effectiveness and potential for personalized medicine and preventive medicine.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s43657-024-00188-4.

PMID:41659741 | PMC:PMC12881232 | DOI:10.1007/s43657-024-00188-4

  •  

Spatial Multi-omics Analyses Reveal Diabetes Promotes Pancreatic Cancer Progression by Stimulating Cholesterol-Induced Neutrophil Extracellular Trap Formation

Cancer Res. 2026 Feb 9. doi: 10.1158/0008-5472.CAN-25-2854. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) patients with diabetes mellitus (DM) exhibit poor clinical outcomes. Metabolic reprogramming of both cancer cells and immune compartments plays a crucial role in shaping the anti-tumor immune response in PDAC. DM-induced metabolic alteration may disrupt the intricate crosstalk between immune cells and tumor-associated immune factors, profoundly influencing PDAC progression. Here, we performed an integrated, spatially resolved multi-omics study to investigate DM-associated, cell-specific metabolic remodeling within the PDAC tumor microenvironment. DM influenced interactions between tumor cells and immune cells, which accelerated PDAC growth in both humans and mice. PDAC patients with DM exhibited higher tumor-stage, poorer differentiation, and worse outcomes. Spatial metabolic and transcriptional profiling revealed that SREBP2-dependent cholesterol biosynthesis exacerbated PDAC progression. Increased cholesterol biosynthesis promoted neutrophil recruitment and accelerated formation of neutrophil extracellular traps (NETs) by stimulating the CXCL1-CXCR1/CXCR2 signaling axis, ultimately promoting PDAC growth. Inhibition of SREBP2, pharmacological blockade of CXCL1, or perturbation of NETs markedly reduced PDAC growth in diabetic mouse models. Together, these multi-omics analyses and follow-up mechanistic studies constitute an integrated approach that elucidates a metabolic mechanism by which diabetes promotes PDAC development by remodeling the tumor immune microenvironment and highlights a potential therapeutic strategy for PDAC with DM.

PMID:41661642 | DOI:10.1158/0008-5472.CAN-25-2854

  •  

Generating High-quality Privacy-preserving Synthetic Data

arXiv:2602.06390v1 Announce Type: cross Abstract: Synthetic tabular data enables sharing and analysis of sensitive records, but its practical deployment requires balancing distributional fidelity, downstream utility, and privacy protection. We study a simple, model agnostic post processing framework that can be applied on top of any synthetic data generator to improve this trade off. First, a mode patching step repairs categories that are missing or severely underrepresented in the synthetic data, while largely preserving learned dependencies. Second, a k nearest neighbor filter replaces synthetic records that lie too close to real data points, enforcing a minimum distance between real and synthetic samples. We instantiate this framework for two neural generative models for tabular data, a feed forward generator and a variational autoencoder, and evaluate it on three public datasets covering credit card transactions, cardiovascular health, and census based income. We assess marginal and joint distributional similarity, the performance of models trained on synthetic data and evaluated on real data, and several empirical privacy indicators, including nearest neighbor distances and attribute inference attacks. With moderate thresholds between 0.2 and 0.35, the post processing reduces divergence between real and synthetic categorical distributions by up to 36 percent and improves a combined measure of pairwise dependence preservation by 10 to 14 percent, while keeping downstream predictive performance within about 1 percent of the unprocessed baseline. At the same time, distance based privacy indicators improve and the success rate of attribute inference attacks remains largely unchanged. These results provide practical guidance for selecting thresholds and applying post hoc repairs to improve the quality and empirical privacy of synthetic tabular data, while complementing approaches that provide formal differential privacy guarantees.
  •  

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

arXiv:2602.06609v1 Announce Type: cross Abstract: Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.
  •  

Yunjue Agent Tech Report: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks

arXiv:2601.18226v2 Announce Type: replace Abstract: Conventional agent systems often struggle in open-ended environments where task distributions continuously drift and external supervision is scarce. Their reliance on static toolsets or offline training lags behind these dynamics, leaving the system's capability boundaries rigid and unknown. To address this, we propose the In-Situ Self-Evolving paradigm. This approach treats sequential task interactions as a continuous stream of experience, enabling the system to distill short-term execution feedback into long-term, reusable capabilities without access to ground-truth labels. Within this framework, we identify tool evolution as the critical pathway for capability expansion, which provides verifiable, binary feedback signals. Within this framework, we develop Yunjue Agent, a system that iteratively synthesizes, optimizes, and reuses tools to navigate emerging challenges. To optimize evolutionary efficiency, we further introduce a Parallel Batch Evolution strategy. Empirical evaluations across five diverse benchmarks under a zero-start setting demonstrate significant performance gains over proprietary baselines. Additionally, complementary warm-start evaluations confirm that the accumulated general knowledge can be seamlessly transferred to novel domains. Finally, we propose a novel metric to monitor evolution convergence, serving as a function analogous to training loss in conventional optimization. We open-source our codebase, system traces, and evolved tools to facilitate future research in resilient, self-evolving intelligence.
  •  

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

arXiv:2602.05183v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.
  •  

Exploring AI-Augmented Sensemaking of Patient-Generated Health Data: A Mixed-Method Study with Healthcare Professionals in Cardiac Risk Reduction

arXiv:2602.05687v2 Announce Type: replace-cross Abstract: Individuals are increasingly generating substantial personal health and lifestyle data, e.g. through wearables and smartphones. While such data could transform preventative care, its integration into clinical practice is hindered by its scale, heterogeneity and the time pressure and data literacy of healthcare professionals (HCPs). We explore how large language models (LLMs) can support sensemaking of patient-generated health data (PGHD) with automated summaries and natural language data exploration. Using cardiovascular disease (CVD) risk reduction as a use case, 16 HCPs reviewed multimodal PGHD in a mixed-methods study with a prototype that integrated common charts, LLM-generated summaries, and a conversational interface. Findings show that AI summaries provided quick overviews that anchored exploration, while conversational interaction supported flexible analysis and bridged data-literacy gaps. However, HCPs raised concerns about transparency, privacy, and overreliance. We contribute empirical insights and sociotechnical design implications for integrating AI-driven summarization and conversation into clinical workflows to support PGHD sensemaking.
  •  

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Nature Medicine, Published online: 09 February 2026; doi:10.1038/s41591-025-04074-y

In a randomized controlled study involving 1,298 participants from a general sample, performance of humans when assisted by a large language model (LLM) was sensibly inferior to that of the LLM alone when assessing ten medical scenarios leading to disease identification and recommendations for treatment.
  •  
❌