❌

Normal view

Tri-Reader: An Open-Access, Multi-Stage AI Pipeline for First-Pass Lung Nodule Annotation in Screening CT

arXiv:2601.19380v1 Announce Type: cross Abstract: Using multiple open-access models trained on public datasets, we developed Tri-Reader, a comprehensive, freely available pipeline that integrates lung segmentation, nodule detection, and malignancy classification into a unified tri-stage workflow. The pipeline is designed to prioritize sensitivity while reducing the candidate burden for annotators. To ensure accuracy and generalizability across diverse practices, we evaluated Tri-Reader on multiple internal and external datasets as compared with expert annotations and dataset-provided reference standards.

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

arXiv:2508.10530v2 Announce Type: replace Abstract: The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize the expected reward that reflects human preferences. Recently, Direct Preference Optimization~(DPO) was proposed as a LM alignment method that directly optimize the policy from static preference data, and further improved by incorporating on-policy sampling~(i.e., preference candidates generated during the training loop) for better LM alignment. However, we show on-policy data is not always optimal, with systematic effectiveness difference emerging between static and on-policy preference candidates. For example, on-policy data can result in a $3\times$ effectiveness compared with static data for Llama-3, and a $0.4\times$ effectiveness for Zephyr. To explain the phenomenon, we propose the alignment stage assumption, which divides the alignment process into two distinct stages: the preference injection stage, which benefits from diverse data, and the preference fine-tuning stage, which favors high-quality data. Through theoretical and empirical analysis, we characterize these stages and propose an effective algorithm to identify the boundaries between them. We perform experiments on $5$ models~(Llama, Zephyr, Phi-2, Qwen, Pythia) and $2$ alignment methods~(DPO, SLiC-HF) to show the generalizability of alignment stage assumption and the effectiveness of the boundary measurement algorithm.

Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning

arXiv:2510.02091v4 Announce Type: replace Abstract: Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance loss. However, such claims are typically drawn from narrow evaluations and may overlook important aspects of model behavior. In this work, we present a systematic study of depth utilization across diverse dimensions, including evaluation protocols, task categories, and model architectures. Our analysis confirms that very deep layers are generally less effective than earlier ones, but their contributions vary substantially with the evaluation setting. Under likelihood-based metrics without generation, pruning most layers preserves performance, with only the initial few being critical. By contrast, generation-based evaluation uncovers indispensable roles for middle and deeper layers in enabling reasoning and maintaining long-range coherence. We further find that knowledge and retrieval are concentrated in shallow components, whereas reasoning accuracy relies heavily on deeper layers -- yet can be reshaped through distillation. These results highlight that depth usage in LLMs is highly heterogeneous and context-dependent, underscoring the need for task-, metric-, and model-aware perspectives in both interpreting and compressing large models.

Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery

arXiv:2601.12542v2 Announce Type: replace Abstract: Artificial intelligence systems for scientific discovery have demonstrated remarkable potential, yet existing approaches remain largely proprietary and operate in batch-processing modes requiring hours per research cycle, precluding real-time researcher guidance. This paper introduces Deep Research, a multi-agent system enabling interactive scientific investigation with turnaround times measured in minutes. The architecture comprises specialized agents for planning, data analysis, literature search, and novelty detection, unified through a persistent world state that maintains context across iterative research cycles. Two operational modes support different workflows: semi-autonomous mode with selective human checkpoints, and fully autonomous mode for extended investigations. Evaluation on the BixBench computational biology benchmark demonstrated state-of-the-art performance, achieving 48.8% accuracy on open response and 64.4% on multiple-choice evaluation, exceeding existing baselines by 14 to 26 percentage points. Analysis of architectural constraints, including open access literature limitations and challenges inherent to automated novelty assessment, informs practical deployment considerations for AI-assisted scientific workflows.

AI-generated data contamination erodes pathological variability and diagnostic reliability

arXiv:2601.12946v3 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.

Products, Performance, and Technological Development of Ambulatory Oxygen Therapy Devices: Scoping Review

Background: Ambulatory oxygen therapy is prescribed for patients with chronic lung diseases who experience exertional hypoxemia. However, available devices may not adequately meet user requirements, and their performance characteristics are heterogeneous. Objective: This study aims to identify devices available for delivery of ambulatory oxygen therapy, the technologies that they use to generate oxygen, the performance characteristics of each device, and the development status. Methods: We used medical and engineering databases to identify peer-reviewed papers (eg, MEDLINE, IEEE). Gray literature was used to identify additional descriptions of ambulatory oxygen devices in military medicine, space exploration, or patents. The last search was conducted in September 2025. Documents that described a device that can deliver oxygen in an ambulatory context (defined as weighing less than 10 kg) and were written in English were included. Search results were screened for inclusion by 2 independent reviewers. Data were synthesized by descriptively mapping the performance of each product, the technology used, and the development status of emerging technologies. Results: From 9702 records identified, a total of 166 met eligibility criteria (106 scientific publications and 60 gray literature). We identified 33 portable oxygen concentrators (POCs; 29 commercially available), 10 oxygen cylinders, and 6 portable liquid oxygen (LOX) devices. The POC products showed a trade-off between portability and oxygen delivery capacity (maximum flow rate ranging from 2.0 to 6.0 L/min; device weight ranging from 1.0 to 9.1 kg). Pressure swing adsorption with zeolite was the most common oxygen generation technology in POCs on the market. The mean maximum continuous operating time of POCs was 3.8 hours. Two prototype POCs (maximum flow rate of 4-6 L/min and device weight of 8-9 kg) were developed for space exploration using modified adsorbents. LOX devices were the lightest and had the longest continuous operating time. Innovations in delivery included the downsizing of a POC by using nanozeolite as an adsorbent and pulse oximeter oxygen saturation (SpO2)–targeted automatic titration of oxygen delivery based on the user’s SpO2. Conclusions: This scoping review is the first study to integrate medical, engineering, and gray literature on ambulatory oxygen devices and their development. Although prior literature has narratively explained the products and technologies, no previous research has systematically investigated them. This review showed that POCs available to consumers may not meet the needs of patients in terms of flow rate, portability, and operating time. LOX devices offered superior performance but are limited by high costs. Limitations of this review include the difficulty of comparing product performance across oxygen delivery settings and that the records were largely obtained from English-language sources. Innovation in ambulatory oxygen technology has been limited over the past decade, highlighting urgent need for research and development of new lightweight devices with higher oxygen delivery. Clinical Trial: OSF Registries 10.17605/OSF.IO/QS7FX; https://osf.io/qs7fx

Exploring the key molecular mechanisms and immune microenvironment of oxidative stress-related pathways in pancreatic neuroendocrine tumor combining scRNA-seq and bulk RNA

Discov Oncol. 2026 Jan 27. doi: 10.1007/s12672-026-04515-1. Online ahead of print.

ABSTRACT

BACKGROUND: Pancreatic neuroendocrine tumor (pNET) is a heterogeneous tumor originating from pancreatic endocrine cells. Emerging evidence suggests that oxidative stress plays a crucial role in pNET pathogenesis, yet the precise molecular mechanisms and their interplay with the tumor microenvironment remain unclear. This study aims to systematically elucidate how oxidative stress-related pathways drive pNET progression through an integrated multi-omics approach.

METHODS: We designed a three-tier analytical strategy to address interconnected scientific questions. First, to identify which oxidative stress-related genes are dysregulated in pNET, we performed differential expression analysis and weighted gene co-expression network analysis (WGCNA) on the GSE73338 dataset (63 pNET samples, 5 controls), intersecting the. results with oxidative stress gene sets to obtain 71 candidate genes. Second, to understand the functional implications of these genes, we conducted GO/KEGG enrichment analysis and constructed protein-protein interaction (PPI) networks, from which we identified BCL2L1 and PHGDH as key hub genes using three independent algorithms. We then assessed their diagnostic value through ROC analysis and built a prognostic nomogram model. Third, to explore how these key genes influence the tumor microenvironment, we performed immune infiltration analysis using CIBERSORTx. Fourth, to reveal upstream regulatory mechanisms, we constructed ceRNA networks and predicted transcription factors. Fifth, to identify potential therapeutic interventions, we conducted drug prediction and molecular docking analyses. Finally, to validate our findings at cellular resolution and understand cellular heterogeneity, we analyzed single-cell RNA sequencing data from GSE256136 (20 samples), identifying cell types, quantifying cell-cell communications, and confirming key gene expression patterns across different cell populations.

RESULTS: Our systematic analysis revealed that oxidative stress-related genes in pNET were significantly enriched in the PI3K-Akt signaling pathway, cysteine and methionine metabolism, and HIF-1 signaling pathway. BCL2L1 and PHGDH emerged as central regulators with excellent diagnostic performance (AUC > 0.9). Immune infiltration analysis demonstrated significant alterations in activated dendritic cells, memory B cells, and resting NK cells, which correlated strongly with BCL2L1 and PHGDH expression, suggesting these genes link oxidative stress to immune dysfunction. The ceRNA network centered on KCNQ1OT1 and hsa-miR-15a-5p revealed multi-layered transcriptional and post-transcriptional regulation. Drug prediction identified sertindole and cabozantinib as promising therapeutic candidates. Single-cell analysis identified 11 cell types and confirmed that endocrine cells are the primary site of BCL2L1 and PHGDH dysregulation, with extensive crosstalk between endocrine cells and T cells potentially mediating immune evasion.

CONCLUSION: Through integrated multi-omics analysis, we established that oxidative stress pathways may drive pNET progression through a coordinated mechanism involving metabolic reprogramming (via BCL2L1 and PHGDH downregulation), immune microenvironment remodeling (through altered dendritic cell and NK cell function), and complex regulatory networks. BCL2L1 and PHGDH represent potential diagnostic biomarkers and candidate therapeutic targets that require experimental validation, providing new directions for precision medicine in pNET.

PMID:41591671 | DOI:10.1007/s12672-026-04515-1

China’s innovation in translational medicine: rethinking early-stage clinical development

Nature Biotechnology, Published online: 27 January 2026; doi:10.1038/s41587-025-02998-x

As pressure mounts globally on drug pricing and development cost continues to rise, clinicians and translational scientists in biotech, academia and biopharma companies are re-evaluating when, where and how to launch early clinical programs. These initial patient data become critical to de-risk development programs and allow developers to deploy their limited time and resources on the most promising drugs. We evaluate four fundamental shifts in drug development that appear to be unfolding and may well become critical to future global biopharma success: use of large-scale high quality cohort studies, sponsor-driven investigator-initiated trials, the integration of affordable artificial intelligence with extensive high quality data registries, and China’s focus on precision medicine. —
❌