❌

Reading view

The Path Ahead for Agentic AI: Challenges and Opportunities

arXiv:2601.02749v1 Announce Type: new Abstract: The evolution of Large Language Models (LLMs) from passive text generators to autonomous, goal-driven systems represents a fundamental shift in artificial intelligence. This chapter examines the emergence of agentic AI systems that integrate planning, memory, tool use, and iterative reasoning to operate autonomously in complex environments. We trace the architectural progression from statistical models to transformer-based systems, identifying capabilities that enable agentic behavior: long-range reasoning, contextual awareness, and adaptive decision-making. The chapter provides three contributions: (1) a synthesis of how LLM capabilities extend toward agency through reasoning-action-reflection loops; (2) an integrative framework describing core components perception, memory, planning, and tool execution that bridge LLMs with autonomous behavior; (3) a critical assessment of applications and persistent challenges in safety, alignment, reliability, and sustainability. Unlike existing surveys, we focus on the architectural transition from language understanding to autonomous action, emphasizing the technical gaps that must be resolved before deployment. We identify critical research priorities, including verifiable planning, scalable multi-agent coordination, persistent memory architectures, and governance frameworks. Responsible advancement requires simultaneous progress in technical robustness, interpretability, and ethical safeguards to realize potential while mitigating risks of misalignment and unintended consequences.
  •  

Organoids in translation: a bench-to-bedside framework for pancreatic cancer precision medicine

J Transl Med. 2026 Jan 6. doi: 10.1186/s12967-025-07596-8. Online ahead of print.

ABSTRACT

INTRODUCTION: Pancreatic ductal adenocarcinoma (PDAC) is one of the most lethal malignancies with a 5-year survival rate of < 13%. Standard treatments such as FOLFIRINOX or gemcitabine/nab-paclitaxel yield modest response rates, underscoring the urgent need for precision oncology approaches. Patient-derived organoids (PDOs) preserve the genomic, phenotypic, and histopathological features of the source tumor and offer a promising platform for drug screening, biomarker development, and personalized therapy. However, a systematic evaluation of their translational capacities is lacking.

METHODS: A systematic review was conducted according to the PRISMA 2020 guidelines (PROSPERO registration pending) using PubMed, EMBASE, and Cochrane CENTRAL (December 10, 2024) to identify English-language PDAC PDO studies that incorporated therapeutic testing. Ninety-five studies met the inclusion criteria. Data extraction captured >75 variables per study, including spanning culture methodology, therapeutic profiling, biomarker integration, and clinical correlation. A 13-domain weighted Translatability Scoring Framework adapted from Wehling et al. assessed predictive validity, biomarker strength, pharmacogenetics, and clinical trial alignment. Scores ranged from 0 to 5 and were categorized as good (>4.0), moderate (3.0-4.0), or low (<3.0) translational potential.

RESULTS: Of the 95 studies, 70.5% have been published since 2021, reflecting the rapid growth in this field. The mean PDO generation success rate was 89.7%, with the primary tumor tissue being the predominant source (48.4%). Only 24.8% were directly linked to clinical trials and 5.3% incorporated multi-omic profiling. The median translatability score was 3.13 (range, 1.72-4.59): 45.3% of the studies had low translatability, 50.5% moderate, and only 4.2% had good translational potential. High-scoring studies consistently combine multi-omic biomarker platforms, in vivo validation, clinical outcome correlation, and prospective trial integration. Conversely, the weakest domains were pharmacogenetics, endpoint strategies, and biomarker validation, limiting their overall clinical relevance.

CONCLUSIONS: PDOs have demonstrated strong feasibility and in vitro clinical correlation in PDAC; however, their clinical translation remains constrained by limited multi-omic integration, absence of pharmacogenomic modeling, and sparse clinical trial embedding. Standardization of protocols, adoption of harmonized and clinically relevant endpoints, and systematic incorporation of biomarker-driven co-clinical trial frameworks are urgently needed to transition PDOs from promising experimental surrogates to validating precision oncology tools capable of informing therapeutic decision-making in PDAC.

PMID:41495743 | DOI:10.1186/s12967-025-07596-8

  •  

STAT+: FDA announces sweeping changes to oversight of wearables, AI-enabled devices

LAS VEGAS — The Food and Drug Administration announced Tuesday that it will ease regulation of digital health products, following through on the Trump administration’s promises to deregulate artificial intelligence and promote its widespread use.

FDA Commissioner Marty Makary indicated that one of the agency’s priorities is fostering an environment that’s good for investors, and that FDA regulation needs to move “at Silicon Valley speed.” He announced the changes during an address to conference attendees at the Consumer Electronics Show.

The agency will soften its approach to the regulation of clinical decision support software, which include AI-enabled products that help doctors navigate diagnoses and treatment options. The agency previously considered products that delivered a single recommendation as FDA-regulated medical devices. Now, those products can enter the market without FDA review as long as they fulfill the agency’s other criteria for escaping regulation. 

Continue to STAT+ to read the full story…

© ANDREW CABALLERO-REYNOLDS/AFP via Getty Images

  •  

Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models

arXiv:2601.01321v1 Announce Type: new Abstract: Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.
  •  

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

arXiv:2601.01718v1 Announce Type: new Abstract: We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks while maintaining competitive capabilities on general-purpose tasks. To address the overthinking phenomenon commonly observed in Large Reasoning Models (LRMs), we propose Reflection-aware Adaptive Policy Optimization (RAPO), a novel RL training algorithm that effectively regulates overthinking behaviors. In enterprise-oriented tasks such as retrieval-augmented generation (RAG), complex table understanding, and summarization, Yuan3.0 Flash consistently achieves superior performance. Moreover, it also demonstrates strong reasoning capabilities in domains such as mathematics, science, etc., attaining accuracy comparable to frontier model while requiring only approximately 1/4 to 1/2 of the average tokens. Yuan3.0 Flash has been fully open-sourced to facilitate further research and real-world deployment: https://github.com/Yuan-lab-LLM/Yuan3.0.
  •  

MACA: A Framework for Distilling Trustworthy LLMs into Efficient Retrievers

arXiv:2601.00926v1 Announce Type: cross Abstract: Modern enterprise retrieval systems must handle short, underspecified queries such as ``foreign transaction fee refund'' and ``recent check status''. In these cases, semantic nuance and metadata matter but per-query large language model (LLM) re-ranking and manual labeling are costly. We present Metadata-Aware Cross-Model Alignment (MACA), which distills a calibrated metadata aware LLM re-ranker into a compact student retriever, avoiding online LLM calls. A metadata-aware prompt verifies the teacher's trustworthiness by checking consistency under permutations and robustness to paraphrases, then supplies listwise scores, hard negatives, and calibrated relevance margins. The student trains with MACA's MetaFusion objective, which combines a metadata conditioned ranking loss with a cross model margin loss so it learns to push the correct answer above semantically similar candidates with mismatched topic, sub-topic, or entity. On a proprietary consumer banking FAQ corpus and BankFAQs, the MACA teacher surpasses a MAFA baseline at Accuracy@1 by five points on the proprietary set and three points on BankFAQs. MACA students substantially outperform pretrained encoders; e.g., on the proprietary corpus MiniLM Accuracy@1 improves from 0.23 to 0.48, while keeping inference free of LLM calls and supporting retrieval-augmented generation.
  •  

OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment

arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via semantic search engine; (3) constructing a hierarchical taxonomy of core-task-related work and performing contribution-level full-text comparisons against each contribution; and (4) synthesizing all analyses into a structured novelty report with explicit citations and evidence snippets. Unlike naive LLM-based approaches, \textsc{OpenNovelty} grounds all assessments in retrieved real papers, ensuring verifiable judgments. We deploy our system on 500+ ICLR 2026 submissions with all reports publicly available on our website, and preliminary analysis suggests it can identify relevant prior work, including closely related papers that authors may overlook. OpenNovelty aims to empower the research community with a scalable tool that promotes fair, consistent, and evidence-backed peer review.
  •  

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

arXiv:2601.01627v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety of LLMs for Japanese healthcare. Our benchmark is based on 67 guidelines from the Japan Medical Association and contains over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Using a dual-LLM scoring protocol, we evaluate 27 models and find that commercial models maintain robust safety while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.0, $p
  •  

EHRSummarizer: A Privacy-Aware, FHIR-Native Architecture for Structured Clinical Summarization of Electronic Health Records

arXiv:2601.01668v1 Announce Type: cross Abstract: Clinicians routinely navigate fragmented electronic health record (EHR) interfaces to assemble a coherent picture of a patient's problems, medications, recent encounters, and longitudinal trends. This work describes EHRSummarizer, a privacy-aware, FHIR-native reference architecture that retrieves a targeted set of high-yield FHIR R4 resources, normalizes them into a consistent clinical context package, and produces structured summaries intended to support structured chart review. The system can be configured for data minimization, stateless processing, and flexible deployment, including local inference within an organization's trust boundary. To mitigate the risk of unsupported or unsafe behavior, the summarization stage is constrained to evidence present in the retrieved context package, is intended to indicate missing or unavailable domains where feasible, and avoids diagnostic or treatment recommendations. Prototype demonstrations on synthetic and test FHIR environments illustrate end-to-end behavior and output formats; however, this manuscript does not report clinical outcomes or controlled workflow studies. We outline an evaluation plan centered on faithfulness, omission risk, temporal correctness, usability, and operational monitoring to guide future institutional assessments.
  •  

How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System

arXiv:2508.17215v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) augmented with Retrieval-Augmented Generation (RAG) are increasingly employed in medical AI to enhance factual grounding through external clinical image-text retrieval. However, this reliance creates a significant attack surface. We propose MedThreatRAG, a novel multimodal poisoning framework that systematically probes vulnerabilities in medical RAG systems by injecting adversarial image-text pairs. A key innovation of our approach is the construction of a simulated semi-open attack environment, mimicking real-world medical systems that permit periodic knowledge base updates via user or pipeline contributions. Within this setting, we introduce and emphasize Cross-Modal Conflict Injection (CMCI), which embeds subtle semantic contradictions between medical images and their paired reports. These mismatches degrade retrieval and generation by disrupting cross-modal alignment while remaining sufficiently plausible to evade conventional filters. While basic textual and visual attacks are included for completeness, CMCI demonstrates the most severe degradation. Evaluations on IU-Xray and MIMIC-CXR QA tasks show that MedThreatRAG reduces answer F1 scores by up to 27.66% and lowers LLaVA-Med-1.5 F1 rates to as low as 51.36%. Our findings expose fundamental security gaps in clinical RAG systems and highlight the urgent need for threat-aware design and robust multimodal consistency checks. Finally, we conclude with a concise set of guidelines to inform the safe development of future multimodal medical RAG systems.
  •  

Digital Twins as Funhouse Mirrors: Five Key Distortions

arXiv:2509.19088v4 Announce Type: replace-cross Abstract: Scientists and practitioners are aggressively moving to deploy digital twins - LLM-based models of real individuals - across social science and policy research. We conducted 19 pre-registered studies with 164 diverse outcomes (e.g., attitudes towards hiring algorithms, intention to share misinformation) and compared human responses to those of their digital twins (trained on each person's previous answers to over 500 questions). We find that digital twins' answers are only modestly more accurate than those from the homogeneous base LLM and correlate weakly with human responses (average r = 0.20). We document five ways in which digital twins distort human behavior: (i) stereotyping, (ii) insufficient individuation, (iii) representation bias, (iv) ideological biases, and (v) hyper-rationality. Together, our results caution against the premature deployment of digital twins, which may systematically misrepresent human cognition and undermine both scientific understanding and practical applications.
  •  

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer

npj Digital Medicine, Published online: 31 December 2025; doi:10.1038/s41746-025-02260-3

A clinically validated 3D deep learning approach for quantifying vascular invasion in pancreatic cancer
  •  

Autologous multiantigen-targeted T cell therapy for pancreatic cancer: a phase 1/2 trial

Nature Medicine, Published online: 02 January 2026; doi:10.1038/s41591-025-04043-5

Results of the phase 1/2 TACTOPS trial show that autologous T cell therapy targeting PRAME, SSX2, MAGEA4, Survivin and NY-ESO-1 in patients with pancreatic ductal adenocarcinoma is feasible and safe, and leads to encouraging clinical responses and evidence of antigen spreading in responders.
  •  

Quantifying the global eco-footprint of wearable healthcare electronics

Nature, Published online: 31 December 2025; doi:10.1038/s41586-025-09819-w

An integrated systems engineering framework based on life-cycle inventories is used to quantify the global eco-footprint of wearable healthcare electronics and identify effective mitigation strategies.
  •  

Machine Learning-Based Multi-Omics Integration for Identification of Hepatocellular Carcinoma Biomarkers in an Egyptian Cohort

J Proteome Res. 2025 Dec 30. doi: 10.1021/acs.jproteome.5c00741. Online ahead of print.

ABSTRACT

Hepatocellular carcinoma (HCC) ranks among the most common causes of cancer-related deaths globally. The high incidence of HCC is largely linked to chronic hepatitis virus infections, liver cirrhosis, and exposure to carcinogenic substances. Egypt has one of the world's highest burdens of HCC, with liver cirrhosis from chronic hepatitis C virus (HCV) infection as the primary risk factor. Malignant conversion of cirrhosis to HCC is often fatal in part because adequate biomarkers are not available for diagnosis of HCC in the early stage. Therefore, there is a critical need for more effective biomarkers to detect HCC at an early stage, when therapeutic intervention is more likely to be successful. Multiomics integration has emerged as a powerful strategy to uncover biomarkers and better understand the molecular underpinnings of complex diseases such as HCC. This study summarizes findings from multiple untargeted and targeted mass spectrometry-based analyses of proteins, N-linked glycans, and metabolites performed on blood samples from HCC cases and cirrhotic cohorts recruited in Egypt. Integrative analysis using machine learning methods is performed to identify a panel of multiomics features that differentiates HCC cases from the high-risk population of cirrhotic patients with liver cirrhosis.

PMID:41467861 | DOI:10.1021/acs.jproteome.5c00741

  •  
❌