❌

Normal view

Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study

Background: Gastroesophageal reflux disease (GERD) symptoms significantly affect patients’ quality of life. Patient-reported outcome (PRO) instruments for symptoms measurement in GERD patients is advocated by regulatory authority. Current tools for GERD symptoms evaluation are limited and the results can be biased by the recall bias. To better characterize the GERD symptoms, an e-diary was developed for daily GERD symptom monitoring. Objective: To build up and optimize a PRO-based e-diary, and to investigate the effect of symptom frequency on adherence. Methods: The GERD e-diary evaluated 8 daytime (acid regurgitation, cough, heartburn, sour taste in the mouth, hiccups, hoarseness, dysphagia, and chest pain) and 2 nighttime symptoms (acid regurgitation and cough) for consecutive 8 weeks. The adherence of e-diary, defined as daily completing rate of e-diary, was evaluated and optimized from First Stage to Third Stage with no reminder implemented in First Stage, sending reminding SMS (Short Message Service) text messaging upon detecting missing data in Second Stage (no reminder during the first 3 to 5 days after enrollment), and immediate installation of reminding system at enrollment in Third Stage. GERD symptom frequency was obtained by summation of the symptomatic days in each week. A multiple regression analysis was performed to examine the effects of system optimization and GERD symptom frequency on patient adherence, while controlling for potential confounding variables. Results: 138 GERD patients (M/F=70/68; age: mean 52.9, SD 12.3 years) were recruited. At First Stage, the adherence was 47.2%, 40% and 57.6% for nighttime, daytime and overall symptom. System optimization significantly improved adherence with increased adherence of nighttime symptoms by 12.5% (P=.005) and 10.9% (P=.01), daytime symptom by 21.7% (P

STAT+: FDA announces sweeping changes to oversight of wearables, AI-enabled devices

7 January 2026 at 04:49

LAS VEGAS — The Food and Drug Administration announced Tuesday that it will ease regulation of digital health products, following through on the Trump administration’s promises to deregulate artificial intelligence and promote its widespread use.

FDA Commissioner Marty Makary indicated that one of the agency’s priorities is fostering an environment that’s good for investors, and that FDA regulation needs to move “at Silicon Valley speed.” He announced the changes during an address to conference attendees at the Consumer Electronics Show.

The agency will soften its approach to the regulation of clinical decision support software, which include AI-enabled products that help doctors navigate diagnoses and treatment options. The agency previously considered products that delivered a single recommendation as FDA-regulated medical devices. Now, those products can enter the market without FDA review as long as they fulfill the agency’s other criteria for escaping regulation. 

Continue to STAT+ to read the full story…

© ANDREW CABALLERO-REYNOLDS/AFP via Getty Images

California lawmaker proposes a four-year ban on AI chatbots in kids’ toys

7 January 2026 at 04:22
“Our children cannot be used as lab rats for Big Tech to experiment on,” Senator Steve Padilla said. He just introduced a bill to ban AI chatbots in toys until safety regulations are developed.  

The role of PCMT1 in prognosis tumor immune microenvironment and therapeutic responses across cancers

Discov Oncol. 2026 Jan 5. doi: 10.1007/s12672-025-04366-2. Online ahead of print.

ABSTRACT

BACKGROUND: Emerging evidence highlights the overexpression of Protein-L-isoaspartate (D-aspartate) O-methyltransferase (PCMT1) in multiple malignancies. However, its pan-cancer prognostic significance, tumor immune microenvironment (TIME) interactions, and therapeutic implications remain underexplored.

METHODS: Multi-omics data were integrated from UCSC Xena, GTEx, UALCAN, and published cohorts. PCMT1 expression patterns were systematically analyzed across 33 cancer types. Associations between PCMT1 and clinical outcomes, immune infiltration, immune checkpoint genes (ICGs), tumor mutation burden (TMB), microsatellite instability (MSI), and drug sensitivity were evaluated using bioinformatics pipelines.

RESULTS: Our pan-cancer analysis revealed differential expression patterns of PCMT1 across various malignancies, with significant upregulation in 20 cancer types and downregulation in 3 cancer types. Notably, PCMT1 overexpression was predominantly observed in epithelial-origin tumors, such as ACC (adrenocortical carcinoma), BRCA (breast invasive carcinoma), COAD (colon adenocarcinoma), and LUAD (lung adenocarcinoma). Survival analysis demonstrated that elevated PCMT1 expression was significantly correlated with unfavorable prognosis in multiple epithelial tumors, particularly in BRCA, esophageal carcinoma (ESCA), head and neck squamous cell carcinoma (HNSC), liver hepatocellular carcinoma (LIHC), and mesothelioma (MESO). Furthermore, comprehensive analysis identified significant associations between PCMT1 expression and various tumor microenvironment features, including immune scores, six distinct immune cell types, four immunosuppressive cell populations, cancer-associated fibroblasts (CAFs)-related markers, and immunosuppressive factors. PCMT1 expression also showed significant correlations with tumor mutation burden (TMB), microsatellite instability (MSI), DNA stemness score (DNAss), and RNA stemness score (RNAss). Particularly noteworthy was the strong positive correlation between PCMT1 expression and CAFs infiltration, along with their associated factors. These findings were further validated in independent immunotherapy cohorts, where PCMT1 consistently demonstrated immunosuppressive characteristics.

CONCLUSION: Multi-omics analysis suggests that PCMT1 may serve as a potential prognostic biomarker and a novel immunotherapy target for pan-cancer.

PMID:41491065 | DOI:10.1007/s12672-025-04366-2

PRIME: an interpretable artificial intelligence model based on liquid biopsy improves prediction of progression risk in non-small cell lung cancer

Mil Med Res. 2026 Jan 6;12(1):94. doi: 10.1186/s40779-025-00679-z.

ABSTRACT

BACKGROUND: Despite the predictive impact of circulating tumor DNA (ctDNA) minimal residual disease (MRD), accurate prediction of failure risk after curative-intent treatments for early-stage or localized non-small cell lung cancer (NSCLC) patients to guide personalized therapy remains challenging. This study aimed to develop and validate an interpretable artificial intelligence-assisted model using global data resources.

METHODS: Liquid biopsy data, blood-based genomic alterations, clinicopathological features, and survival outcomes of stage I-III NSCLC patients who underwent surgery or definitive chemoradiotherapy were collected from 6 cohorts. PRIME (Progression Risk prediction by Interpretable Machine learning on ctDNA-MRD, Mutations, and clinical-therapeutic features) was trained by 6 machine learning algorithms across 4 cohorts and validated in 2 independent cohorts. Model performance was evaluated by the area under the curve (AUC) and interpreted by SHapley Additive exPlanations (SHAP). Whole-exome sequencing (WES) or whole-genome sequencing (WGS) of tumor tissue from 430 stage II-III NSCLC patients and RNA-sequencing (RNA-seq) data from 1149 subjects, sourced from The Cancer Genome Atlas, were used to validate the prognostic effect of mutations identified in peripheral blood and investigate the underlying mechanisms.

RESULTS: A global dataset encompassing 781 blood samples from 493 patients was analyzed. Clinical stage, pre-treatment ctDNA, post-treatment MRD, blood-based Kelch-like ECH-associated protein 1 (KEAP1), serine/threonine kinase 11 (STK11), and cyclin-dependent kinase inhibitor 2A (CDKN2A) mutations, and treatment modality were significantly associated with the risk of disease progression and were thereby included in the model training. WES/WGS and RNA-seq confirmed the poor prognostic effect of KEAP1, STK11, and CDKN2A mutations, which were characterized by the suppressive tumor microenvironment and attenuated humoral immunity. The neural network (NN) model exhibited optimal prediction of treatment failure risk in the training (AUC = 0.85, 95% CI 0.81-0.89) and validation sets (AUC = 0.82, 95% CI 0.74-0.89). SHAP analysis indicated that MRD (+0.306), treatment modality (+0.128), and pre-treatment ctDNA (+0.043) ranked in the top 3 contributions. NN-PRIME outperformed single liquid biopsy biomarkers and clinical-therapeutic signatures, and demonstrated consistent robustness across different clinical scenarios. High-risk patients identified by NN-PRIME had poorer prognoses but derived significant benefits from adjuvant therapy after surgery.

CONCLUSIONS: As an interpretable model integrating readily-accessible and crucial clinical-genomic predictors, PRIME achieves enhanced performance, allowing for early outcome prediction, refined risk stratification, and personalized clinical decision-making.

PMID:41491583 | PMC:PMC12771999 | DOI:10.1186/s40779-025-00679-z

Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models

arXiv:2601.01321v1 Announce Type: new Abstract: Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

arXiv:2601.01330v1 Announce Type: new Abstract: Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM routing and aggregation at scale and identify three key bottlenecks: (1) current train-free routers are limited by a query-based paradigm focusing solely on textual similarity; (2) recent aggregation methods remain largely static, failing to select appropriate aggregators for different tasks;(3) the complementarity of routing and aggregation remains underutilized. To address these problems, we introduce JiSi, a novel framework designed to release the full potential of LLMs' collaboration through three innovations: (1) Query-Response Mixed Routing capturing both semantic information and problem difficulty; (2) Support-Set-based Aggregator Selection jointly evaluating the aggregation and domain capacity of aggregators; (3) Adaptive Routing-Aggregation Switch dynamically leveraging the advantages of routing and aggregation. Comprehensive experiments on nine benchmarks demonstrate that JiSi can surpass Gemini-3-Pro with only 47% costs by orchestrating ten open-source LLMs, while outperforming mainstream baselines. It suggests that collective intelligence represents a novel path towards Artificial General Intelligence (AGI).

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

arXiv:2601.01718v1 Announce Type: new Abstract: We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks while maintaining competitive capabilities on general-purpose tasks. To address the overthinking phenomenon commonly observed in Large Reasoning Models (LRMs), we propose Reflection-aware Adaptive Policy Optimization (RAPO), a novel RL training algorithm that effectively regulates overthinking behaviors. In enterprise-oriented tasks such as retrieval-augmented generation (RAG), complex table understanding, and summarization, Yuan3.0 Flash consistently achieves superior performance. Moreover, it also demonstrates strong reasoning capabilities in domains such as mathematics, science, etc., attaining accuracy comparable to frontier model while requiring only approximately 1/4 to 1/2 of the average tokens. Yuan3.0 Flash has been fully open-sourced to facilitate further research and real-world deployment: https://github.com/Yuan-lab-LLM/Yuan3.0.

The Qualitative Laboratory: Theory Prototyping and Hypothesis Generation with Large Language Models

arXiv:2601.00797v1 Announce Type: cross Abstract: A central challenge in social science is to generate rich qualitative hypotheses about how diverse social groups might interpret new information. This article introduces and illustrates a novel methodological approach for this purpose: sociological persona simulation using Large Language Models (LLMs), which we frame as a "qualitative laboratory". We argue that for this specific task, persona simulation offers a distinct advantage over established methods. By generating naturalistic discourse, it overcomes the lack of discursive depth common in vignette surveys, and by operationalizing complex worldviews through natural language, it bypasses the formalization bottleneck of rule-based agent-based models (ABMs). To demonstrate this potential, we present a protocol where personas derived from a sociological theory of climate reception react to policy messages. The simulation produced nuanced and counter-intuitive hypotheses - such as a conservative persona's rejection of a national security frame - that challenge theoretical assumptions. We conclude that this method, used as part of a "simulation then validation" workflow, represents a superior tool for generating deeply textured hypotheses for subsequent empirical testing.

MACA: A Framework for Distilling Trustworthy LLMs into Efficient Retrievers

arXiv:2601.00926v1 Announce Type: cross Abstract: Modern enterprise retrieval systems must handle short, underspecified queries such as ``foreign transaction fee refund'' and ``recent check status''. In these cases, semantic nuance and metadata matter but per-query large language model (LLM) re-ranking and manual labeling are costly. We present Metadata-Aware Cross-Model Alignment (MACA), which distills a calibrated metadata aware LLM re-ranker into a compact student retriever, avoiding online LLM calls. A metadata-aware prompt verifies the teacher's trustworthiness by checking consistency under permutations and robustness to paraphrases, then supplies listwise scores, hard negatives, and calibrated relevance margins. The student trains with MACA's MetaFusion objective, which combines a metadata conditioned ranking loss with a cross model margin loss so it learns to push the correct answer above semantically similar candidates with mismatched topic, sub-topic, or entity. On a proprietary consumer banking FAQ corpus and BankFAQs, the MACA teacher surpasses a MAFA baseline at Accuracy@1 by five points on the proprietary set and three points on BankFAQs. MACA students substantially outperform pretrained encoders; e.g., on the proprietary corpus MiniLM Accuracy@1 improves from 0.23 to 0.48, while keeping inference free of LLM calls and supporting retrieval-augmented generation.

Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code

arXiv:2601.01215v1 Announce Type: cross Abstract: Large language models (LLMs) can generate programs that pass unit tests, but passing tests does not guarantee reliable runtime behavior. We find that different correct solutions to the same task can show very different memory and performance patterns, which can lead to hidden operational risks. We present a framework to measure execution-time memory stability across multiple correct generations. At the solution level, we introduce Dynamic Mean Pairwise Distance (DMPD), which uses Dynamic Time Warping to compare the shapes of memory-usage traces after converting them into Monotonic Peak Profiles (MPPs) to reduce transient noise. Aggregating DMPD across tasks yields a model-level Model Instability Score (MIS). Experiments on BigOBench and CodeContests show substantial runtime divergence among correct solutions. Instability often increases with higher sampling temperature even when pass@1 improves. We also observe correlations between our stability measures and software engineering indicators such as cognitive and cyclomatic complexity, suggesting links between operational behavior and maintainability. Our results support stability-aware selection among passing candidates in CI/CD to reduce operational risk without sacrificing correctness. Artifacts are available.

OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment

arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via semantic search engine; (3) constructing a hierarchical taxonomy of core-task-related work and performing contribution-level full-text comparisons against each contribution; and (4) synthesizing all analyses into a structured novelty report with explicit citations and evidence snippets. Unlike naive LLM-based approaches, \textsc{OpenNovelty} grounds all assessments in retrieved real papers, ensuring verifiable judgments. We deploy our system on 500+ ICLR 2026 submissions with all reports publicly available on our website, and preliminary analysis suggests it can identify relevant prior work, including closely related papers that authors may overlook. OpenNovelty aims to empower the research community with a scalable tool that promotes fair, consistent, and evidence-backed peer review.

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

arXiv:2601.01627v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety of LLMs for Japanese healthcare. Our benchmark is based on 67 guidelines from the Japan Medical Association and contains over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Using a dual-LLM scoring protocol, we evaluate 27 models and find that commercial models maintain robust safety while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.0, $p

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

arXiv:2601.01701v1 Announce Type: cross Abstract: Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently, with the advent of digital twins and data-driven decision-making, several statistical and machine-learning methods have been proposed. However, these methods face several challenges, such as dependence on only real sensor datasets, limited labeled data, high false alarm rates, and privacy concerns. To address these problems, we propose a suite of digital twin-integrated federated learning (DTFL) methods that enhance global model performance while preserving data privacy and communication efficiency. Specifically, we present five novel approaches: Digital Twin-Based Meta-Learning (DTML), Federated Parameter Fusion (FPF), Layer-wise Parameter Exchange (LPE), Cyclic Weight Adaptation (CWA), and Digital Twin Knowledge Distillation (DTKD). Each method introduces a unique mechanism to combine synthetic and real-world knowledge, balancing generalization with communication overhead. We conduct an extensive experiment using a publicly available cyber-physical anomaly detection dataset. For a target accuracy of 80%, CWA reaches the target in 33 rounds, FPF in 41 rounds, LPE in 48 rounds, and DTML in 87 rounds, whereas the standard FedAvg baseline and DTKD do not reach the target within 100 rounds. These results highlight substantial communication-efficiency gains (up to 62% fewer rounds than DTML and 31% fewer than LPE) and demonstrate that integrating DT knowledge into FL accelerates convergence to operationally meaningful accuracy thresholds for IIoT anomaly detection.

Multimodal Fact-Checking: An Agent-based Approach

arXiv:2512.22933v3 Announce Type: replace Abstract: The rapid spread of multimodal misinformation poses a growing challenge for automated fact-checking systems. Existing approaches, including large vision language models (LVLMs) and deep multimodal fusion methods, often fall short due to limited reasoning and shallow evidence utilization. A key bottleneck is the lack of dedicated datasets that provide complete real-world multimodal misinformation instances accompanied by annotated reasoning processes and verifiable evidence. To address this limitation, we introduce RW-Post, a high-quality and explainable dataset for real-world multimodal fact-checking. RW-Post aligns real-world multimodal claims with their original social media posts, preserving the rich contextual information in which the claims are made. In addition, the dataset includes detailed reasoning and explicitly linked evidence, which are derived from human written fact-checking articles via a large language model assisted extraction pipeline, enabling comprehensive verification and explanation. Building upon RW-Post, we propose AgentFact, an agent-based multimodal fact-checking framework designed to emulate the human verification workflow. AgentFact consists of five specialized agents that collaboratively handle key fact-checking subtasks, including strategy planning, high-quality evidence retrieval, visual analysis, reasoning, and explanation generation. These agents are orchestrated through an iterative workflow that alternates between evidence searching and task-aware evidence filtering and reasoning, facilitating strategic decision-making and systematic evidence analysis. Extensive experimental results demonstrate that the synergy between RW-Post and AgentFact substantially improves both the accuracy and interpretability of multimodal fact-checking.

How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System

arXiv:2508.17215v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) augmented with Retrieval-Augmented Generation (RAG) are increasingly employed in medical AI to enhance factual grounding through external clinical image-text retrieval. However, this reliance creates a significant attack surface. We propose MedThreatRAG, a novel multimodal poisoning framework that systematically probes vulnerabilities in medical RAG systems by injecting adversarial image-text pairs. A key innovation of our approach is the construction of a simulated semi-open attack environment, mimicking real-world medical systems that permit periodic knowledge base updates via user or pipeline contributions. Within this setting, we introduce and emphasize Cross-Modal Conflict Injection (CMCI), which embeds subtle semantic contradictions between medical images and their paired reports. These mismatches degrade retrieval and generation by disrupting cross-modal alignment while remaining sufficiently plausible to evade conventional filters. While basic textual and visual attacks are included for completeness, CMCI demonstrates the most severe degradation. Evaluations on IU-Xray and MIMIC-CXR QA tasks show that MedThreatRAG reduces answer F1 scores by up to 27.66% and lowers LLaVA-Med-1.5 F1 rates to as low as 51.36%. Our findings expose fundamental security gaps in clinical RAG systems and highlight the urgent need for threat-aware design and robust multimodal consistency checks. Finally, we conclude with a concise set of guidelines to inform the safe development of future multimodal medical RAG systems.

Digital Twins as Funhouse Mirrors: Five Key Distortions

arXiv:2509.19088v4 Announce Type: replace-cross Abstract: Scientists and practitioners are aggressively moving to deploy digital twins - LLM-based models of real individuals - across social science and policy research. We conducted 19 pre-registered studies with 164 diverse outcomes (e.g., attitudes towards hiring algorithms, intention to share misinformation) and compared human responses to those of their digital twins (trained on each person's previous answers to over 500 questions). We find that digital twins' answers are only modestly more accurate than those from the homogeneous base LLM and correlate weakly with human responses (average r = 0.20). We document five ways in which digital twins distort human behavior: (i) stereotyping, (ii) insufficient individuation, (iii) representation bias, (iv) ideological biases, and (v) hyper-rationality. Together, our results caution against the premature deployment of digital twins, which may systematically misrepresent human cognition and undermine both scientific understanding and practical applications.

Wearable-informed generative digital avatars predict task-conditioned post-stroke locomotion

arXiv:2512.14329v2 Announce Type: replace-cross Abstract: Dynamic prediction of locomotor capacity after stroke could enable more individualized rehabilitation, yet current assessments largely provide static impairment scores and do not indicate whether patients can perform specific tasks such as slope walking or stair climbing. Here, we present a wearable-informed data-physics hybrid generative framework that reconstructs a stroke survivor's locomotor control from wearable inertial sensing and predicts task-conditioned post-stroke locomotion in new environments. From a single 20 m level-ground walking trial recorded by five IMUs, the framework personalizes a physics-based digital avatar using a healthy-motion prior and hybrid imitation learning, generating dynamically feasible, patient-specific movements for inclined walking and stair negotiation. Across 11 stroke inpatients, predicted postures reached 82.2% similarity for slopes and 69.9% for stairs, substantially exceeding a physics-only baseline. In a multicentre pilot randomized study (n = 21; 28 days), access to scenario-specific locomotion predictions to support task selection and difficulty titration was associated with larger gains in Fugl-Meyer lower-extremity scores than standard care (mean change 6.0 vs 3.7 points; $p
❌