❌

Reading view

Trustworthy Blockchain-based Federated Learning for Electronic Health Records: Securing Participant Identity with Decentralized Identifiers and Verifiable Credentials

arXiv:2602.02629v1 Announce Type: cross Abstract: The digitization of healthcare has generated massive volumes of Electronic Health Records (EHRs), offering unprecedented opportunities for training Artificial Intelligence (AI) models. However, stringent privacy regulations such as GDPR and HIPAA have created data silos that prevent centralized training. Federated Learning (FL) has emerged as a promising solution that enables collaborative model training without sharing raw patient data. Despite its potential, FL remains vulnerable to poisoning and Sybil attacks, in which malicious participants corrupt the global model or infiltrate the network using fake identities. While recent approaches integrate Blockchain technology for auditability, they predominantly rely on probabilistic reputation systems rather than robust cryptographic identity verification. This paper proposes a Trustworthy Blockchain-based Federated Learning (TBFL) framework integrating Self-Sovereign Identity (SSI) standards. By leveraging Decentralized Identifiers (DIDs) and Verifiable Credentials (VCs), our architecture ensures only authenticated healthcare entities contribute to the global model. Through comprehensive evaluation using the MIMIC-IV dataset, we demonstrate that anchoring trust in cryptographic identity verification rather than behavioral patterns significantly mitigates security risks while maintaining clinical utility. Our results show the framework successfully neutralizes 100% of Sybil attacks, achieves robust predictive performance (AUC = 0.954, Recall = 0.890), and introduces negligible computational overhead (
  •  

Digital intervention <i>mylovia</i> improves sexual functioning in women with sexual dysfunction in randomized controlled trial

npj Digital Medicine, Published online: 03 February 2026; doi:10.1038/s41746-026-02385-z

Digital intervention mylovia improves sexual functioning in women with sexual dysfunction in randomized controlled trial
  •  

Barriers to Digital Health Adoption in Older Adults: Scoping Review Informed by Innovation Resistance Theory

Background: The transformation of digital health technologies has reshaped how healthcare is delivered, particularly in primary care. However, despite the advantages of these innovations, older adults remain among the most resistant users. Traditional technology adoption models may not fully capture the complexity of this reluctance, which is shaped not only by usability challenges but also by emotional, psychological, and identity-related concerns. Innovation Resistance Theory (IRT) offers a complementary framework focused on understanding barriers to adoption rather than solely on facilitators. Objective: To map and synthesize evidence on older adults’ resistance to digital health technologies in primary care through the lens of IRT, and to examine how empirically observed resistance factors align with, extend, or refine IRT’s functional and psychological barriers. Methods: A scoping review combined with concept-driven thematic synthesis was conducted. Empirical studies published between 2014 and 2025 were identified through systematic searches across five databases: PubMed, CINAHL, Ovid Medline, Web of Science, and Scopus. Inclusion criteria focused on studies examining barriers or resistance to digital health use among older adults aged 60 and above in primary care settings. The search was guided by terms related to “older adults”, “digital health”, “eHealth”, “telemedicine”, and “technology resistance”. After screening and reviewing the full texts, data were extracted into a structured matrix, and findings were organized according to the five dimensions of the IRT: usage, value, risk, tradition, and image barriers. Results: Of 4,976 identified records, seventeen studies met the inclusion criteria. Functional barriers included usability challenges, interface complexity, and age-related impairments. Psychological resistance was frequently linked to emotional discomfort, symbolic misalignment, and concerns about the loss of relational care. Value and risk concerns included distrust in diagnostics accuracy, concerns regarding privacy and data security, and skepticism about care quality. Traditional preferences for face-to-face interactions and generational digital divides further reinforced image-based resistance. A key finding was the interaction between barriers, where low self-efficacy and technology anxiety create feedback loops that reinforce avoidance behaviors. Conclusions: Resistance to digital health among older adults is not simply a lack of adoption but a complex, emotionally grounded process involving functional, psychological, and identity-based barriers. Interventions must go beyond technical usability to rebuild emotional trust, preserve the relational aspects of care, and align digital solutions with the values and expectations of older adults. Innovation Resistance Theory offers a comprehensive framework for understanding these multifaceted dynamics and serves as a valuable guide for policy development, user-centered design, and future research Clinical Trial: None
  •  

Integrative proteogenomics maps multifactorial aetiology, progression and therapeutic vulnerabilities in gastric cancer

Gut. 2026 Jan 30:gutjnl-2025-337247. doi: 10.1136/gutjnl-2025-337247. Online ahead of print.

ABSTRACT

BACKGROUND: Gastric cancer, with disproportionately higher incidence in East Asia, arises from complex host-microbiome-environment interactions beyond Helicobacter pylori (HP) infection. However, the molecular architecture linking environmental carcinogens, microbial succession and host response remains unclear.

OBJECTIVE: To delineate multifactorial aetiologies and clinically actionable subtypes/biomarkers of gastric cancer through integrative proteogenomic, microbial and environmental exposure profiling.

DESIGN: We established a multiomics atlas of paired tumour, adjacent mucosa tissues and blood from 154 treatment-naïve Taiwanese patients, integrating whole-exome sequencing, RNA-seq, proteome and phosphoproteome profiling with carcinogen signatures, HP status, microbiome composition and refined anatomical mapping. Cell-based functional assays tested carcinogen effects. Microbial subtype was assessed in an independent cohort.

RESULTS: A polycyclic-aromatic-hydrocarbon signature, dibenz[a,h]acridine, emerged as a high-risk exposure promoting invasion, immune suppression and poor survival, significantly exceeding nitrosamine-linked risk in this cohort. Multilayer integration defined three initiation ecologies: HP-driven inflammatory, non-HP microbiome-enriched immune-silent and HP-free microbially depleted states. Among HP-negative tumours, a Streptococcus-enriched subtype associated with tight-junction (CLDN18.2/ZO-1/OCLN) disruption and epithelial-mesenchymal transition, whereas a subset of clinically aggressive cases retained CLDN18.2-high epithelial-stable subtype for therapeutic accessibility. An independent cohort revealed gastric juice-derived Streptococcus anginosus abundance inversely correlated with tight-junction proteins. Anatomical mapping reveals location-specific, sex-specific, subtype-specific oncogenic networks and kinase activity, including CDK4 activation in clinical biomarker-negative tumours. Decision-tree models combining exposure and proteome-immune states refined recurrence and survival prediction beyond stage.

CONCLUSION: This proteogenomic framework defines exposure-informed and microbiome-informed gastric cancer subtypes, providing a molecular schema for patient stratification, prevention and actionable therapeutic vulnerabilities.

PMID:41617485 | DOI:10.1136/gutjnl-2025-337247

  •  

Liquid biopsy biomarkers for accurate detection of malignant pulmonary nodules: a meta-analytic approach

Discov Oncol. 2026 Jan 29;17(1):178. doi: 10.1007/s12672-025-03646-1.

ABSTRACT

Pulmonary nodules are a common radiological finding that can be classified as either benign or Malignant, with significant clinical implications. The early detection of malignant nodules is critically important for improving the prognosis of lung cancer, which remains the leading cause of cancer-related mortality worldwide. Traditional imaging techniques have Limitations in accurately classifying pulmonary nodules. Liquid biopsy, a minimally invasive method that evaluates circulating components in the Blood, presents promising diagnostic potential in this context. This study aims to evaluate the diagnostic capacity of multiple liquid biopsy biomarkers for early and accurate differentiation between benign and Malignant pulmonary nodules. Accordingly, we conducted a comprehensive study involving a meta-analysis, selecting 16 eligible studies that utilised liquid biopsy to assess various circulating biomarkers in the diagnostic yield. The most significant results were linked to circulating free DNA (cfDNA). However, other components, including circulating tumour cells (CTCs), microRNAs/pfeRNAs, extracellular vesicles (EVs), serological markers, and imaging techniques, also provided valuable information. Similarly, integrating multi-omics data with machine learning models has been shown to enhance the ability to differentiate between benign and malignant pulmonary nodules, thereby supporting early diagnosis and improved management for patients with lung cancer.

PMID:41612093 | PMC:PMC12855667 | DOI:10.1007/s12672-025-03646-1

  •  

Products, Performance, and Technological Development of Ambulatory Oxygen Therapy Devices: Scoping Review

Background: Ambulatory oxygen therapy is prescribed for patients with chronic lung diseases who experience exertional hypoxemia. However, available devices may not adequately meet user requirements, and their performance characteristics are heterogeneous. Objective: This study aims to identify devices available for delivery of ambulatory oxygen therapy, the technologies that they use to generate oxygen, the performance characteristics of each device, and the development status. Methods: We used medical and engineering databases to identify peer-reviewed papers (eg, MEDLINE, IEEE). Gray literature was used to identify additional descriptions of ambulatory oxygen devices in military medicine, space exploration, or patents. The last search was conducted in September 2025. Documents that described a device that can deliver oxygen in an ambulatory context (defined as weighing less than 10 kg) and were written in English were included. Search results were screened for inclusion by 2 independent reviewers. Data were synthesized by descriptively mapping the performance of each product, the technology used, and the development status of emerging technologies. Results: From 9702 records identified, a total of 166 met eligibility criteria (106 scientific publications and 60 gray literature). We identified 33 portable oxygen concentrators (POCs; 29 commercially available), 10 oxygen cylinders, and 6 portable liquid oxygen (LOX) devices. The POC products showed a trade-off between portability and oxygen delivery capacity (maximum flow rate ranging from 2.0 to 6.0 L/min; device weight ranging from 1.0 to 9.1 kg). Pressure swing adsorption with zeolite was the most common oxygen generation technology in POCs on the market. The mean maximum continuous operating time of POCs was 3.8 hours. Two prototype POCs (maximum flow rate of 4-6 L/min and device weight of 8-9 kg) were developed for space exploration using modified adsorbents. LOX devices were the lightest and had the longest continuous operating time. Innovations in delivery included the downsizing of a POC by using nanozeolite as an adsorbent and pulse oximeter oxygen saturation (SpO2)–targeted automatic titration of oxygen delivery based on the user’s SpO2. Conclusions: This scoping review is the first study to integrate medical, engineering, and gray literature on ambulatory oxygen devices and their development. Although prior literature has narratively explained the products and technologies, no previous research has systematically investigated them. This review showed that POCs available to consumers may not meet the needs of patients in terms of flow rate, portability, and operating time. LOX devices offered superior performance but are limited by high costs. Limitations of this review include the difficulty of comparing product performance across oxygen delivery settings and that the records were largely obtained from English-language sources. Innovation in ambulatory oxygen technology has been limited over the past decade, highlighting urgent need for research and development of new lightweight devices with higher oxygen delivery. Clinical Trial: OSF Registries 10.17605/OSF.IO/QS7FX; https://osf.io/qs7fx
  •  

OpenAI and Anthropic Introduce Healthcare-Focused AI Platforms

OpenAI and Anthropic have announced new healthcare-oriented AI offerings that extend their models beyond general conversational use and into regulated clinical and life sciences environments. Both releases emphasize technical integration, interoperability, and governance, reflecting a shift toward AI systems designed to operate directly within existing healthcare infrastructure.

By Robert Krzaczyński
  •  

High-Fidelity Longitudinal Patient Simulation Using Real-World Data

arXiv:2601.17310v1 Announce Type: new Abstract: Simulation is a powerful tool for exploring uncertainty. Its potential in clinical medicine is transformative and includes personalized treatment planning and virtual clinical trials. However, simulating patient trajectories is challenging because of complex biological and sociocultural influences. Here, we show that real-world clinical records can be leveraged to empirically model patient timelines. We developed a generative simulator model that takes a patient's history as input and synthesizes fine-grained, realistic future trajectories. The model was pretrained on more than 200 million clinical records. It produced high-fidelity future timelines, closely matching event occurrence rates, laboratory test results, and temporal dynamics in real patient future data. It also accurately estimated future event probabilities, with observed-to-expected ratios consistently near 1.0 across diverse outcomes and time horizons. Our results reveal the untapped value of real-world data in electronic health records and introduce a scalable framework for in silico modeling of clinical care.
  •  

Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach

arXiv:2601.17303v1 Announce Type: cross Abstract: As Industrial Internet of Things (IIoT) environments expand to include tens of thousands of connected devices. The centralization of security monitoring architectures creates serious latency issues that savvy attackers can exploit to compromise an entire manufacturing ecosystem. This paper outlines a new, decentralized multi-agent swarm (DMAS) architecture that includes autonomous artificial intelligence (AI) agents at each edge gateway, functioning as a distributed digital "immune system" for IIoT networks. Instead of using a traditional static firewall approach, the DMAS agents communicate via a lightweight peer-to-peer protocol to cooperatively detect anomalous behavior across the IIoT network without sending data to a cloud infrastructure. The authors also outline a consensus-based threat validation (CVT) process in which agents vote on the threat level of an identified threat, enabling instant quarantine of a compromised node or nodes. The authors conducted experiments on a testbed that simulated an innovative factory environment with 2000 IIoT devices and found that the DMAS demonstrated sub-millisecond response times (average of 0.85ms), 97.3% accuracy in detecting malicious activity under high load, and 87% accuracy in detecting zero-day attacks. All significantly higher than baseline values for both centralized and edge computing. Additionally, the proposed architecture can prevent real-time cascading failures in industrial control systems and reduce network bandwidth use by 89% compared to cloud-based solutions.
  •  

Coronary Artery Segmentation and Vessel-Type Classification in X-Ray Angiography

arXiv:2601.17429v1 Announce Type: cross Abstract: X-ray coronary angiography (XCA) is the clinical reference standard for assessing coronary artery disease, yet quantitative analysis is limited by the difficulty of robust vessel segmentation in routine data. Low contrast, motion, foreshortening, overlap, and catheter confounding degrade segmentation and contribute to domain shift across centers. Reliable segmentation, together with vessel-type labeling, enables vessel-specific coronary analytics and downstream measurements that depend on anatomical localization. From 670 cine sequences (407 subjects), we select a best frame near peak opacification using a low-intensity histogram criterion and apply joint super-resolution and enhancement. We benchmark classical Meijering, Frangi, and Sato vesselness filters under per-image oracle tuning, a single global mean setting, and per-image parameter prediction via Support Vector Regression (SVR). Neural baselines include U-Net, FPN, and a Swin Transformer, trained with coronary-only and merged coronary+catheter supervision. A second stage assigns vessel identity (LAD, LCX, RCA). External evaluation uses the public DCA1 cohort. SVR per-image tuning improves Dice over global means for all classical filters (e.g., Frangi: 0.759 vs. 0.741). Among deep models, FPN attains 0.914+/-0.007 Dice (coronary-only), and merged coronary+catheter labels further improve to 0.931+/-0.006. On DCA1 as a strict external test, Dice drops to 0.798 (coronary-only) and 0.814 (merged), while light in-domain fine-tuning recovers to 0.881+/-0.014 and 0.882+/-0.015. Vessel-type labeling achieves 98.5% accuracy (Dice 0.844) for RCA, 95.4% (0.786) for LAD, and 96.2% (0.794) for LCX. Learned per-image tuning strengthens classical pipelines, while high-resolution FPN models and merged-label supervision improve stability and external transfer with modest adaptation.
  •  

"Rebuilding" Statistics in the Age of AI: A Town Hall Discussion on Culture, Infrastructure, and Training

arXiv:2601.17510v1 Announce Type: cross Abstract: This article presents the full, original record of the 2024 Joint Statistical Meetings (JSM) town hall, "Statistics in the Age of AI," which convened leading statisticians to discuss how the field is evolving in response to advances in artificial intelligence, foundation models, large-scale empirical modeling, and data-intensive infrastructures. The town hall was structured around open panel discussion and extensive audience Q&A, with the aim of eliciting candid, experience-driven perspectives rather than formal presentations or prepared statements. This document preserves the extended exchanges among panelists and audience members, with minimal editorial intervention, and organizes the conversation around five recurring questions concerning disciplinary culture and practices, data curation and "data work," engagement with modern empirical modeling, training for large-scale AI applications, and partnerships with key AI stakeholders. By providing an archival record of this discussion, the preprint aims to support transparency, community reflection, and ongoing dialogue about the evolving role of statistics in the data- and AI-centric future.
  •  

GenAI-Net: A Generative AI Framework for Automated Biomolecular Network Design

arXiv:2601.17582v1 Announce Type: cross Abstract: Biomolecular networks underpin emerging technologies in synthetic biology-from robust biomanufacturing and metabolic engineering to smart therapeutics and cell-based diagnostics-and also provide a mechanistic language for understanding complex dynamics in natural and ecological systems. Yet designing chemical reaction networks (CRNs) that implement a desired dynamical function remains largely manual: while a proposed network can be checked by simulation, the reverse problem of discovering a network from a behavioral specification is difficult, requiring substantial human insight to navigate a vast space of topologies and kinetic parameters with nonlinear and possibly stochastic dynamics. Here we introduce GenAI-Net, a generative AI framework that automates CRN design by coupling an agent that proposes reactions to simulation-based evaluation defined by a user-specified objective. GenAI-Net efficiently produces novel, topologically diverse solutions across multiple design tasks, including dose responses, complex logic gates, classifiers, oscillators, and robust perfect adaptation in deterministic and stochastic settings (including noise reduction). By turning specifications into families of circuit candidates and reusable motifs, GenAI-Net provides a general route to programmable biomolecular circuit design and accelerates the translation from desired function to implementable mechanisms.
  •  

The Limits of AI Data Transparency Policy: Three Disclosure Fallacies

arXiv:2601.18127v1 Announce Type: cross Abstract: Data transparency has emerged as a rallying cry for addressing concerns about AI: data quality, privacy, and copyright chief among them. Yet while these calls are crucial for accountability, current transparency policies often fall short of their intended aims. Similar to nutrition facts for food, policies aimed at nutrition facts for AI currently suffer from a limited consideration of research on effective disclosures. We offer an institutional perspective and identify three common fallacies in policy implementations of data disclosures for AI. First, many data transparency proposals exhibit a specification gap between the stated goals of data transparency and the actual disclosures necessary to achieve such goals. Second, reform attempts exhibit an enforcement gap between required disclosures on paper and enforcement to ensure compliance in fact. Third, policy proposals manifest an impact gap between disclosed information and meaningful changes in developer practices and public understanding. Informed by the social science on transparency, our analysis identifies affirmative paths for transparency that are effective rather than merely symbolic.
  •  

Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts

arXiv:2508.19008v2 Announce Type: replace Abstract: Building on a human-led thematic analysis of life-story interviews with inpatients with Borderline Personality Disorder, this study examines the capacity of large language models (OpenAI's GPT, Google's Gemini, and Anthropic's Claude) to support qualitative clinical analysis. The models were evaluated through a mixed procedure. Study A involved blinded and non-blinded expert judges in phenomenology and clinical psychology. Assessments included semantic congruence, Jaccard coefficients for overlap of outputs, multidimensional validity ratings of credibility, coherence, and the substantiveness of results, and their grounding in qualitative data. In Study B, neural methods were used to embed the theme descriptions created by humans and the models in a two-dimensional vector space to provide a computational measure of the difference between human and model semantics and linguistic style. In Study C, complementary non-expert evaluations were conducted to examine the influence of thematic verbosity on the perception of human authorship and content validity. Results of all three studies revealed variable overlap with the human analysis, with models being partly indistinguishable from, and also identifying themes originally omitted by, human researchers. The findings highlight both the variability and potential of AI-augmented thematic qualitative analysis to mitigate human interpretative bias and enhance sensitivity.
  •  

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

arXiv:2409.07314v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To bridge the gap between theoretical capability and verified utility, we introduce MEDIC, a comprehensive evaluation framework establishing leading indicators across various clinical dimensions. Beyond standard question-answering, we assess operational capabilities using deterministic execution protocols and a novel Cross-Examination Framework (CEF), which quantifies information fidelity and hallucination rates without reliance on reference texts. Our evaluation across a heterogeneous task suite exposes critical performance trade-offs: we identify a significant knowledge-execution gap, where proficiency in static retrieval does not predict success in operational tasks such as clinical calculation or SQL generation. Furthermore, we observe a divergence between passive safety (refusal) and active safety (error detection), revealing that models fine-tuned for high refusal rates often fail to reliably audit clinical documentation for factual accuracy. These findings demonstrate that no single architecture dominates across all dimensions, highlighting the necessity of a portfolio approach to clinical model deployment. As part of this investigation, we released a public leaderboard on Hugging Face.\footnote{https://huggingface.co/spaces/m42-health/MEDIC-Benchmark}
  •  

Multimodal Cancer Modeling in the Age of Foundation Model Embeddings

arXiv:2505.07683v4 Announce Type: replace-cross Abstract: The Cancer Genome Atlas (TCGA) has enabled novel discoveries and served as a large-scale reference dataset in cancer through its harmonized genomics, clinical, and imaging data. Numerous prior studies have developed bespoke deep learning models over TCGA for tasks such as cancer survival prediction. A modern paradigm in biomedical deep learning is the development of foundation models (FMs) to derive feature embeddings agnostic to a specific modeling task. Biomedical text especially has seen growing development of FMs. While TCGA contains free-text data as pathology reports, these have been historically underutilized. Here, we investigate the ability to train classical machine learning models over multimodal, zero-shot FM embeddings of cancer data. We demonstrate the ease and additive effect of multimodal fusion, outperforming unimodal models. Further, we show the benefit of including pathology report text and rigorously evaluate the effect of model-based text summarization and hallucination. Overall, we propose an embedding-centric approach to multimodal cancer modeling.
  •  

Feasibility, Acceptability, and Perspectives Regarding the Use of Activity Tracking Wearable Devices Among Home Health Aides: Mixed Methods Study

Background: Home health aides and attendants (HHAs) provide in-home care to the growing population of older adults who want to age in place. Despite their vital role in patient care, HHAs are an underserved and vulnerable population of health care professionals who often experience poor health themselves. Activity tracking devices offer a promising way to improve HHAs’ health-related awareness and promote health behavior change, particularly regarding physical activity and sleep quality, 2 areas in which the workforce struggles. Objective: This study aimed to understand how feasible it is for HHAs to use activity tracking devices and assess their perceptions of such devices for improving their health. Specifically, we conducted (1) a field study to assess the use, feasibility, and acceptability of these devices among HHAs and (2) a qualitative study to understand HHAs’ perspectives on and reactions to activity trackers on and off the job. Methods: We partnered with the 1199 Service Employees International Union Training and Employment Fund to conduct a field study with home care agency–employed HHAs working in New York City, New York. Participants wore activity tracking devices for 4 weeks that collected data on physical activity and sleep. The HHAs were subsequently interviewed on their experiences with and attitudes toward the devices and asked to reflect on personalized visualizations of their data to prompt them to think aloud. Quantitative data were analyzed using descriptive statistics. Qualitative data were analyzed using grounded theory. Results: A total of 17 HHAs participated; their mean age was 48.7 (SD 12.2) years, 15 (88%) were women, 11 (65%) identified as Black, 5 (29%) identified as Hispanic or Latinx, and they had worked as HHAs for a mean of 11.7 (SD 7.5) years. In total, 94% (n=16) of the HHAs wore their activity trackers for the full 28-day study period. Participants took a mean of 10,230 (SD 3586) daily steps during the study period and slept for a mean of 6.27 (SD 0.58) hours per night. Overall, 4 key themes emerged: (1) activity tracking devices enhanced participants’ health awareness by providing empirical data for self-reflection; (2) this increased awareness led to positive behavior changes, including setting and achieving health-related goals; (3) HHAs believed that these devices could improve not only their own health but also that of their patients through positive behavior changes; and (4) despite this optimism, participants emphasized that their ability to modify sleep and activity patterns was constrained by social and occupational determinants, with sleep improvements being particularly challenging. Conclusions: Our findings suggest that appropriately designed personal tracking interventions could offer a promising approach to supporting positive health-related changes in this historically overlooked workforce, potentially improving their well-being and the quality of care they provide to their patients.
  •  

PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning

arXiv:2601.16414v1 Announce Type: cross Abstract: Difficulty replicating baselines, high computational costs, and required domain expertise create persistent barriers to clinical AI research. To address these challenges, we introduce PyHealth 2.0, an enhanced clinical deep learning toolkit that enables predictive modeling in as few as 7 lines of code. PyHealth 2.0 offers three key contributions: (1) a comprehensive toolkit addressing reproducibility and compatibility challenges by unifying 15+ datasets, 20+ clinical tasks, 25+ models, 5+ interpretability methods, and uncertainty quantification including conformal prediction within a single framework that supports diverse clinical data modalities - signals, imaging, and electronic health records - with translation of 5+ medical coding standards; (2) accessibility-focused design accommodating multimodal data and diverse computational resources with up to 39x faster processing and 20x lower memory usage, enabling work from 16GB laptops to production systems; and (3) an active open-source community of 400+ members lowering domain expertise barriers through extensive documentation, reproducible research contributions, and collaborations with academic health systems and industry partners, including multi-language support via RHealth. PyHealth 2.0 establishes an open-source foundation and community advancing accessible, reproducible healthcare AI. Available at pip install pyhealth.
  •  

The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems

arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed because authority and verification capacity do not coincide. We show that this is not a process deviation or technical defect, but a structural property of deployments where decision generation throughput exceeds bounded human verification capacity. We identify a scaling limit under standard deployment assumptions, including parallel agent generation, CI-based validation, and individualized human approval gates. Beyond a throughput threshold, verification ceases to function as a decision criterion and is replaced by ritualized approval based on proxy signals. Personalized responsibility becomes structurally unattainable in this regime. We further characterize a CI amplification dynamic, whereby increasing automated validation coverage raises proxy signal density without restoring human capacity. Under fixed time and attention constraints, this accelerates cognitive offloading in the broad sense and widens the gap between formal approval and epistemic understanding. Additional automation therefore amplifies, rather than mitigates, the responsibility vacuum. We conclude that unless organizations explicitly redesign decision boundaries or reassign responsibility away from individual decisions toward batch- or system-level ownership, responsibility vacuum remains an invisible but persistent failure mode in scaled agent deployments.
  •  
❌