❌

Reading view

The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems

arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed because authority and verification capacity do not coincide. We show that this is not a process deviation or technical defect, but a structural property of deployments where decision generation throughput exceeds bounded human verification capacity. We identify a scaling limit under standard deployment assumptions, including parallel agent generation, CI-based validation, and individualized human approval gates. Beyond a throughput threshold, verification ceases to function as a decision criterion and is replaced by ritualized approval based on proxy signals. Personalized responsibility becomes structurally unattainable in this regime. We further characterize a CI amplification dynamic, whereby increasing automated validation coverage raises proxy signal density without restoring human capacity. Under fixed time and attention constraints, this accelerates cognitive offloading in the broad sense and widens the gap between formal approval and epistemic understanding. Additional automation therefore amplifies, rather than mitigates, the responsibility vacuum. We conclude that unless organizations explicitly redesign decision boundaries or reassign responsibility away from individual decisions toward batch- or system-level ownership, responsibility vacuum remains an invisible but persistent failure mode in scaled agent deployments.
  •  

Large-scale spatial profiling of the tumor microenvironment

Multiplex immunofluorescence proteomics is a powerful method in spatial biology to decipher cell types, states, and architecture in tissues. In this issue of Cell, Valanarasu et al. develop GigaTIME, an expansive population-scale analysis of the tumor immune microenvironment of over 14,000 patients across 24 cancer types, enabling clinical discovery and patient stratification.
  •  

Large language models improve transferability of electronic health record-based predictions across countries and coding systems

npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5

Large language models improve transferability of electronic health record-based predictions across countries and coding systems
  •  

Multimodal AI generates virtual population for tumor microenvironment modeling

GigaTIME leverages multimodal AI to generate virtual multiplex immunofluorescence (mIF) profiles from standard H&E slides, enabling comprehensive tumor immune microenvironment modeling across a large (>14,000) and diverse patient population. This virtual approach unlocks new opportunities for large-scale clinical discoveries that were previously hindered by the scarcity of mIF data.
  •  

Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward

arXiv:2601.13122v1 Announce Type: new Abstract: Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vulnerabilities of modern general-purpose AI along eight widely accepted responsible AI (RAI) principles (fairness, privacy, explainability, robustness, safety, truthfulness, governance, and sustainability) and compare how they are non-existent or less severe and easily mitigable in traditional task-specific counterparts. We argue that this is due to the non-deterministically high Degree of Freedom in output (DoFo) of general-purpose AI (unlike the deterministically constant or low DoFo of traditional task-specific AI systems), and there is a need to rethink our approach to RAI for general-purpose AI. Following this, we derive C2V2 (Control, Consistency, Value, Veracity) desiderata to meet the RAI requirements for future general-purpose AI systems, and discuss how recent efforts in AI alignment, retrieval-augmented generation, reasoning enhancements, etc. fare along one or more of the desiderata. We believe that the goal of developing responsible general-purpose AI can be achieved by formally modeling application- or domain-dependent RAI requirements along C2V2 dimensions, and taking a system design approach to suitably combine various techniques to meet the desiderata.
  •  

Medication counseling with large language models: balancing flexibility and rigidity

arXiv:2601.11544v1 Announce Type: cross Abstract: The introduction of large language models (LLMs) has greatly enhanced the capabilities of software agents. Instead of relying on rule-based interactions, agents can now interact in flexible ways akin to humans. However, this flexibility quickly becomes a problem in fields where errors can be disastrous, such as in a pharmacy context, but the opposite also holds true; a system that is too inflexible will also lead to errors, as it can become too rigid to handle situations that are not accounted for. Work using LLMs in a pharmacy context have adopted a wide scope, accounting for many different medications in brief interactions -- our strategy is the opposite: focus on a more narrow and long task. This not only enables a greater understanding of the task at hand, but also provides insight into what challenges are present in an interaction of longer nature. The main challenge, however, remains the same for a narrow and wide system: it needs to strike a balance between adherence to conversational requirements and flexibility. In an effort to strike such a balance, we present a prototype system meant to provide medication counseling while juggling these two extremes. We also cover our design in constructing such a system, with a focus on methods aiming to fulfill conversation requirements, reduce hallucinations and promote high-quality responses. The methods used have the potential to increase the determinism of the system, while simultaneously not removing the dynamic conversational abilities granted by the usage of LLMs. However, a great deal of work remains ahead, and the development of this kind of system needs to involve continuous testing and a human-in-the-loop. It should also be evaluated outside of commonly used benchmarks for LLMs, as these do not adequately capture the complexities of this kind of conversational system.
  •  

Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology

arXiv:2601.11567v1 Announce Type: cross Abstract: Small open-source medical large language models (LLMs) offer promising opportunities for low-resource deployment and broader accessibility. However, their evaluation is often limited to accuracy on medical multiple choice question (MCQ) benchmarks, and lacks evaluation of consistency, robustness, or reasoning behavior. We use MCQ coupled to human evaluation and clinical review to assess six small open-source medical LLMs (HuatuoGPT-o1 (Chen 2024), Diabetica-7B, Diabetica-o1 (Wei 2024), Meditron3-8B (Sallinen2025), MedFound-7B (Liu 2025), and ClinicaGPT-base-zh (Wang 2023)) in pediatric endocrinology. In deterministic settings, we examine the effect of prompt variation on models' output and self-assessment bias. In stochastic settings, we evaluate output variability and investigate the relationship between consistency and correctness. HuatuoGPT-o1-8B achieved the highest performance. The results show that high consistency across the model response is not an indicator of correctness, although HuatuoGPT-o1-8B showed the highest consistency rate. When tasked with selecting correct reasoning, both HuatuoGPT-o1-8B and Diabetica-o1 exhibit self-assessment bias and dependency on the order of the candidate explanations. Expert review of incorrect reasoning rationales identified a mix of clinically acceptable responses and clinical oversight. We further show that system-level perturbations, such as differences in CUDA builds, can yield statistically significant shifts in model output despite stable accuracy. This work demonstrates that small, semantically negligible prompt perturbations lead to divergent outputs, raising concerns about reproducibility of LLM-based evaluations and highlights the output variability under different stochastic regimes, emphasizing the need of a broader diagnostic framework to understand potential pitfalls in real-world clinical decision support scenarios.
  •  

A Cloud-based Multi-Agentic Workflow for Science

arXiv:2601.12607v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex decisions limits their utility. LLM-based agents bridge this gap due to their ability to call external resources and tools and thus are now rapidly gaining popularity. However, coming up with a workflow that can balance the models, cloud providers, and external resources is very challenging, making implementing an agentic system more of a hindrance than a help. In this work, we present a domain-agnostic, model-independent workflow for an agentic framework that can act as a scientific assistant while being run entirely on cloud. Built with a supervisor agent marshaling an array of agents with individual capabilities, our framework brings together straightforward tasks like literature review and data analysis with more complex ones like simulation runs. We describe the framework here in full, including a proof-of-concept system we built to accelerate the study of Catalysts, which is highly important in the field of Chemistry and Material Science. We report the cost to operate and use this framework, including the breakdown of the cost by services use. We also evaluate our system on a custom-curated synthetic benchmark and a popular Chemistry benchmark, and also perform expert validation of the system. The results show that our system is able to route the task to the correct agent 90% of the time and successfully complete the assigned task 97.5% of the time for the synthetic tasks and 91% of the time for real-world tasks, while still achieving better or comparable accuracy to most frontier models, showing that this is a viable framework for other scientific domains to replicate.
  •  

Probabilistic Analysis of Copyright Disputes and Generative AI Safety

arXiv:2410.00475v5 Announce Type: replace-cross Abstract: This paper presents a probabilistic approach to analyzing copyright infringement disputes. Evidentiary principles shaped by case law are formalized in probabilistic terms, and the ``inverse ratio rule'' -- a controversial legal doctrine adopted by some courts -- is examined. Although this rule has faced significant criticism, a formal proof demonstrates its validity, provided it is properly defined. The probabilistic approach is further employed to study the copyright safety of generative AI. Specifically, the Near Access-Free (NAF) condition, previously proposed as a strategy for mitigating the heightened copyright infringement risks of generative AI, is evaluated. The analysis reveals limitations in its justifiability and efficacy.
  •  

Conformal Prediction-Driven Adaptive Sampling for Digital Water Twins

arXiv:2511.05610v2 Announce Type: replace-cross Abstract: Digital Twins (DTs) for Water Distribution Networks (WDNs) require accurate state estimation with limited sensors. Uniform sampling often wastes resources across nodes with different uncertainty. We propose an adaptive framework combining LSTM forecasting and Conformal Prediction (CP) to estimate node-wise uncertainty and focus sensing on the most uncertain points. Marginal CP is used for its low computational cost, suitable for real-time DTs. Experiments on Hanoi, Net3, and CTOWN show 33--34\% lower demand error than uniform sampling at 40\% coverage and maintain 89.4--90.2\% empirical coverage with only 5--10\% extra computation.
  •  

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

arXiv:2601.11505v1 Announce Type: cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D management datasets. Current datasets differ substantially in structure and are time-consuming to access and process, which impedes data integration and reduces the comparability and generalizability of algorithmic developments. This work aims to establish a unified and accessible data resource for T1D algorithm development. Multiple publicly available T1D datasets were consolidated into a unified resource, termed the MetaboNet dataset. Inclusion required the availability of both continuous glucose monitoring (CGM) data and corresponding insulin pump dosing records. Additionally, auxiliary information such as reported carbohydrate intake and physical activity was retained when present. The MetaboNet dataset comprises 3135 subjects and 1228 patient-years of overlapping CGM and insulin data, making it substantially larger than existing standalone benchmark datasets. The resource is distributed as a fully public subset available for immediate download at https://metabo-net.org/ , and with a Data Use Agreement (DUA)-restricted subset accessible through their respective application processes. For the datasets in the latter subset, processing pipelines are provided to automatically convert the data into the standardized MetaboNet format. A consolidated public dataset for T1D research is presented, and the access pathways for both its unrestricted and DUA-governed components are described. The resulting dataset covers a broad range of glycemic profiles and demographics and thus can yield more generalizable algorithmic performance than individual datasets.
  •  

Contaminating plasmid sequences and disrupted vector genomes in the liver following adeno-associated virus gene therapy

Nature Medicine, Published online: 16 January 2026; doi:10.1038/s41591-025-04073-z

Analyses of liver biopsies from a child with spinal muscular atrophy treated with adeno-associated virus gene therapy who developed hepatitis reveal contaminating manufacturing plasmids and disrupted vector genomes, possibly resulting from recombination events.
  •  

Human-AI Co-design for Clinical Prediction Models

arXiv:2601.09072v1 Announce Type: new Abstract: Developing safe, effective, and practically useful clinical prediction models (CPMs) traditionally requires iterative collaboration between clinical experts, data scientists, and informaticists. This process refines the often small but critical details of the model building process, such as which features/patients to include and how clinical categories should be defined. However, this traditional collaboration process is extremely time- and resource-intensive, resulting in only a small fraction of CPMs reaching clinical practice. This challenge intensifies when teams attempt to incorporate unstructured clinical notes, which can contain an enormous number of concepts. To address this challenge, we introduce HACHI, an iterative human-in-the-loop framework that uses AI agents to accelerate the development of fully interpretable CPMs by enabling the exploration of concepts in clinical notes. HACHI alternates between (i) an AI agent rapidly exploring and evaluating candidate concepts in clinical notes and (ii) clinical and domain experts providing feedback to improve the CPM learning process. HACHI defines concepts as simple yes-no questions that are used in linear models, allowing the clinical AI team to transparently review, refine, and validate the CPM learned in each round. In two real-world prediction tasks (acute kidney injury and traumatic brain injury), HACHI outperforms existing approaches, surfaces new clinically relevant concepts not included in commonly-used CPMs, and improves model generalizability across clinical sites and time periods. Furthermore, HACHI reveals the critical role of the clinical AI team, such as directing the AI agent to explore concepts that it had not previously considered, adjusting the granularity of concepts it considers, changing the objective function to better align with the clinical objectives, and identifying issues of data bias and leakage.
  •  

Circulating metabolites, genetics and lifestyle factors in relation to future risk of type 2 diabetes

Nat Med. 2026 Jan 14. doi: 10.1038/s41591-025-04105-8. Online ahead of print.

ABSTRACT

The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.

PMID:41535386 | DOI:10.1038/s41591-025-04105-8

  •  

Multi-omics to study chronic respiratory diseases and viral infections

Eur Respir Rev. 2026 Jan 14;35(179):240286. doi: 10.1183/16000617.0286-2024. Print 2026 Jan.

ABSTRACT

Despite recent advances, the underlying mechanisms of the development and progression of many chronic respiratory diseases remain to be elucidated. Factors such as heterogeneity and complexity of human diseases and difficulty interpreting large datasets hinder research into chronic respiratory diseases. Omics assesses the changes in specific biological entities, such as mRNA expression, epigenetics/epigenomics, genomics, proteomics, metagenomics and metabolomics, and provides valuable insights into the roles of these processes in chronic respiratory diseases. High-throughput omics at bulk, single-cell and spatial levels empower the exploration of disease-related changes through untargeted data-driven statistical methods. Multi-omics is the exploration and integration of multiple biological processes, which compared to a single-omics, can provide a substantially greater and more holistic overview of the pathogenic mechanisms that underpin complex diseases. Multi-omics analysis can comprehensively characterise the mechanisms that drive chronic respiratory diseases, capturing unique biological signatures and cellular interactions at different omics levels. Use of these methods has begun to identify key factors and biomarkers in chronic respiratory diseases. Here, we review current omics approaches and highlight recent advances in respiratory research achieved using multi-omics and integrative methods. Our review provides a valuable resource for researchers and clinicians in this area.

PMID:41534886 | DOI:10.1183/16000617.0286-2024

  •  

Complement-secreting CAFs are associated with better prognosis in pancreatic cancer: single-cell multiomics

Gut. 2026 Jan 13:gutjnl-2025-335683. doi: 10.1136/gutjnl-2025-335683. Online ahead of print.

ABSTRACT

BACKGROUND: Accumulating evidence has demonstrated that distinct tumour-promoting and tumour-restraining cancer-associated fibroblast (CAF) subtypes coexist in pancreatic ductal adenocarcinoma.

OBJECTIVE: To develop targeted CAF therapeutic strategies by reprogramming tumour-promoting CAF subtypes.

DESIGN: We leveraged multiomics technologies to systematically identify and characterise CAF subtypes transcriptionally, epigenetically and spatially and correlate them with clinicopathological features.

RESULTS: We found that complement-secreting CAFs (csCAFs), initially identified by our group and inflammatory CAFs (iCAFs) share significant overlap in their transcriptional profiles and chromatin accessibility. iCAFs specifically express transcription factors from the heme and oxidative homeostasis pathway and the activator protein 1 family, which are both involved in cellular response to oxidative stress. Notably, the composition of csCAFs among all CAFs declined during pancreatic carcinogenesis, while trajectory analysis showed that csCAFs could potentially differentiate into iCAFs. Spatially resolved analysis indicated that tumour regions with a higher csCAF composition were associated with lower levels of TGF-β ligands, fewer M2 tumour-associated macrophages and increased levels of lipid mediators. Additionally, we identified a spatially defined CXCL12-CXCR4 ligand-receptor interaction between csCAFs and T cells, but in distinct patterns between different metastatic organs. Patients with a higher composition of csCAFs have significantly longer overall survival and recurrence-free survival through multiplex immunohistochemistry and bulk RNA-seq deconvolution.

CONCLUSION: Our study demonstrates that csCAFs may represent an early-stage iCAF subtype and suggests a promising strategy for reprogramming iCAFs into csCAFs.

PMID:41534892 | DOI:10.1136/gutjnl-2025-335683

  •  

<em>Helicobacter pylori</em> and Cancer: What's the Link?

Clin Exp Gastroenterol. 2026 Jan 7;19:1-11. doi: 10.2147/CEG.S495588. eCollection 2026.

ABSTRACT

Helicobacter pylori (H. pylori) is a human bacterial pathogen that causes one of the most common chronic bacterial infections worldwide. The microorganism has been classified by the International Agency for Research on Cancer as a Group I carcinogen. While the etiological link to gastric cancer is well established, the precise molecular and cellular mechanisms driving this transformation are highly complex and incompletely understood. Fundamentally, the infection results from the chronic presence of acute on chronic gastric mucosal inflammation. H. pylori pathogenicity is increased by bacterial virulence factors including the cytotoxin-associated gene A (CagA) and Vacuolating cytotoxin A (VacA) which may interfere with the host's cell communication and create a pro-tumorigenic microenvironment. Host microRNAs (miRNAs) may amplify these effects by modulating immune responses, enhancing oncogenic signalling. Despite the proven benefits of H. pylori eradication in reducing cancer risk, especially in high-incidence regions, rising antibiotic resistance and host-related variables impede its global implementation. Recent advances in genomics and multi-omics profiling potentially offer new opportunities for targeted prevention. Moreover, emerging evidence suggests H. pylori may also negatively influence immunotherapy outcomes, underscoring its broader relevance in cancer treatment planning. By synthesizing molecular insights, epidemiological trends, and clinical data, this narrative review examines the multifaceted pathways through which H. pylori contributes to gastric carcinogenesis, integrating current knowledge on microbial virulence, host signalling disruption, immune modulation, and epigenetic remodelling.

PMID:41531650 | PMC:PMC12791163 | DOI:10.2147/CEG.S495588

  •  
❌