❌

Reading view

Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study

Background: Embedding models are critical components of Retrieval Augmented Generation (RAG) systems for retrieving and searching unstructured medical data. However, existing models are predominantly trained on publicly available English datasets, limiting their effectiveness in non-English health care settings. More importantly, these models lack training on real-world clinical documents, leading to inaccurate context retrieval when integrated into RAG systems for health care applications. This gap is particularly pronounced in specialized medical documentation containing domain-specific terminology, abbreviations, and nuanced clinical language. Objective: This retrospective study aimed to develop and validate embedding models specifically trained on real-world clinical documents from multiple medical specialties to improve medical information retrieval (IR) and RAG system performance in both German and English language contexts. Methods: We fine-tuned embedding models, so-called sentence transformers, using the multilingual-e5-large architecture as a foundation. Training data consisted of approximately 11 million question-answer pairs synthetically generated from 400,000 diverse clinical documents from a large German tertiary hospital, spanning 163,840 patients and 282,728 clinical cases between 2018 and 2023. The large language model generated medically relevant questions and corresponding answers for each document. The dataset was additionally pseudonymized and translated into English to aim for broader applicability. Models were evaluated in 2 distinct scenarios: IR using questions with multiple relevant passages, and RAG system performance in both cross-patient and patient-centered contexts. Results: In the IR evaluation, the fine-tuned miracle model achieved a mAP@100 of 0.27, outperforming the multilingual-e5-large baseline (0.14) and state-of-the-art models such as bge-m3 (0.11). In the RAG evaluation, the model demonstrated robust performance comparable with the baseline in the constrained patient-centered scenario (BERTScore F1 0.781 vs 0.778) and showed moderate improvements in the unconstrained cross-patient setting (BLEURT 0.56 vs 0.53). Notably, the model trained on pseudonymized data achieved comparable retrieval performance (mAP@100 0.25) and the highest scores for patient-centered contextual precision (0.93). Performance gains were robust in the German dataset, while the translated English model demonstrated promising results as a proof of concept for cross-lingual transfer. Conclusions: By leveraging a comprehensive real-world dataset spanning multiple medical specialties and using large language models for synthetic question generation, we successfully created and validated domain-specific embedding models. These models can improve medical IR in large-scale search spaces and perform competitively in constrained RAG applications. By publishing the models trained on pseudonymized data, other health care institutions can integrate or adapt these embedding models to their needs. This work establishes a reproducible framework for developing domain-specific clinical embedding models, with the potential to improve data retrieval in medical settings.
  •  

Systems Biology and Multi-Omics in Asthma and COPD: A Systematic Review of Computational Approaches (2010-2024)

J Asthma Allergy. 2026 Mar 19;19:575312. doi: 10.2147/JAA.S575312. eCollection 2026.

ABSTRACT

Systems biology approaches have contributed to advancing our understanding of complex respiratory diseases including asthma and chronic obstructive pulmonary disease (COPD). This systematic review evaluates the application of systems biology methodologies in respiratory medicine, focusing on multi-omics data integration and computational techniques for biomarker discovery and mechanistic understanding. Following PRISMA 2020 guidelines, we conducted a comprehensive literature search across Web of Science and Scopus databases, identifying 117 peer-reviewed documents published from 2010 to 2024. The review methodology employed bibliometric analysis combined with qualitative synthesis of included studies. Results demonstrate steady growth in systems biology applications for asthma and COPD research, with publication rates increasing by approximately 0.5 articles per year (R2 = 0.73, p < 0.001). Bibliometric analysis identified five major research clusters: systems biology as a foundational methodological framework (Basic Theme), COPD-focused research as the most developed area (Motor Theme), gene expression analysis, disease classification approaches, and specialized lung disease investigations (Niche Theme). Multi-omics integration studies achieved 82-91% accuracy in disease classification tasks, with transcriptomics-based asthma endotyping validated in over 1500 patients across multiple cohorts. Network analysis approaches identified hub genes (IL-6, TNF-α, MMP9) replicated across three independent studies. Machine learning applications demonstrated 80-90% accuracy for diagnostic and prognostic tasks, though external validation remains limited, with only 15% of reviewed studies including independent validation cohorts. Significant challenges persist in data integration, computational reproducibility, and clinical translation. Most studies employed modest sample sizes (median n=89), and population diversity was limited, with 89% conducted in European-ancestry populations. This review provides a comprehensive assessment of systems biology progress in respiratory medicine, identifies methodological gaps, and highlights the need for standardized protocols, larger collaborative studies, and rigorous external validation to advance clinical implementation of systems biology findings in asthma and COPD management.

PMID:41878747 | PMC:PMC13007689 | DOI:10.2147/JAA.S575312

  •  

Systems Biology and Multi-Omics in Asthma and COPD: A Systematic Review of Computational Approaches (2010-2024)

J Asthma Allergy. 2026 Mar 19;19:575312. doi: 10.2147/JAA.S575312. eCollection 2026.

ABSTRACT

Systems biology approaches have contributed to advancing our understanding of complex respiratory diseases including asthma and chronic obstructive pulmonary disease (COPD). This systematic review evaluates the application of systems biology methodologies in respiratory medicine, focusing on multi-omics data integration and computational techniques for biomarker discovery and mechanistic understanding. Following PRISMA 2020 guidelines, we conducted a comprehensive literature search across Web of Science and Scopus databases, identifying 117 peer-reviewed documents published from 2010 to 2024. The review methodology employed bibliometric analysis combined with qualitative synthesis of included studies. Results demonstrate steady growth in systems biology applications for asthma and COPD research, with publication rates increasing by approximately 0.5 articles per year (R2 = 0.73, p < 0.001). Bibliometric analysis identified five major research clusters: systems biology as a foundational methodological framework (Basic Theme), COPD-focused research as the most developed area (Motor Theme), gene expression analysis, disease classification approaches, and specialized lung disease investigations (Niche Theme). Multi-omics integration studies achieved 82-91% accuracy in disease classification tasks, with transcriptomics-based asthma endotyping validated in over 1500 patients across multiple cohorts. Network analysis approaches identified hub genes (IL-6, TNF-α, MMP9) replicated across three independent studies. Machine learning applications demonstrated 80-90% accuracy for diagnostic and prognostic tasks, though external validation remains limited, with only 15% of reviewed studies including independent validation cohorts. Significant challenges persist in data integration, computational reproducibility, and clinical translation. Most studies employed modest sample sizes (median n=89), and population diversity was limited, with 89% conducted in European-ancestry populations. This review provides a comprehensive assessment of systems biology progress in respiratory medicine, identifies methodological gaps, and highlights the need for standardized protocols, larger collaborative studies, and rigorous external validation to advance clinical implementation of systems biology findings in asthma and COPD management.

PMID:41878747 | PMC:PMC13007689 | DOI:10.2147/JAA.S575312

  •  

Sketching a Space of Brain States

arXiv:2603.22296v1 Announce Type: new Abstract: Brain functional connectivity alterations, that is, pathological changes in the signal exchange between areas of the brain, occur in several neurological diseases, including neurodegenerative and neuropsychiatric ones. They consist in changes in how brain functional networks operate. By conceptualising a brain space as a space whose points are connectome configurations representing brain functional states, changes in brain network functionality can be represented by paths between these points. Paths from a healthy state to a diseased one, or between diseased states as instances of disease progression, are modelled as the action of the Krankheit-Operator, which produces changes from a brain functional state to another. This study proposes a formal representation of the space of brain states and presents its computational definition. References to patients affected by Parkinson's disease, schizophrenia, and Alzheimer-Perusini's disease are included to discuss the proposed approach and possible developments of the research toward a generalisation.
  •  

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

arXiv:2603.22529v1 Announce Type: cross Abstract: Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limitation: they focus entirely on web-based interaction and perception, lacking grounding in the user's real-world physical surroundings. This limitation prevents evaluation in crucial scenarios, such as when an agent must use egocentric visual perception (e.g., via AR glasses) to recognize an object in the user's surroundings and then complete a related task online. To address this gap, we introduce Ego2Web, the first benchmark designed to bridge egocentric video perception and web agent execution. Ego2Web pairs real-world first-person video recordings with web tasks that require visual understanding, web task planning, and interaction in an online environment for successful completion. We utilize an automatic data-generation pipeline combined with human verification and refinement to curate well-constructed, high-quality video-task pairs across diverse web task types, including e-commerce, media retrieval, knowledge lookup, etc. To facilitate accurate and scalable evaluation for our benchmark, we also develop a novel LLM-as-a-Judge automatic evaluation method, Ego2WebJudge, which achieves approximately 84% agreement with human judgment, substantially higher than existing evaluation methods. Experiments with diverse SoTA agents on our Ego2Web show that their performance is weak, with substantial headroom across all task categories. We also conduct a comprehensive ablation study on task design, highlighting the necessity of accurate video understanding in the proposed task and the limitations of current agents. We hope Ego2Web can be a critical new resource for developing truly capable AI assistants that can seamlessly see, understand, and act across the physical and digital worlds.
  •  

YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception

arXiv:2603.23037v1 Announce Type: cross Abstract: The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their confidence scores in visually degraded or ambiguous scenes. To address this limitation, a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features. The additive spline-based structure of the Kolmogorov-Arnold network enables direct visualisation of each feature's influence. This produces smooth and transparent functional mappings that reveal when the model's confidence is well supported and when it is unreliable. Experiments on both Common Objects in Context (COCO), and images from the University of Bath campus demonstrate that the framework accurately identifies low-trust predictions under blur, occlusion, or low texture. This provides actionable insights for filtering, review, or downstream risk mitigation. Furthermore, a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene. This tool enables a lightweight multimodal interface without affecting the interpretability layer. The resulting system delivers interpretable object detection with trustworthy confidence estimates. It offers a powerful tool for transparent and practical perception component for autonomous and multimodal artificial intelligence applications.
  •  

Assessing the Robustness of Climate Foundation Models under No-Analog Distribution Shifts

arXiv:2603.23043v1 Announce Type: cross Abstract: The accelerating pace of climate change introduces profound non-stationarities that challenge the ability of Machine Learning based climate emulators to generalize beyond their training distributions. While these emulators offer computationally efficient alternatives to traditional Earth System Models, their reliability remains a potential bottleneck under "no-analog" future climate states, which we define here as regimes where external forcing drives the system into conditions outside the empirical range of the historical training data. A fundamental challenge in evaluating this reliability is data contamination; because many models are trained on simulations that already encompass future scenarios, true out-of-distribution (OOD) performance is often masked. To address this, we benchmark the OOD robustness of three state-of-the-art architectures: U-Net, ConvLSTM, and the ClimaX foundation model specifically restricted to a historical-only training regime (1850-2014). We evaluate these models using two complementary strategies: (i) temporal extrapolation to the recent climate (2015-2023) and (ii) cross-scenario forcing shifts across divergent emission pathways. Our analysis within this experimental setup reveals an accuracy vs. stability trade-off: while the ClimaX foundation model achieves the lowest absolute error, it exhibits higher relative performance changes under distribution shifts, with precipitation errors increasing by up to 8.44% under extreme forcing scenarios. These findings suggest that when restricted to historical training dynamics, even high-capacity foundation models are sensitive to external forcing trajectories. Our results underscore the necessity of scenario-aware training and rigorous OOD evaluation protocols to ensure the robustness of climate emulators under a changing climate.
  •  

Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy

arXiv:2603.23146v1 Announce Type: cross Abstract: The widespread adoption of Large Language Models (LLMs) has made the detection of AI-Generated text a pressing and complex challenge. Although many detection systems report high benchmark accuracy, their reliability in real-world settings remains uncertain, and their interpretability is often unexplored. In this work, we investigate whether contemporary detectors genuinely identify machine authorship or merely exploit dataset-specific artefacts. We propose an interpretable detection framework that integrates linguistic feature engineering, machine learning, and explainable AI techniques. When evaluated on two prominent benchmark corpora, namely PAN CLEF 2025 and COLING 2025, our model trained on 30 linguistic features achieves leaderboard-competitive performance, attaining an F1 score of 0.9734. However, systematic cross-domain and cross-generator evaluation reveals substantial generalisation failure: classifiers that excel in-domain degrade significantly under distribution shift. Using SHAP- based explanations, we show that the most influential features differ markedly between datasets, indicating that detectors often rely on dataset-specific stylistic cues rather than stable signals of machine authorship. Further investigation with in-depth error analysis exposes a fundamental tension in linguistic-feature-based AI text detection: the features that are most discriminative on in-domain data are also the features most susceptible to domain shift, formatting variation, and text-length effects. We believe that this knowledge helps build AI detectors that are robust across different settings. To support replication and practical use, we release an open-source Python package that returns both predictions and instance-level explanations for individual texts.
  •  

Emergence of Fragility in LLM-based Social Networks: the Case of Moltbook

arXiv:2603.23279v1 Announce Type: cross Abstract: The rapid diffusion of large language models and the growth in their capability has enabled the emergence of online environments populated by autonomous AI agents that interact through natural language. These platforms provide a novel empirical setting for studying collective dynamics among artificial agents. In this paper we analyze the interaction network of Moltbook, a social platform composed entirely of LLM based agents, using tools from network science. The dataset comprises 39,924 users, 235,572 posts, and 1,540,238 comments collected through web scraping. We construct a directed weighted network in which nodes represent agents and edges represent commenting interactions. Our analysis reveals strongly heterogeneous connectivity patterns characterized by heavy tailed degree and activity distributions. At the mesoscale, the network exhibits a pronounced core periphery organization in which a very small structural core (0.9% of nodes) concentrates a large fraction of connectivity. Robustness experiments show that the network is relatively resilient to random node removal but highly vulnerable to targeted attacks on highly connected nodes, particularly those with high out degree. These findings indicate that the interaction structure of AI agent social systems may develop strong centralization and structural fragility, providing new insights into the collective organization of LLM native social environments.
  •  

Failure of contextual invariance in gender inference with large language models

arXiv:2603.23485v1 Announce Type: cross Abstract: Standard evaluation practices assume that large language model (LLM) outputs are stable under contextually equivalent formulations of a task. Here, we test this assumption in the setting of gender inference. Using a controlled pronoun selection task, we introduce minimal, theoretically uninformative discourse context and find that this induces large, systematic shifts in model outputs. Correlations with cultural gender stereotypes, present in decontextualized settings, weaken or disappear once context is introduced, while theoretically irrelevant features, such as the gender of a pronoun for an unrelated referent, become the most informative predictors of model behaviour. A Contextuality-by-Default analysis reveals that, in 19--52\% of cases across models, this dependence persists after accounting for all marginal effects of context on individual outputs and cannot be attributed to simple pronoun repetition. These findings show that LLM outputs violate contextual invariance even under near-identical syntactic formulations, with implications for bias benchmarking and deployment in high-stakes settings.
  •  

A transformer architecture alteration to incentivise externalised reasoning

arXiv:2603.21376v2 Announce Type: replace Abstract: We propose a new architectural change, and post-training pipeline, for making LLMs more verbose reasoners by teaching a model to truncate forward passes early. We augment an existing transformer architecture with an early-exit mechanism at intermediate layers and train the model to exit at shallower layers when the next token can be predicted without deep computation. After a calibration stage, we incentivise the model to exit as early as possible while maintaining task performance using reinforcement learning. We provide preliminary results to this effect for small reasoning models, showing that they learn to adaptively reduce computations across tokens. We predict that, applied at the right scale, our approach can minimise the amount of excess computation that reasoning models have at their disposal to perform non-myopic planning using their internal activations, reserving this only for difficult-to-predict tokens.
  •  

GUIrilla: A Scalable Framework for Automated Desktop UI Exploration

arXiv:2510.16051v2 Announce Type: replace-cross Abstract: The performance and generalization of foundation models for interactive systems critically depend on the availability of large-scale, realistic training data. While recent advances in large language models (LLMs) have improved GUI understanding, progress in desktop automation remains constrained by the scarcity of high-quality, publicly available desktop interaction data, particularly for macOS. We introduce GUIRILLA, a scalable data crawling framework for automated exploration of desktop GUIs. GUIRILLA is not an autonomous agent; instead, it systematically collects realistic interaction traces and accessibility metadata intended to support the training, evaluation, and stabilization of downstream foundation models and GUI agents. The framework targets macOS, a largely underrepresented platform in existing resources, and organizes explored interfaces into hierarchical MacApp Trees derived from accessibility states and user actions. As part of this work, we release these MacApp Trees as a reusable structural representation of macOS applications, enabling downstream analysis, retrieval, testing, and future agent training. We additionally release macapptree, an open-source library for reproducible accessibility-driven GUI data collection, along with the full framework implementation to support open research in desktop autonomy.
  •  

When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift

arXiv:2603.04648v2 Announce Type: replace-cross Abstract: Real-world reinforcement learning systems must operate under distributional drift in their observation streams, yet most policy architectures implicitly assume fully observed and noise-free states. We study robustness of Proximal Policy Optimization (PPO) under temporally persistent sensor failures that induce partial observability and representation shift. To respond to this drift, we augment PPO with temporal sequence models, including Transformers and State Space Models (SSMs), to enable policies to infer missing information from history and maintain performance. Under a stochastic sensor failure process, we prove a high-probability bound on infinite-horizon reward degradation that quantifies how robustness depends on policy smoothness and failure persistence. Empirically, on MuJoCo continuous-control benchmarks with severe sensor dropout, we show Transformer-based sequence policies substantially outperform MLP, RNN, and SSM baselines in robustness, maintaining high returns even when large fractions of sensors are unavailable. These results demonstrate that temporal sequence reasoning provides a principled and practical mechanism for reliable operation under observation drift caused by sensor unreliability.
  •  

A fast starburst wind consumes most of the energy from supernovae

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10231-1

Starburst galaxies are seen to host galaxy-scale winds, which are super-fast and could be powered entirely by the thermal pressure of gas heated by supernovae.
  •  

Androgen activity in the male embryonic hindbrain drives lethal PFA ependymoma

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10264-6

Androgen activity in the male embryonic hindbrain prolongs hindbrain differentiation in male individuals and drives sex differences in the incidence and prognosis of posterior fossa type A (PFA) ependymoma, an aggressive childhood brain tumour.
  •  

Pembrolizumab and olaparib in homologous-recombination-deficient metastatic pancreatic cancer: the phase 2 POLAR trial

Nature Medicine, Published online: 25 March 2026; doi:10.1038/s41591-026-04299-5

Results of the phase 2 POLAR trial show that biomarker-guided treatment in patients with metastatic pancreatic cancer based on homologous repair deficiency leads to encouraging clinical response rates in immune cell-infiltrated tumors.
  •  

Parasites trigger epithelial cell crosstalk to drive gut–brain signalling

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10281-5

Paracrine signalling between tuft cells and enterochromaffin cells is a key mode of immune–sensory and gut–brain communication, and accounts for the pattern of gastrointestinal symptoms that occurs during parasite infections.
  •  

Disequilibrium response to tapping crustal magma reveals storage conditions

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10317-w

Magma drilling data from Krafla volcano, Iceland, are used to reconstruct in situ lithostatic magmatic conditions using disequilibrium simulations that provide a method for improving the understanding of magma storage conditions and evolution.
  •  
❌