❌

Normal view

A Workflow for Full Traceability of AI Decisions

arXiv:2511.11275v2 Announce Type: replace Abstract: An ever increasing number of high-stake decisions are made or assisted by automated systems employing brittle artificial intelligence technology. There is a substantial risk that some of these decision induce harm to people, by infringing their well-being or their fundamental human rights. The state-of-the-art in AI systems makes little effort with respect to appropriate documentation of the decision process. This obstructs the ability to trace what went into a decision, which in turn is a prerequisite to any attempt of reconstructing a responsibility chain. Specifically, such traceability is linked to a documentation that will stand up in court when determining the cause of some AI-based decision that inadvertently or intentionally violates the law. This paper takes a radical, yet practical, approach to this problem, by enforcing the documentation of each and every component that goes into the training or inference of an automated decision. As such, it presents the first running workflow supporting the generation of tamper-proof, verifiable and exhaustive traces of AI decisions. In doing so, we expand the DBOM concept into an effective running workflow leveraging confidential computing technology. We demonstrate the inner workings of the workflow in the development of an app to tell poisonous and edible mushrooms apart, meant as a playful example of high-stake decision support.

Fair human-centric image dataset for ethical AI benchmarking

Nature, Published online: 05 November 2025; doi:10.1038/s41586-025-09716-2

The Fair Human-Centric Image Benchmark (FHIBE, pronounced ‘Feebee’)—an image dataset that implements best practices for consent, privacy, compensation, safety, diversity and utility—can be used responsibly as a fairness evaluation dataset for many human-centric computer vision applications.

An In-depth Study of LLM Contributions to the Bin Packing Problem

arXiv:2510.27353v1 Announce Type: new Abstract: Recent studies have suggested that Large Language Models (LLMs) could provide interesting ideas contributing to mathematical discovery. This claim was motivated by reports that LLM-based genetic algorithms produced heuristics offering new insights into the online bin packing problem under uniform and Weibull distributions. In this work, we reassess this claim through a detailed analysis of the heuristics produced by LLMs, examining both their behavior and interpretability. Despite being human-readable, these heuristics remain largely opaque even to domain experts. Building on this analysis, we propose a new class of algorithms tailored to these specific bin packing instances. The derived algorithms are significantly simpler, more efficient, more interpretable, and more generalizable, suggesting that the considered instances are themselves relatively simple. We then discuss the limitations of the claim regarding LLMs' contribution to this problem, which appears to rest on the mistaken assumption that the instances had previously been studied. Our findings instead emphasize the need for rigorous validation and contextualization when assessing the scientific value of LLM-generated outputs.

Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models

arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who might fine-tune these models for malicious use. To address this gap, we propose \eval, a framework to evaluate the robustness of procedures that are intended to reduce the dual-use capabilities of bio-foundation models. \eval assesses models' virus understanding through three lenses, including sequence modeling, mutational effects prediction, and virulence prediction. Our results show that current filtering practices may not be particularly effective: Excluded knowledge can be rapidly recovered in some cases via fine-tuning, and exhibits broader generalizability in sequence modeling. Furthermore, dual-use signals may already reside in the pretrained representations, and can be elicited via simple linear probing. These findings highlight the challenges of data filtering as a standalone procedure, underscoring the need for further research into robust safety and security strategies for open-weight bio-foundation models.

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.

Impact and Implications of Generative AI for Enterprise Architects in Agile Environments: A Systematic Literature Review

arXiv:2510.22003v1 Announce Type: cross Abstract: Generative AI (GenAI) is reshaping enterprise architecture work in agile software organizations, yet evidence on its effects remains scattered. We report a systematic literature review (SLR), following established SLR protocols of Kitchenham and PRISMA, of 1,697 records, yielding 33 studies across enterprise, solution, domain, business, and IT architect roles. GenAI most consistently supports (i) design ideation and trade-off exploration; (ii) rapid creation and refinement of artifacts (e.g., code, models, documentation); and (iii) architectural decision support and knowledge retrieval. Reported risks include opacity and bias, contextually incorrect outputs leading to rework, privacy and compliance concerns, and social loafing. We also identify emerging skills and competencies, including prompt engineering, model evaluation, and professional oversight, and organizational enablers around readiness and adaptive governance. The review contributes with (1) a mapping of GenAI use cases and risks in agile architecting, (2) implications for capability building and governance, and (3) an initial research agenda on human-AI collaboration in architecture. Overall, the findings inform responsible adoption of GenAI that accelerates digital transformation while safeguarding architectural integrity.

Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents

arXiv:2510.22620v1 Announce Type: cross Abstract: AI agents powered by large language models (LLMs) are being deployed at scale, yet we lack a systematic understanding of how the choice of backbone LLM affects agent security. The non-deterministic sequential nature of AI agents complicates security modeling, while the integration of traditional software with AI components entangles novel LLM vulnerabilities with conventional security risks. Existing frameworks only partially address these challenges as they either capture specific vulnerabilities only or require modeling of complete agents. To address these limitations, we introduce threat snapshots: a framework that isolates specific states in an agent's execution flow where LLM vulnerabilities manifest, enabling the systematic identification and categorization of security risks that propagate from the LLM to the agent level. We apply this framework to construct the $\operatorname{b}^3$ benchmark, a security benchmark based on 194331 unique crowdsourced adversarial attacks. We then evaluate 31 popular LLMs with it, revealing, among other insights, that enhanced reasoning capabilities improve security, while model size does not correlate with security. We release our benchmark, dataset, and evaluation code to facilitate widespread adoption by LLM providers and practitioners, offering guidance for agent developers and incentivizing model developers to prioritize backbone security improvements.

Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation

arXiv:2510.23241v1 Announce Type: cross Abstract: In this work, we introduce Progressive Growing of Patch Size, an automatic curriculum learning approach for 3D medical image segmentation. Our approach progressively increases the patch size during model training, resulting in an improved class balance for smaller patch sizes and accelerated convergence of the training process. We evaluate our curriculum approach in two settings: a resource-efficient mode and a performance mode, both regarding Dice score performance and computational costs across 15 diverse and popular 3D medical image segmentation tasks. The resource-efficient mode matches the Dice score performance of the conventional constant patch size sampling baseline with a notable reduction in training time to only 44%. The performance mode improves upon constant patch size segmentation results, achieving a statistically significant relative mean performance gain of 1.28% in Dice Score. Remarkably, across all 15 tasks, our proposed performance mode manages to surpass the constant patch size baseline in Dice Score performance, while simultaneously reducing training time to only 89%. The benefits are particularly pronounced for highly imbalanced tasks such as lesion segmentation tasks. Rigorous experiments demonstrate that our performance mode not only improves mean segmentation performance but also reduces performance variance, yielding more trustworthy model comparison. Furthermore, our findings reveal that the proposed curriculum sampling is not tied to a specific architecture but represents a broadly applicable strategy that consistently boosts performance across diverse segmentation models, including UNet, UNETR, and SwinUNETR. In summary, we show that this simple yet elegant transformation on input data substantially improves both Dice Score performance and training runtime, while being compatible across diverse segmentation backbones.

Alternatives to animal testing are the future — it’s time that journals, funders and scientists embrace them

Nature, Published online: 20 October 2025; doi:10.1038/d41586-025-03344-6

Biomedical research techniques that don’t involve the use of animals are gaining momentum, but those using innovative approaches still face resistance from some quarters.

Framework for the Development and Delivery of Digital Peer Support Programs: Qualitative Study on in-Person and Digital Delivery for People With Cardiovascular Disease

Background: Peer support (sharing experiences/support with others with the same condition) improves health outcomes among people with cardiovascular disease (CVD), including self-management behaviours and self-efficacy. However, current peer support interventions are diverse. Evidence is lacking on peer support attenders perceptions of benefits and the elements that are considered priorities, especially for digital interventions. Objective: The study objectives were to 1) describe perceived benefits and recommendations for CVD peer support programs from people attending in-person peer support, 2) identify priorities for digital peer support from consumers and clinicians testing a peer support app prototype, and 3) develop a framework to inform future peer support intervention development. Methods: Qualitative methodology was used across two components to address the objectives of this study. In Component 1, semi-structured focus groups were conducted with attenders of established in-person CVD peer support groups, exploring the perceived benefits of peer support and recommendations for future programs. In Component 2, semi-structured interactive workshops with consumers with CVD and semi-structured online interviews with CVD clinicians/researchers were undertaken seeking feedback and recommendations for digital peer support using an exploratory digital CVD peer support application prototype. Data were recorded digitally, transcribed verbatim, and analysed thematically. Findings from both components were iteratively synthesised to inform a digital peer support development framework. Results: In Component 1, 22 participants (age range 29-84 years, male 45%) took part in focus groups. The overarching theme was that peer support provides benefits through sharing experiences. Five themes were refined and defined; (i) peer support provides a way of coping, (ii) peers learn from each other, (iii) peers understand what each other are going through, (iv) the peer community uplifts mood and build confidence, and (v) awareness, flexibility and resources are important for engagement. In Component 2, five participants (age range 55-74 years, male 60%) attended two workshops and eight clinicians/researchers (age range 30-65 years, male 10%) were interviewed. Three themes were refined and defined: (i) autonomy is essential to promote engagement, (ii) safeguarding is important to both users and clinicians, and (iii) interfaces that are simple, easy to use and visually attractive enable use. Priorities identified from both components included greater peer support awareness and uptake, flexibility with timing and family participation, healthcare professional involvement, provision of resources, autonomous features enabling choice, checklists and clinician moderation for safeguarding, and simple to use interfaces. Conclusions: Participants in peer support programs derive benefit from sharing their experience of living with CVD which enable coping, learning, feeling understood and a sense of community. Priorities were synthesised to create a framework for digital peer support development for future peer support with recommendations to focus on six key areas: uptake, flexibility, resources, autonomy, safeguarding and interface.

An eyecare foundation model for clinical assistance: a randomized controlled trial

Nature Medicine, Published online: 28 August 2025; doi:10.1038/s41591-025-03900-7

Trained and validated on multimodal data from 14.5 million images from multicountry datasets, a foundation model is shown to increase diagnostic and referral accuracy of clinicians when used as an assistant in a trial involving 16 ophthalmologists and 668 patients.

GeneBits: ultra-sensitive tumour-informed ctDNA monitoring of treatment response and relapse in cancer patients

J Transl Med. 2025 Aug 27;23(1):964. doi: 10.1186/s12967-025-06993-3.

ABSTRACT

BACKGROUND: Circulating tumour DNA (ctDNA) in liquid biopsies has emerged as a powerful biomarker in cancer patients. Its relative abundance in cell-free DNA serves as a proxy for the overall tumour burden. Here we present GeneBits, a method for cancer therapy monitoring and relapse detection. GeneBits employs tumour-informed enrichment panels targeting 20-100 somatic single-nucleotide variants (SNVs) in plasma-derived DNA, combined with ultra-deep sequencing and unique molecular barcoding. In conjunction with the newly developed computational method umiVar, GeneBits enables accurate detection of molecular residual disease and early relapse identification.

RESULTS: To assess the performance of GeneBits and umiVar, we conducted benchmarking experiments using three different commercial cell-free DNA reference standards. These standards were tested with targeted next-generation sequencing (NGS) workflows from both IDT and Twist, allowing us to evaluate the consistency and accuracy of our approach across different oligo-enrichment strategies. GeneBits achieved comparable depth of coverage across all target sites, demonstrating robust performance independent of the enrichment kit used. For duplex reads with ≥ 4x UMI-family size, umiVar achieved exceptionally low error rates, ranging from 7.4×10-7 to 7.5×10-5. Even when including mixed consensus reads (duplex & simplex), error rates remained low, between 6.1×10-6 and 9×10-5. Furthermore, umiVar enabled variant detection at a limit of detection as low as 0.0017%, with no false positive calls in mutation-free reference samples. In a reanalysed melanoma cohort, variant allele frequency kinetics closely mirrored imaging results, confirming the clinical relevance of our method.

CONCLUSION: GeneBits and umiVar enable highly accurate therapy and relapse monitoring in plasma as well as identification of molecular residual disease within four weeks of tumour surgery or biopsy. By leveraging small, tumour-informed sequencing panels, GeneBits provides a targeted, cost-effective, and scalable approach for ctDNA-based cancer monitoring. The benchmarking experiments using multiple commercial cell-free DNA reference standards confirmed the high sensitivity and specificity of GeneBits and umiVar, making them valuable tools for precision oncology. UmiVar is available at https://github.com/imgag/umiVar .

PMID:40866952 | PMC:PMC12382282 | DOI:10.1186/s12967-025-06993-3

Integrative genomic identification of therapeutic targets for pancreatic cancer

Cell Rep. 2025 Aug 21;44(9):116191. doi: 10.1016/j.celrep.2025.116191. Online ahead of print.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) is a deadly disease, and new therapeutic strategies are urgently needed. Here, we conduct an integrative, genome-scale examination of genetic dependencies and cell surface targets using CRISPR-Cas screening and multi-omic data, including single-nucleus and spatial transcriptomic data from patient tumors. We systematically identify clinically tractable and biomarker-linked PDAC dependencies, including CDS2 as a synthetic lethal target in cancer cells expressing signatures of epithelial-to-mesenchymal transition. We examine biomarkers and co-dependencies of the KRAS oncogene, defining gene expression signatures of sensitivity and resistance associated with response to pharmacological inhibition of KRAS. mRNA and protein profiling reveal cell surface protein-encoding genes with robust expression in patient tumors and minimal expression in non-malignant tissues. Furthermore, we define intratumoral and interpatient heterogeneity of target gene expression and identify orthogonal targets that suggest combinatorial strategies. Collectively, this work identifies multiple targets that may inform therapeutic strategies for patients with PDAC.

PMID:40848256 | DOI:10.1016/j.celrep.2025.116191

❌