❌

Normal view

Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models

arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who might fine-tune these models for malicious use. To address this gap, we propose \eval, a framework to evaluate the robustness of procedures that are intended to reduce the dual-use capabilities of bio-foundation models. \eval assesses models' virus understanding through three lenses, including sequence modeling, mutational effects prediction, and virulence prediction. Our results show that current filtering practices may not be particularly effective: Excluded knowledge can be rapidly recovered in some cases via fine-tuning, and exhibits broader generalizability in sequence modeling. Furthermore, dual-use signals may already reside in the pretrained representations, and can be elicited via simple linear probing. These findings highlight the challenges of data filtering as a standalone procedure, underscoring the need for further research into robust safety and security strategies for open-weight bio-foundation models.

Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study

arXiv:2510.21389v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must discern \textit{when} and \textit{why} to trust algorithmic recommendations. This work presents an application-grounded user study with eight professional sleep medicine practitioners, who score nocturnal arousal events in polysomnographic data under three conditions: (i) manual scoring, (ii) black-box (BB) AI assistance, and (iii) transparent white-box (WB) AI assistance. Assistance is provided either from the \textit{start} of scoring or as a post-hoc quality-control (\textit{QC}) review. We systematically evaluate how the type and timing of assistance influence event-level and clinically most relevant count-based performance, time requirements, and user experience. When evaluated against the clinical standard used to train the AI, both AI and human-AI teams significantly outperform unaided experts, with collaboration also reducing inter-rater variability. Notably, transparent AI assistance applied as a targeted QC step yields median event-level performance improvements of approximately 30\% over black-box assistance, and QC timing further enhances count-based outcomes. While WB and QC approaches increase the time required for scoring, start-time assistance is faster and preferred by most participants. Participants overwhelmingly favor transparency, with seven out of eight expressing willingness to adopt the system with minor or no modifications. In summary, strategically timed transparent AI assistance effectively balances accuracy and clinical efficiency, providing a promising pathway toward trustworthy AI integration and user acceptance in clinical workflows.

A Definition of AGI

arXiv:2510.18212v2 Announce Type: replace Abstract: The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

Codon specific readthrough as a mechanism of BRCA2 restoration in acquired PARP inhibitor and chemotherapy resistance

Nucleic Acids Res. 2025 Oct 14;53(19):gkaf990. doi: 10.1093/nar/gkaf990.

ABSTRACT

BRCA2 mutations contribute to the pathogenesis and treatment sensitivity of a subset of ovarian, breast, prostate, and pancreatic cancers. When these cancers become therapy resistant, secondary mutations that restore the BRCA2 open reading frame are found in half the cases, but other causes of resistance remain incompletely understood. Here, we identified translational readthrough of a premature termination codon (PTC) as a cause of resistance to poly(ADP-ribose) polymerase inhibitors (PARPis) and cisplatin in cells derived from the BRCA2-mutated ovarian cancer line PEO1 by PARPi selection. Despite persistence of the signature 4965C > G (p.Y1655X) BRCA2 mutation, low-level expression of full-length BRCA2 protein was detectable in these cells by immunoblotting and tandem mass spectrometry. Either BRCA2 knockdown or gene interruption 5' or 3' to the PTC restored treatment sensitivity, implicating BRCA2 in the resistance. Reporter assays demonstrated UAG-selective readthrough in the resistant clones but not parental cells. Moreover, custom searching of global proteomic data indicated readthrough of stop codons, particularly UAGs, in additional proteins in the resistant clones. Finally, multi-omic analysis identified multiple changes in the nonsense-mediated decay and termination machineries that favor readthrough. Accordingly, the present results identify PTC readthrough as a potential mechanism of drug resistance in cells with BRCA2 nonsense mutations.

PMID:41099700 | PMC:PMC12526053 | DOI:10.1093/nar/gkaf990

❌