❌

Normal view

Europe’s cyber agency blames hacking gangs for massive data breach and leak

3 April 2026 at 23:50
CERT-EU blamed the cybercrime group TeamPCP for the recent hack on the European Commission, and said the notorious ShinyHunters gang was responsible for leaking the stolen data online.

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics

arXiv:2604.01375v1 Announce Type: new Abstract: Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubrics using downstream signals such as reinforcement learning outcomes, there remains no principled way to diagnose rubric quality issues from such aggregated or downstream signals alone. To address this gap, we introduce RIFT: RubrIc Failure mode Taxonomy, a taxonomy for systematically characterizing failure modes in rubric composition and design. RIFT consists of eight failure modes organized into three high-level categories: Reliability Failures, Content Validity Failures, and Consequential Validity Failures. RIFT is developed using grounded theory by iteratively annotating rubrics drawn from five diverse benchmarks spanning general instruction following, code generation, creative writing, and expert-level deep research, until no new failure modes are identified. We evaluate the consistency of the taxonomy by measuring agreement among independent human annotators, observing fair agreement overall (87% pairwise agreement and 0.64 average Cohen's kappa). Finally, to support scalable diagnosis, we propose automated rubric quality metrics and show that they align with human failure-mode annotations, achieving up to 0.86 F1.

Qiana: A First-Order Formalism to Quantify over Contexts and Formulas with Temporality

arXiv:2604.01952v1 Announce Type: new Abstract: We introduce Qiana, a logic framework for reasoning on formulas that are true only in specific contexts. In Qiana, it is possible to quantify over both formulas and contexts to express, e.g., that ``everyone knows everything Alice says''. Qiana also permits paraconsistent logics within contexts, so that contexts can contain contradictions. Furthermore, Qiana is based on first-order logic, and is finitely axiomatizable, so that Qiana theories are compatible with pre-existing first-order logic theorem provers. We show how Qiana can be used to represent temporality, event calculus, and modal logic. We also discuss different design alternatives of Qiana.

From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems

arXiv:2604.02198v1 Announce Type: new Abstract: While Artificial Intelligence (AI) offers transformative potential for operational performance, its deployment in safety-critical domains such as aviation requires strict adherence to rigorous certification standards. Current EASA guidelines mandate demonstrating complete coverage of the AI/ML constituent's Operational Design Domain (ODD) -- a requirement that demands proof that no critical gaps exist within defined operational boundaries. However, as systems operate within high-dimensional parameter spaces, existing methods struggle to provide the scalability and formal grounding necessary to satisfy the completeness criterion. Currently, no standardized engineering method exists to bridge the gap between abstract ODD definitions and verifiable evidence. This paper addresses this void by proposing a method that integrates parameter discretization, constraint-based filtering, and criticality-based dimension reduction into a structured, multi-step ODD coverage verification process. Grounded in gathered simulation data from prior research on AI-based mid-air collision avoidance research, this work demonstrates a systematic engineering approach to defining and achieving coverage metrics that satisfy EASA's demand for completeness. Ultimately, this method enables the validation of ODD coverage in higher dimensions, advancing a Safety-by-Design approach while complying with EASA's standards.

A deep learning pipeline for PAM50 subtype classification using histopathology images and multi-objective patch selection

arXiv:2604.01798v1 Announce Type: cross Abstract: Breast cancer is a highly heterogeneous disease with diverse molecular profiles. The PAM50 gene signature is widely recognized as a standard for classifying breast cancer into intrinsic subtypes, enabling more personalized treatment strategies. In this study, we introduce a novel optimization-driven deep learning framework that aims to reduce reliance on costly molecular assays by directly predicting PAM50 subtypes from H&E-stained whole-slide images (WSIs). Our method jointly optimizes patch informativeness, spatial diversity, uncertainty, and patch count by combining the non-dominated sorting genetic algorithm II (NSGA-II) with Monte Carlo dropout-based uncertainty estimation. The proposed method can identify a small but highly informative patch subset for classification. We used a ResNet18 backbone for feature extraction and a custom CNN head for classification. For evaluation, we used the internal TCGA-BRCA dataset as the training cohort and the external CPTAC-BRCA dataset as the test cohort. On the internal dataset, an F1-score of 0.8812 and an AUC of 0.9841 using 627 WSIs from the TCGA-BRCA cohort were achieved. The performance of the proposed approach on the external validation dataset showed an F1-score of 0.7952 and an AUC of 0.9512. These findings indicate that the proposed optimization-guided, uncertainty-aware patch selection can achieve high performance and improve the computational efficiency of histopathology-based PAM50 classification compared to existing methods, suggesting a scalable imaging-based replacement that has the potential to support clinical decision-making.

Towards Transparent and Efficient Anomaly Detection in Industrial Processes through ExIFFI

arXiv:2405.01158v4 Announce Type: replace-cross Abstract: Anomaly Detection (AD) is crucial in industrial settings to streamline operations by detecting underlying issues. Conventional methods merely label observations as normal or anomalous, lacking crucial insights. In Industry 5.0, interpretable outcomes become desirable to enable users to understand the rational under model decisions. This paper presents the first industrial application of ExIFFI, a recent approach for fast, efficient explanations for the Extended Isolation Forest (EIF) AD method. ExIFFI is tested on four industrial datasets, demonstrating superior explanation effectiveness, computational efficiency and improved raw anomaly detection performances. ExIFFI reaches over then 90\% of average precision on all the benchmarks considered in the study and overperforms state-of-the-art Explainable Artificial Intelligence (XAI) approaches in terms of the feature selection proxy task metric which was specifically introduced to quantitatively evaluate model explanations.

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

arXiv:2603.27044v2 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space. A recent framework, which we refer to as Action-based Policy Compression (APC), mitigates this issue by compressing the parameter space $\Theta$ into a low-dimensional latent manifold $\mathcal Z$ using a learned generative mapping $g:\mathcal Z \to \Theta$. However, its performance is severely constrained by relying on immediate action-matching as a reconstruction loss, a myopic proxy for behavioral similarity that suffers from compounding errors across sequential decisions. To overcome this bottleneck, we introduce Occupancy-based Policy Compression (OPC), which enhances APC by shifting behavior representation from immediate action-matching to long-horizon state-space coverage. Specifically, we propose two principal improvements: (1) we curate the dataset generation with an information-theoretic uniqueness metric that delivers a diverse population of policies; and (2) we propose a fully differentiable compression objective that directly minimizes the divergence between the true and reconstructed mixture occupancy distributions. These modifications force the generative model to organize the latent space around true functional similarity, promoting a latent representation that generalizes over a broad spectrum of behaviors while retaining most of the original parameter space's expressivity. Finally, we empirically validate the advantages of our contributions across multiple continuous control benchmarks.

Target product profiles for treatments to delay or prevent symptomatic Alzheimer’s disease

Nature Medicine, Published online: 03 April 2026; doi:10.1038/s41591-026-04305-w

To accelerate therapeutic development and equip stakeholders with clear benchmarks, the authors outline target product profiles for therapies designed to delay or prevent the onset of clinical symptoms of Alzheimer’s disease.

Single-cell and spatial profiling in cancer biology and clinical oncology

Nature Cancer, Published online: 03 April 2026; doi:10.1038/s43018-026-01142-1

Izar and colleagues review the insights into cancer biology gained via single-cell analyses and spatial profiling and overview the challenges and opportunities associated with the implementation of these approaches to guide clinical discovery.

Ethical Handling of Occupational Health and Safety Data in the Fire Service: Empirical Interview and Focus Group Study of Firefighter and Fire Service Leadership Privacy Preferences

Background: There are ongoing efforts to collect larger and higher-quality amounts of occupational health and safety data to better understand and prevent injuries and fatalities among high-risk workers, such as firefighters. Digital health systems including wearable technologies, mobile apps, or internet-based data collection platforms could collect large amounts of sensitive data, but there is little evidence on worker and employer perspectives on data privacy in the fire service. Objective: Our study examined firefighters’ and fire service leadership’s preferences regarding occupational health and safety data privacy. Methods: We conducted interviews and focus groups with career firefighters in Maryland and Virginia; interviews with union representatives and department-level leaders in each state; and interviews with national-level fire service leaders in advocacy, government, and research organizations (March to November 2023). Interviews and focus groups were audio recorded and transcribed. We analyzed transcripts using thematic analysis. Results: The sample included 31 career firefighters, 2 union leaders, 11 national leaders, and 21 department-level leaders (65 total participants from 35 interviews and 4 focus groups). We identified 4 themes: acceptability of data access, sharing, and reporting practices; data sharing and access preferences; appropriate use of firefighter data; and the need for improved communication. Leaders described firefighters’ concerns about job loss and loss of privacy. Firefighters expressed general preferences that their data be deidentified and not shared widely, and they identified mental health data as important but particularly sensitive information. Firefighters also expressed frustration about sharing data with researchers or their departments without knowing the purpose or outcomes. Both firefighters and leaders emphasized the need for enhanced communication and translation of data for firefighters. Conclusions: Fire service leaders held more concerns about the use and sharing of occupational health and safety data than firefighters, but both groups identified ways to further safeguard firefighter data and improve communication about health and safety data. Future fire service data collection should incorporate privacy protections, such as limiting the collection of identifiable information and restricting data access. Data collection should be accompanied by clear communication about the purpose of the data collection, how firefighter data will be used and accessed, and the interpretation of the results. Future digital health interventions should integrate these data privacy protections to respect firefighter preferences and contribute to acceptability and uptake.

Proteogenomic Analysis of Coronary Artery Calcification in Human Populations

Arterioscler Thromb Vasc Biol. 2026 Apr 2. doi: 10.1161/ATVBAHA.125.324171. Online ahead of print.

ABSTRACT

BACKGROUND: Joint use of multiple molecular layers can be useful to prioritize targets for mechanistic studies. Application of coronary disease in large populations is an emerging field.

METHODS: We used reported circulating proteomic data (Somascan aptamer-based) from ≈3000 individuals in the CARDIA study (Coronary Artery Risk Development in Young Adults), measuring association with prevalent and 10-year incident coronary artery calcium (CAC) score. We used a multiparametric approach to prioritize circulating protein-CAC associations via genomics of circulating protein levels and coronary artery transcription.

RESULTS: Proteins linked to prevalent/incident CAC in CARDIA implicated pathogenic mechanisms of vascular disease, including fibrosis and inflammation (GDF-15 [growth/differentiation factor 15], CDCP1 [CUB domain-containing protein 1], GSN [gelsolin], THBS2 [thrombospondin-2], chemokines, RNAS6), oxidative lipid metabolism (CILP2), extracellular matrix remodeling and signaling (MMPs [matrix metalloproteinases], TIMP-1, integrins), calcification (Notch 1, ARHGAP36 [Rho GTPase-activating protein 36]), and metabolism (GIP [gastric inhibitory polypeptide]), as well as new proteins not previously reported. Using protein-wide association study genetic approaches, several targets with nominal evidence in CAC proteomics were associated with atherosclerosis or myocardial infarction in over 300K individuals, including PCSK9 (proprotein convertase subtilisin/kexin type 9) and APO C1. Finally, the coronary artery-specific transcriptome-wide association study of CAC yielded genes with previously implicated mechanistic roles in vascular homeostasis, inflammation, and metabolism, as well as genes without previously described function in CAC. Overlap across CAC proteomics and transcriptome-wide association study highlighted genes involved in vascular inflammation (S100A9), cardiac development (HES1), vessel wall structure (SPARCL1), and vascular dysfunction or plaque (NOTCH3, TNFSF12, S100A12).

CONCLUSIONS: These results report population-level multiomics in human coronary calcification, presenting a method to identify disease-relevant targets through integration of human genetic approaches with multiomics.

PMID:41924874 | DOI:10.1161/ATVBAHA.125.324171

Proteogenomic Analysis of Coronary Artery Calcification in Human Populations

Arterioscler Thromb Vasc Biol. 2026 Apr 2. doi: 10.1161/ATVBAHA.125.324171. Online ahead of print.

ABSTRACT

BACKGROUND: Joint use of multiple molecular layers can be useful to prioritize targets for mechanistic studies. Application of coronary disease in large populations is an emerging field.

METHODS: We used reported circulating proteomic data (Somascan aptamer-based) from ≈3000 individuals in the CARDIA study (Coronary Artery Risk Development in Young Adults), measuring association with prevalent and 10-year incident coronary artery calcium (CAC) score. We used a multiparametric approach to prioritize circulating protein-CAC associations via genomics of circulating protein levels and coronary artery transcription.

RESULTS: Proteins linked to prevalent/incident CAC in CARDIA implicated pathogenic mechanisms of vascular disease, including fibrosis and inflammation (GDF-15 [growth/differentiation factor 15], CDCP1 [CUB domain-containing protein 1], GSN [gelsolin], THBS2 [thrombospondin-2], chemokines, RNAS6), oxidative lipid metabolism (CILP2), extracellular matrix remodeling and signaling (MMPs [matrix metalloproteinases], TIMP-1, integrins), calcification (Notch 1, ARHGAP36 [Rho GTPase-activating protein 36]), and metabolism (GIP [gastric inhibitory polypeptide]), as well as new proteins not previously reported. Using protein-wide association study genetic approaches, several targets with nominal evidence in CAC proteomics were associated with atherosclerosis or myocardial infarction in over 300K individuals, including PCSK9 (proprotein convertase subtilisin/kexin type 9) and APO C1. Finally, the coronary artery-specific transcriptome-wide association study of CAC yielded genes with previously implicated mechanistic roles in vascular homeostasis, inflammation, and metabolism, as well as genes without previously described function in CAC. Overlap across CAC proteomics and transcriptome-wide association study highlighted genes involved in vascular inflammation (S100A9), cardiac development (HES1), vessel wall structure (SPARCL1), and vascular dysfunction or plaque (NOTCH3, TNFSF12, S100A12).

CONCLUSIONS: These results report population-level multiomics in human coronary calcification, presenting a method to identify disease-relevant targets through integration of human genetic approaches with multiomics.

PMID:41924874 | DOI:10.1161/ATVBAHA.125.324171

Semaglutide on liver fibrosis and heart outcomes in patients at high risk of liver fibrosis: a prespecified analysis of the SELECT randomized trial

Nature Medicine, Published online: 02 April 2026; doi:10.1038/s41591-026-04281-1

A prespecified analysis from the SELECT trial showed that semaglutide reduces major adverse cardiovascular events by 20% compared with placebo, particularly in patients at high risk of fibrosis, as indicated by the Fibrosis-4 index.

The MicrobeAtlas database: Global trends and insights into Earth’s microbial ecosystems

MicrobeAtlas (www.microbeatlas.org) is an integrated, reference-based resource for truly planet-wide microbiomics, analyzing hundreds of thousands of microbial lineages across diverse environments, conditions, and technologies.

Large-scale proteomics across neurological disorders uncovers biomarker panel and targets in multiple sclerosis

Deep proteome profiling of over 5,000 cerebrospinal fluid samples by mass spectrometry maps protein alterations across major neurological disorders, resolving key sources of variation as well as shared and disease-specific signatures. This framework yields a 22-protein assay that improves the differential diagnosis of multiple sclerosis from other inflammatory conditions, particularly in diagnostically challenging oligoclonal band-negative individuals.
❌