❌

Reading view

HGKAN: Hypergraph Kolmogorov-Arnold Networks for Interpretable Prediction of Herb–Symptom Associations

With the increasing availability of large-scale traditional Chinese medicine (TCM) data, accurate prediction of herb–symptom associations (HSAs) has become a crucial task in natural drug discovery. Existing computational approaches mainly employ graph neural networks (GNNs) to model herb–symptom relationships in biological information networks. However, these methods are limited to low-order pairwise associations, failing to capture herbs’ high-order effects and providing limited interpretability. To address these challenges, we model HSAs as a hierarchical hypergraph, where ingredients and targets form two distinct node layers, herbs and symptoms are represented as hyperedges connecting multiple nodes, and inter-layer edges capture ingredient–target interactions. Based on this structure, we propose hypergraph Kolmogorov-Arnold networks (HGKAN) comprising two modules: a binary interaction KAN (BiKAN) with bidirectional cross-attention for encoding pairwise ingredient–target interactions, and two high-order interaction KAN (HiKAN) branches for node-hyperedge message passing and selective high-order feature aggregation. Extensive experiments on two public TCM datasets show that HGKAN significantly outperforms 6 state-of-the-art baselines. Case studies further highlight the model’s interpretability and provide mechanistic insights into herbal treatment.
  •  

Efficient Sleep Staging With Bayesian Uncertainty-Guided Active Learning

Automated sleep staging is essential for large-scale and home-based sleep monitoring; however, in routine clinical practice, sleep annotation remains largely dependent on experienced experts performing time-consuming and labor-intensive manual scoring. Existing automatic systems often struggle to adapt reliably to new subjects, limiting their clinical adoption and reinforcing the reliance on expert review. This creates a strong demand for adaptive and efficient sleep staging systems that can substantially reduce annotation workload while preserving expert-level accuracy. We propose BayesSleepNet, a novel framework that integrates Bayesian uncertainty quantification with active learning for adaptive sleep staging. BayesSleepNet employs principled Bayesian modeling by placing distributions over network weights and performing Monte Carlo sampling at inference, enabling explicit quantification of model (epistemic) uncertainty. These uncertainty estimates drive a two-stage sample selection strategy that first fine-tunes the model using representative epochs and subsequently prioritizes persistently uncertain samples for expert review. Across four public sleep datasets, BayesSleepNet consistently improves performance—by 7.60% in accuracy, 8.27% in macro-F1, and 0.104 in Cohen's $\kappa$—while requiring manual annotation of only 20% of data from new subjects. Despite its adaptive learning capability, BayesSleepNet remains computationally lightweight, using substantially fewer parameters than representative high-capacity state-of-the-art models. These results demonstrate the clinical promise of uncertainty-aware active learning as a practical and cost-efficient paradigm for semi-automated sleep staging.
  •  

A Lightweight Curriculum and Contrastive Learning Framework for Protein—Protein Interaction Prediction

Protein—protein interaction (PPI) prediction is essential for understanding cellular functions and enabling applications in drug development and disease research. PPI networks exhibit multi-scale structural and semantic heterogeneity and introduce representational bias under topological imbalances, leading to insufficient exploitation of subgraph-level semantic complexity. Moreover, many existing approaches rely heavily on external data during modeling, resulting in high computational costs for large-scale inferences. In this paper, we propose JCCLPPI, a joint curriculum- and contrastive-learning framework for lightweight PPI prediction. First, we implement a PPI network-structure encoding module designed to mitigate topological bias and learn biologically interpretable representations without relying on handcrafted features or external annotations. Next, we propose a motif-based curriculum learning module that incrementally introduces training samples according to their structural–semantic complexity, thereby enhancing model robustness to long-tail distributions and structural heterogeneity. Finally, our approach incorporates two graph neural network–based modules during PPI inference to perform local structural modeling and global context encoding, facilitating multi-scale feature extraction. Experiments conducted on two widely used human PPI benchmark datasets, SHS27k and SHS148k, demonstrate that JCCLPPI improves model generalization. It achieves an approximate 4% increase in micro-F1 score compared to state-of-the-art methods, while also improving computational efficiency by reducing memory consumption by 76% and inference time by 35%. Furthermore, JCCLPPI provides a scalable basis for therapeutic target prioritization and early-stage drug discovery.
  •  

Adaptive Feature Selection With Hierarchical Learning for Drug-Target Interaction Prediction

Accurate prediction of drug–target interactions (DTIs) is essential for drug discovery and repurposing. Although deep learning has driven substantial progress, critical limitations remain: a singular focus on intermolecular associations results in suboptimal representation learning, and the failure to leverage key features during interactions constrains further performance gains. Here, we propose ASHL-DTI, a novel framework that integrates hierarchical learning with adaptive feature selection to significantly boost both feature quality and model generalizability. Specifically, the hierarchical learning component captures multi-level intramolecular associations to learn more discriminative representations. Simultaneously, we incorporate an adaptive Top-k selection mechanism to retain the most predictive features, facilitating effective interaction between drugs and targets. Experimental results across multiple public benchmark datasets demonstrate that ASHL-DTI achieves superior performance compared with state-of-the-art approaches. Moreover, ASHL-DTI exhibits strong generalization ability in predicting novel drug–target pairs, underscoring its potential in drug discovery.
  •  

Development of a Multimodal Obstructive Sleep Apnea Diagnostic Prediction Model Using Two-Dimensional Facial Images and Clinical Data

Obstructive sleep apnea (OSA) is a prevalent disorder in middle-aged and obese men with increased risk of cardiovascular disease, metabolic dysfunction, and neurocognitive impairments. Delays in diagnosis and treatment of patients with OSA increase long-term all-cause mortality. In this study, we developed a multimodal artificial intelligence (AI) model that utilizes two-dimensional facial photograph, cephalometric radiograph, and clinical data to enhance OSA screening. Clinical parameters of sex, age, body mass index, neck circumference, and abdominal circumference and sleep questionnaires were assessed. Data from 710 patients who underwent polysomnography were used to train and validate a deep learning model combining ShuffleNet-V2 for image feature extraction and a deep neural network for clinical data analysis. The model was evaluated using five-fold cross-validation and a holdout test set. Performance metrics were area under the curve (AUC), sensitivity, specificity, accuracy, and F1 score. The trimodal AI model significantly outperformed uni- and bimodal approaches, achieving an AUC of 0.859 in distinguishing moderate-severe OSA from normal-mild OSA. Subgroup analyses of data from obese, elderly, and male patients showed higher classification accuracy, and data from patients with smaller abdominal circumferences had the lowest sensitivity compared to other subgroups. The Grad-CAM analysis demonstrated that the model focused on airway structures and low part of face, aligning with clinical expectations. This study presents a novel AI-driven OSA screening approach that integrates data from facial images and clinical data. This AI-based OSA screening model may facilitate early diagnosis of OSA and improve patient outcomes.
  •  

MPAExpo-LM: Fine-Tuned Large Language Model for Mycophenolic Acid Exposure Estimation After Renal Transplantation

An accurate estimation of mycophenolic acid (MPA) exposure after renal transplantation is important for reducing acute rejection in patients treated with mycophenolate mofetil (MMF) or enteric-coated mycophenolate sodium (MPS). We propose MPAExpo-LM, a fine-tuned large language model with a differential attention denoising module for this task. Experiments on real-world datasets demonstrate that MPAExpo-LM achieves superior results compared to existing methods in the task, offering a new pathway for precise immunosuppressive therapy. It attains an RMSE of 8.78 mg·h/L using three sampling points (C2, C4, C8) and 7.43 mg·h/L with four points (C0.5, C2, C6, C8). Furthermore, we conducted additional fine-tuning by incorporating a limited amount of data from external data source from another hospital. In the subsequent multi-center validation, the model demonstrated robust generalization, achieving RMSEs of 7.342 mg·h/L (three points) and 7.281 mg·h/L (four points). Via reinforcement fine-tuning, the model demonstrates strong generalization capability in cross-hospital tests.
  •  

Wavelet-Driven Spatial Frequency Mamba Network for Spine Image Segmentation

Accurate spine segmentation is crucial for diagnosing and treating various spine diseases. Recently, Mamba-based methods have been widely applied in medical image segmentation. However, the spatial domain scanning strategy of Mamba fails to fully capture the fine anatomical structures and global dependencies of the spine. Moreover, existing frequency-enhanced methods often suffer from the loss of spatial localization. To address these challenges, we propose the Wavelet-Driven Spatial Frequency Mamba Network (WDSFM-Net). Specifically, we integrate the Discrete Wavelet Transform (DWT) into Mamba to construct the Spatial-Frequency Mamba Block (SFMB). By decomposing features into distinct frequency subbands, SFMB explicitly captures global structural context from low-frequency components while enhancing local anatomical details through high-frequency components. To accommodate specific spinal morphology, the Global Strip Pooling Attention (GSPA) module aggregates directional contexts to model the elongated and anisotropic spinal anatomy, while the Multi-Scale Attention Enhancement (MSAE) module employs multi-scale convolutions to adapt to significant vertebral scale variations. Additionally, we introduce a Dual-Domain Loss (DDL) function, which optimizes both spatial and frequency domain representations for robust training. We evaluated our WDSFM-Net on two public spine MRI datasets. The results show that the WDSFM-Net outperforms other state-of-the-art methods, achieving average Dice similarity coefficients of 0.8885 and 0.8669 in the Spider and MRSpine datasets, respectively.
  •  

Automatic Segmentation of Placenta From MR Images Using a Novel BiGC U-Net

Accurate segmentation of the placenta in Magnetic Resonance (MR) images is required for quantitative techniques such as texture and shape analysis, which have been proposed to improve placenta accreta spectrum (PAS) diagnostic rates. However, it is challenging due to the low contrast, image noise, blurred boundaries, and manual annotation variability. Hence, we proposed an enhanced U-Net architecture named BiGC U-Net, to automatically segment placental MR images. This deep learning (DL) structure incorporates a bidirectional gated convolutional module (BiGC) to capture complementary spatial dependencies, a hierarchical regularization mechanism (HRM) to enhance crosslayer semantic consistency, and an innovative data augmentation strategy to synthesize new images. The performance of BiGC U-Net was evaluated on three placental MR datasets: (1) public, (2) Sheffield Teaching Hospitals (STH) and (3) combined (public + STH + augmented), and compared against existing DL models including U-Net, Attention U-Net, ResNet, UNet++, TransUNet, nnUNet, and SSM-Mamba. The BiGC U-Net exhibited the best Dice similarity coefficient (90.74 ± 0.44), 95th percentile Hausdorff distance (3.84 mm ± 0.53 mm) and relative volume difference (9.06 ± 0.41) in comparison with other DL models in the combined dataset. These findings indicate the effectiveness and robustness of the BiGC U-Net in accurately and automatically segmenting the placenta.
  •  

Explainable Convolutional Channel Ranking (ECCR) for EEG-Based Detection of Idiopathic Absence Seizures

Accurate identification of EEG electrodes associated with epilepsy is essential for developing real-time diagnostic applications. This paper introduces the Explainable Convolutional Channel Ranking (ECCR) method for identifying diagnostically relevant EEG channels for Idiopathic Absence Seizure (IAS) detection by analyzing channel-specific feature contributions learned by a convolutional neural network (CNN). Unlike traditional saliency-based approaches that focus only on highly activated regions or pool contributions across seizure types and spatial locations, ECCR retains channel-specific contribution patterns and shows that channels with moderate contribution levels offer the most discriminative and physiologically consistent information. This finding suggests that channels with very high saliency are often affected by noise or subject-specific artifacts, while medium-contribution channels capture more stable seizure-related information dynamics. In 10-fold cross-validation, the ECCR-guided CNN achieved 82.21% accuracy and 92.01% sensitivity, while leave-one-subject-out (LOSO) validation yielded 73.78% accuracy, demonstrating improved subject-independent performance under a leakage-controlled protocol; ECCR consistently selected fronto–central, temporo–parietal, and occipital regions, reducing 29 channels to 7 in the subject-dependent evaluation. A validation using a Random Forest classifier confirmed that ECCR-selected channels provided stronger detection power than those excluded. These findings suggest that ECCR can guide the design of compact, interpretable EEG systems, supporting more reliable deep learning solutions for IAS diagnosis.
  •  

Reinforcement Learning-Enhanced Dual-View GAT-Based Multi-Task Learning for Non-Coding RNA-Disease Association Prediction

Non-coding RNAs (ncRNAs), particularly long non-coding RNAs (lncRNAs) and microRNAs (miRNAs), are important regulators of gene expression and are closely involved in disease pathogenesis. Therefore, identifying ncRNA–disease associations is essential for clarifying disease mechanisms. Because lncRNAs and miRNAs frequently regulate each other in cells, predicting miRNA-disease associations (MDA) and lncRNA-disease associations (LDA) is inherently interconnected. However, many existing computational approaches still treat these tasks as independent problems, which limits their ability to capture cross-task biological signals. In addition, most models use fixed hyperparameters, which may be suboptimal as training dynamics and data characteristics change. To address these issues, we propose RL-DMGLMD (Reinforcement Learning-enhanced Dual-view Multi-task Graph learning for LncRNA-MiRNA-Disease association prediction). RL-DMGLMD contains three key components: (1) a Soft Actor-Critic (SAC) controller that adaptively tunes hyperparameters during training by monitoring loss, validation performance, and gradient information; (2) a unified multi-task framework that jointly predicts LDA, MDA, and lncRNA-miRNA interactions (LMI) using shared encoders with task-specific decoders to enable knowledge transfer; and (3) a dual-view multi-head Graph Attention Network (GAT) that learns from both heterogeneous interaction graphs and attribute graphs to capture relation-specific importance. RL-DMGLMD achieves AUROC values of 0.9900 / 0.9872 / 0.9867 on Dataset 1 and 0.9946 / 0.9903 / 0.9954 on Dataset 2 for LDA, MDA, and LMI, respectively. These results outperform state-of-the-art baselines and support RL-DMGLMD as a practical tool for biomarker discovery and therapeutic target prioritization.
  •  

Meta-Entity Driven Triplet Mining for Aligning Medical Vision-Language Models

Imaging-based diagnostics rely on evaluation of both medical images and radiology reports, but increasing data volumes strain medical experts, leading to errors and workflow delays. Medical vision-language models (med-VLMs) offer an efficient approach for processing multimodal imaging data, especially for chest X-rays (CXRs), though their success depends on effective image-text alignment. Existing alignment methods for med-VLMs, primarily based on contrastive learning, often focus on coarse-grained disease class separation and overlook fine-grained pathology attributes such as location, size, or severity, which results in suboptimal representations. We introduce MedTrim (Meta-entity-driven Triplet mining), a novel alignment method that improves precision via structured triplet learning guided by meta-entities extracted from radiology reports. Unlike conventional contrastive and triplet frameworks for representational learning that rely on global class labels or implicit similarity references, MedTrim explicitly models hierarchical relationships between pathology attributes to preserve clinically meaningful intra-class variation. To do this, MedTrim leverages a domain-specific ontology to identify adjectival qualifiers and directional descriptors of pathology, a novel entity-aware triplet mining score to capture hierarchical inter-sample similarity, and a multimodal alignment objective that enforces consistency across image-text pairs sharing detailed pathology attributes without compromising within-modality relationships. MedTrim improves performance in downstream retrieval, classification, and generation tasks compared to leading alignment methods.
  •  

Waveformer: Dual-Branch Adaptive Network With Wavelet-Guided Cross-Context Decoding for Colorectal Polyp Segmentation

With the advancement of deep learning, polyp segmentation in endoscopic images has achieved remarkable progress. However, clinical polyps often exhibit variable morphology, blurred boundaries, and low contrast with the intestinal mucosa, hindering accurate lesion localization and edge delineation. Moreover, complex conditions of low light, luminal distortion, and mucosal folds further exacerbate the problem with identification, resulting in frequent misdetections and omissions in computer-aided diagnosis. Accordingly, we propose Waveformer, a local-global co-modeling segmentation network, to improve segmentation accuracy. Concretely, the encoder employs parallel CNN-Transformer branches to synergistically extract detailed and global features, thereby enhancing the completeness and discriminative power of the representation. The decoder integrates a wavelet-based frequency decomposition unit (WFDU), a camouflage identification module (CIM), and an information fusion layer (IFL). These modules collaboratively enhance edge responses and semantic aggregation across scales, significantly boosting the framework’s capability in boundary modeling and lesion discernment. Extensive experiments on CVC-ClinicDB and Kvasir-SEG datasets achieve Dice Similarity Coefficients (DSC) of 95.60% and 94.11%, outperforming fourteen state-of-the-art (SOTA) methods. Cross-dataset evaluations further verify its strong generalization ability, with DSC scores of 81.0% and 79.2%, respectively.
  •  

Awareness- and Arousal-Specific Brain Analysis Predicts Spinal Cord Stimulation Effectiveness in Disorders of Consciousness

Spinal cord stimulation (SCS) is an advanced treatment for disorders of consciousness (DoC), but its success rate varies between 30 and 60%, and consciousness-related biomarkers are urgently needed for SCS assessment and optimization. This paper proposes an awareness- and arousal-specific brain analysis (AAA) method to encode these two consciousness dimensions and construct consciousness-related features to predict SCS outcomes of DoC patients. Firstly, electroencephalogram (EEG) brain signals were collected from twenty-eight DoC patients during SCS treatment. Then, dynamic brain networks are formulated based on sliding-window correlation analysis of weighted phase lag index. Afterwards, a hierarchical network decomposition algorithm was developed to resolve the dynamic networks into consciousness-related networks by elaborating data-driven optimization of non-negative matrix factorization and consciousness-related variability. Further, consciousness features are designed to quantify the activation, interaction, and stability of awareness- and arousal-specific networks, and a support vector machine is trained for classification of SCS treatment effectiveness. Clinical results showed that our method achieved an overall prediction accuracy of 88%, which is significantly better than clinical accuracy of 50% by doctors. Moreover, our method outperformed existing EEG-based approaches in both prediction accuracy and pathology interpretability. The proposed awareness- and arousal-specific brain analysis method establishes a pivotal framework for precise and reliable SCS treatment of DoC.
  •  

TMN-LAKDE: Characterizing Sleep Instability via Prediction Intervals of Dynamic EEG Spectral Networks

Sleep instability is a typical characteristic of insomnia, manifested as the inability of the brain to maintain a stable state, but its precise quantification is still challenging. We assume that sleep instability fundamentally reflects an increase in unpredictability in the evolution of brain network dynamics. To verify this, an interval prediction framework combining the temporal mobile network (TMN) and local adaptive kernel density estimation (LAKDE) is proposed to characterize the sleep instability. Specifically, TMN predicts future network states, while LAKDE module quantifies the uncertainty of these predictions by generating prediction intervals (PIs). Experiments on SIESTA and Sleep-EDF databases have shown that this method can construct well calibrated PIs. The key finding is that the PI normalized average width (PINAW) of subjects with sleep disorders is significantly higher than that of the healthy control group, validating that wider PIs are a mechanistic biomarker of sleep instability. In addition, this study further revealed a significant correlation between PINAW and traditional indicators such as number of sleep stage transitions, indicating that dynamic instability based on prediction uncertainty shares a common physiological basis with sleep fragmentation phenomena, establishing interval prediction as a paradigm for quantifying sleep instability.
  •  

CFRAFN: A Cross-Feature Residual Attention Fusion Network for Major Depressive Disorder Prediction Using Clinical Voice Recordings

Major depressive disorder (MDD) is a prevalent mental disorder with a significant burden on individuals and society, and timely identification and intervention are essential for effective management. Voice data have been used as behavioral indicators of MDD, offering valuable insights into an individual's mental state. In this study, we collected voice data from 221 patients diagnosed with MDD at the inpatient ward of the Department of Psychiatry and Psychosomatics, Zhongda Hospital, Southeast University, alongside 113 healthy controls, to construct the Chinese depressive voice dataset. We proposed the cross-feature residual attention fusion network (CFRAFN), which leverages extended Geneva minimalistic acoustic parameter set features along with high-dimensional embeddings extracted from the pretrained VGGish model to effectively capture MDD-associated phonetic patterns. Specifically, CFRAFN utilizes differentiated residual blocks to maintain training stability in deep hierarchical structure. Furthermore, the self-attention fusion strategy dynamically weighted the significance of each feature modality, ensuring effective feature integration and consequently improving MDD prediction accuracy. Experimental results demonstrated that CFRAFN achieved an excellent predictive performance with an area under the receiver operating characteristic curve of 0.924 in an independent test set, and significantly outperformed 11 baseline models across 5-fold cross-validation.
  •  

Legal and Ethical Considerations for Translating Federated Learning Into Cross-Border Healthcare Innovation

Federated Learning (FL) offers a privacy-enhancing architecture for training artificial intelligence on decentralized healthcare data, yet the prevailing mantra to “move the model, not the data” obscures significant privacy risks, ethical dilemmas, and regulatory conflicts. This article challenges the assumption that FL inherently solves cross-border compliance by analyzing a hypothetical consortium involving the United States, the U.K., EU, China, and Brazil. We identify a specific compliance deadlock arising from the friction between Western rights-based frameworks (HIPAA, GDPR, LGPD) and state-centric security models (China's PIPL/DSL), particularly regarding model inversion attacks and data localization. Moving beyond validatory analysis, we propose a multi-layered Federated Governance Framework to operationalize data diplomacy. This contribution introduces novel legal and structural mechanisms: the Federated Data Sharing & Use Agreement (F-DSA) to codify binding mandates like privacy-enhancing technologies and isolated training; Governance as a Service (GaaS) to neutralize conflicts of interest through third-party administration; and relational mechanisms, including blockchain-based dynamic consent to solve the cascade of consent problem. This framework provides a blueprint for converting theoretical FL potential into a legally compliant, ethically sound, and functioning international Learning Health System (LHS).
  •  

MedSegAgent: A Universal and Scalable Multi-Agent System for Instructive Medical Image Segmentation

Medical image segmentation is vital for clinical diagnosis and treatment; however, current solutions face three major limitations: (1) the lack of a universal framework capable of handling diverse modalities and anatomical targets, (2) the limited scalability to adapt to evolving clinical needs and new datasets, and (3) the lack of instructive interfaces that make models usable for non-expert users. To address these challenges, this paper presents MedSegAgent, a universal and scalable multi-agent system for instructive medical image segmentation. Specifically, MedSegAgent comprises five agents: one query parsing agent that processes natural language requests, three coarse-to-fine filtering agents (modality filtering, anatomical filtering, and label selection) for identifying relevant datasets and label values, and one execution agent responsible for model inference and result integration. Based on this framework, MedSegAgent utilizes 23 diverse datasets and pre-trained models to perform 343 types of segmentation across various modalities and anatomical targets. Experimental results demonstrate that MedSegAgent simplifies model selection while maintaining high performance, accurately identifying matching datasets and labels in 94.27% of queries and locating at least one suitable match in 99.03% of queries. MedSegAgent offers a universal and scalable solution for diverse medical image segmentation tasks, bridging the gap between user-friendly queries and the complexities of model selection and deployment.
  •  

Rethinking Feature Interactions for Medical Image Segmentation: A Unified Hierarchical Aggregation Framework With Boundary Guidance

Medical image segmentation is a crucial task of medical image analysis and computer vision. Medical images, compared to natural ones, contain more complex semantic information, making feature learning more challenging. Existing encoder-decoder architectures are limited by inadequate cross-scale interaction and insufficient boundary modeling in their feature fusion designs. To address this, we propose a Hierarchical Feature Interaction network with Boundary guidance (HFIBNet), which unifies dynamic cross-level feature fusion and explicit edge supervision within a coarse-to-fine segmentation framework. Specifically, we introduce a Boundary Prediction (BP) module to extract boundary-aware features that guide the fusion process. A Cross-Level Feature Fusion (CLFF) module is designed to promote semantic interaction across adjacent encoder stages, while the Edge Feature Aggregation (EFA) module propagates boundary cues hierarchically to enhance structural consistency. Furthermore, a Partially Parallel Decoder (PPD) generates a coarse global prediction, which is progressively refined by a Global-Local Feature Enrichment (GLFE) module, mimicking the clinical annotation workflow from coarse localization to fine delineation. Extensive experiments on ten public medical segmentation datasets across four distinct tasks demonstrate that HFIBNet consistently outperforms existing state-of-the-art methods.
  •  

A Knowledge-Guided Bi-Modal Network for the Classification of Anterior Chamber Angle Images

Glaucoma is a leading cause of irreversible blindness globally. When glaucoma is diagnosed, Anterior Chamber Angle (ACA) evaluation is the necessary step for the prognosis and treatment of glaucoma. However, current clinical evaluation methods are labor intensive and rely on expert judgment, which makes them inefficient. Automating ACA classification based on images using machine learning, especially deep neural networks, holds promise. Yet, image samples alone can’t provide sufficient high-level semantic information on ACA, resulting in suboptimal classification performance. This paper proposes a novel end-to-end knowledge-guided bi-modal network (KGNet) for ACA evaluation. Specifically, we consider two modalities of ACA data: textual domain knowledge and images. We first design a new strategy to refine class-based knowledge into textual descriptions, thereby increasing the diversity of features learned by the model. We then extract two types of representations using two distinct components: 1) a supervised loss is applied to learn modality-specific representations by incorporating domain knowledge; 2) a fusion module that uses knowledge-guided learning to highlight key clinical structures in ACA images leveraging bimodal correlations. Experimental results on an ACA dataset and three public datasets show that our method outperforms several state-of-the-art deep learning models in eye image evaluation, indicating the potential medical interest of our method. Furthermore, our approach improves interpretability by explicitly aligning visual representations with structured clinical knowledge, enabling more structured and clinically grounded explanations than conventional models.
  •  

Pseudo Anomalies and Hard Sample Mining for Ventricular Arrhythmia Anomaly Detection

Ventricular arrhythmias (VA) are among the most prevalent and clinically significant cardiac arrhythmias. Conventional detection methodologies predominantly employ supervised learning approaches that depend on precisely annotated training datasets. However, the morphological similarity between VA waveforms and noise artifacts poses a significant challenge for traditional algorithms in discriminating these clinically distinct categories. In this paper, we propose a novel anomaly detection framework named PHVA (Pseudo-data and Hard-sample mining for VA detection) based on one-class anomaly detection, where only normal ECGs are used for training, while pseudo-anomalies are generated through self-supervised modules. These modules, including pseudo-anomaly generation and hard-sample mining, provide supervisory signals without requiring abnormal labels. Specifically, we introduce a physiology-aware synthetic ECG generation method that captures the beat morphology and waveform characteristics of VA while incorporating realistic noise simulations based on conventional noise models. These pseudo-annotated signals are then used to refine the decision boundary of a time-frequency hypersphere, constructed from normal ECG features in both temporal and spectral domains. Additionally, we employ triplet-loss-based hard-sample mining to improve the model's discriminative power for ventricular fibrillation detection. Extensive experiments on three public ECG datasets demonstrate that our proposed method PHVA achieves superior overall performance, outperforming state-of-the-art anomaly detection baselines by up to 8.7% in AUC (the Area Under a Receiver Operating Characteristic Curve).
  •  
❌