❌

Normal view

Received β€” 13 April 2026 ⏭ IEEE Journal of Biomedical and Health Informatics - new TOC

Text-Driven Weakly Supervised OCT Lesion Segmentation With Structural Guidance

Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of supervised learning for large datasets. Weakly Supervised Semantic Segmentation (WSSS) offers a promising alternative by using weaker forms of supervision, such as image-level labels, to reduce the annotation burden. Despite its advantages, weak supervision inherently carries limited information. We propose a novel WSSS framework with only image-level labels for OCT lesion segmentation that integrates structural and text-driven guidance to produce high-quality, pixel-level pseudo labels. The framework employs two visual processing modules: one that processes the original OCT images and another that operates on layer segmentations augmented with anomalous signals, enabling the model to associate lesions with their corresponding anatomical layers. Complementing these visual cues, we leverage large-scale pretrained models to provide two forms of textual guidance: label-derived descriptions that encode local semantics, and domain-agnostic synthetic descriptions that, although expressed in natural image terms, capture spatial and relational semantics useful for generating globally consistent representations. By fusing these visual and textual features in a multi-modal framework, our method aligns semantic meaning with structural relevance, thereby improving lesion localization and segmentation performance. Experiments on three OCT datasets demonstrate state-of-the-art results, highlighting its potential to advance diagnostic accuracy and efficiency in medical imaging.
Received β€” 8 April 2026 ⏭ IEEE Journal of Biomedical and Health Informatics - new TOC

ZhiFangDanTai: Fine-Tuning Graph-Based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula

Traditional Chinese Medicine (TCM) formulas play a significant role in treating epidemics and complex diseases. Existing models for TCM utilize traditional algorithms or deep learning techniques to analyze formula relationships, yet lack comprehensive results, such as complete formula compositions and detailed explanations. Although recent efforts have used TCM instruction datasets to fine-tune Large Language Models (LLMs) for explainable formula generation, existing datasets lack sufficient details, such as the roles of the formula’s sovereign, minister, assistant, courier; efficacy; contraindications; tongue and pulse diagnosisβ€”β€”limiting the depth of model outputs. To address these challenges, we propose ZhiFangDanTai, a framework combining Graph-based Retrieval-Augmented Generation (GraphRAG) with LLM fine-tuning. ZhiFangDanTai uses GraphRAG to retrieve and synthesize structured TCM knowledge into concise summaries, while also constructing an enhanced instruction dataset to improve LLMs’ ability to integrate retrieved information. Furthermore, we provide novel theoretical proofs demonstrating that integrating GraphRAG with fine-tuning techniques can reduce generalization error and hallucination rates in the TCM formula task. Experimental results on both collected and clinical datasets demonstrate that ZhiFangDanTai achieves significant improvements over state-of-the-art models.

A Local-Global Multi-View Diffusion Variational Graph Auto-Encoder for lncRNA-Protein Interaction Prediction

Long non-coding RNAs (lncRNAs) interact with proteins, influencing cell growth, differentiation, and disease onset. Despite significant advancements in computational methods, current approaches rely heavily on manually engineered features and require improved feature fusion techniques. Furthermore, prior studies have predominantly utilized supervised and semi-supervised learning techniques, which fail to effectively harness limited data from various sources, significantly constraining their generalizability and performance across diverse scenarios. Additionally, existing variational graph auto-encoders (VGAE) do not adequately capture long-range interactions of biomolecules. Therefore, this study introduces the Local-Global Multi-View Diffusion Variational Graph Auto-encoder (LG-MDVGA) for predicting lncRNA-protein interactions (LPIs). LG-MDVGA integrates a feature construction and fusion module that creates parameterized feature matrices for lncRNAs and proteins, which are updated through backpropagation. To capture local features effectively, separate adaptive local multi-modal feature matrices for lncRNAs and proteins are constructed. To fully utilize limited data to capture global data features and enhance predictive accuracy and generalization, LG-MDVGA incorporates a global multi-space collaborative computation by self-supervised learning. In addition, it introduces a diffusion variational graph auto-encoder (DVGA) to address the limitation that traditional VGAE have difficulty in capturing the complex patterns and relationships of LPIs. Experimental results show that LG-MDVGA significantly outperforms current methods and holds potential for discovering new LPIs. Additionally, LG-MDVGA was tested on five datasets involving three other types of biological entities and consistently attained superior performance. This underscores the generalizability and high precision of LG-MDVGA in accurately predicting associations among biological entities.
  • βœ‡IEEE Journal of Biomedical and Health Informatics - new TOC
  • Predicting the Effort Required to Manually Mend Auto-Segmentations
    Auto-segmentation quality or accuracy influences their clinical usefulness. However, currently widely utilized segmentation metrics (e.g., Dice Coefficient (DC) and Hausdorff Distance (HD)) cannot effectively express the manual mending effort required when utilizing auto-segmentation results in clinical practice. In this article, we explore ways of evaluating auto-segmentations with clinical efficiency considerations in mind. The time required for correcting auto-segmentations by experts is reco
     

Predicting the Effort Required to Manually Mend Auto-Segmentations

Auto-segmentation quality or accuracy influences their clinical usefulness. However, currently widely utilized segmentation metrics (e.g., Dice Coefficient (DC) and Hausdorff Distance (HD)) cannot effectively express the manual mending effort required when utilizing auto-segmentation results in clinical practice. In this article, we explore ways of evaluating auto-segmentations with clinical efficiency considerations in mind. The time required for correcting auto-segmentations by experts is recorded to indicate ground-truth mending effort. Extended from our previous work, five explicitly-defined metrics are studied in detail for their ability to predict mending effort. More importantly, we explore the use of deep learning networks to provide an implicit metric, which predict mending effort using auto-segmentation masks and original images as input. A 3-institution evaluation is conducted with 7 different anatomic organs in the setting of auto-contouring for radiation therapy planning. Among the five explicit metrics, one form of the proposed Mendability Index (MIhd) shows the best performance to indicate the mending effort for sparse objects with 6.2-14.4% error, while one form of HD (sHD) performs best when assessing large non-sparse objects. Interestingly, while the explicit metrics all require ground truth segmentations for estimating mending effort, the implicit models obtained via deep learning are effective in predicting mending efforts (with 2.9-12.9% error) without the need for ground-truth segmentations and directly from the given image plus the auto-segmentations. We conclude that once effort-predicting deep models are created, it is feasible to assess the clinical usability of new segmentation models, going beyond bench technical evaluation commonly done via explicit metrics.

A Self-Supervised Diffusion Model With Edge Prior for Unpaired LDCT Denoising

Low-dose computed tomography (LDCT) reduces health risks from radiation exposure but introduces imaging noise and artifacts. While numerous studies have employed deep learning for LDCT image denoising, the field continues to face significant challenges. Recent advancements have seen diffusion models applied to overcome issues of over-smoothness and unstable training inherent in prior deep learning approaches. However, the diffusion models face challenges in direct practical applications due to the extensive sampling steps, significant inference time required, and the need for hard-to-obtain paired data during training. To address these difficulties, this paper introduces a self-supervised diffusion model with edge prior for unpaired LDCT denoising. This method enables denoising within a lower-dimensional space, reducing computational complexity. Our proposed approach enhances denoised image clarity by applying prior edge constraints to compressed encodings; it employs a noise-conditioned encoding strategy to facilitate self-supervised image training, enabling the method to be applicable to unpaired CT data; and it utilizes compressed LDCT encoding as intermediate sampling results during the inference process, thereby accelerating sampling and reducing the time required for inference, making the method more real-time capable. Extensive validation across multiple datasets demonstrates that our method achieves competitive performance against state-of-the-art approaches in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and perceptual quality (LPIPS), while maintaining a practically acceptable inference time.

EAP-LSTM: A Bi-LSTM-Based Deep Learning Framework for Quantitatively Predicting Enhancer Activity in Drosophila and Human Cell Lines

Enhancer activity plays a critical role in gene regulation, influencing various biological processes such as development and disease progression. Accurate prediction of enhancer activity is essential for understanding the mechanisms underlying gene regulation and enhancer function. This study introduces a novel deep learning framework, EAP-LSTM (Enhancer Activity Prediction based on Bi-LSTM), to quantitatively predict enhancer activity across different species and cell lines. The model integrates multiple feature modules, including Word2Vec-based representations of DNA sequences, reverse complement k-mer, mismatch k-mer features, and epigenomic data. Evaluated on six cell lines, including five human cell lines (A549, HCT116, HepG2, K562, and MCF-7) and one Drosophila cell line (S2), EAP-LSTM consistently outperforms state-of-the-art models, such as DeepSTARR and HEAP, in all datasets. For example, on the K562 dataset, EAP-LSTM achieves a Pearson correlation coefficient (PCC) of 0.7944, outperforming DeepSTARR and HEAP by 13.65% and 2.73%, respectively. In addition, EAP-LSTM demonstrates strong performance in small-sample learning scenarios, showing clear improvements compared with baseline models. Furthermore, the study investigates the role of transcription factor binding sites (TFBSs) within enhancer regions, identifying critical motifs associated with enhancer activity. These findings not only improve enhancer prediction accuracy but also provide valuable insights into the molecular mechanisms underlying enhancer function.

P3DL: A Privacy Preserving Personalized Distributed Learning Framework for EEG-Based Cognitive State Identification

Electroencephalography (EEG)-based brain cognitive state identification for the elderly allows timely detection and early intervention of cognitive deterioration. Notably, EEG signals carry a great deal of vital personal information. However, a majority of the existing cognitive evaluations focus on improving the accuracy of EEG decoding and enhancing the performance of identification models, while neglecting the privacy protection of EEG data. To address the risky challenge, we propose a privacy-preserving personalized distributed learning framework (P3DL) for cognitive state identification. Specifically, it consists of the clients and a central server. Each client contains a cognitive model and a score model for identifying cognitive states and quantifying cognitive levels, respectively. The central server can aggregate local models’ parameters from distributed clients, then, update and downstream the global model’s parameters for iterative optimization. A federated dynamic update strategy (FedDBS) is designed to jointly update all global and local models with a supervisory metric. In order to further improve the identification performance and judge the misdiagnosis level, a novel loss function, extreme error Loss (E2Loss), is proposed. Compared with the baseline, experimental results on our self-collected clinical dataset and a public dataset show an average increase in F2Score of 5.58% and 3.31%, and in accuracy of 1.78% and 2.46%, respectively. Furthermore, the scalability of the framework has been proved in the emotion recognition task. Our proposed framework P3DL can not only improve the identification performance, but also protect the privacy of EEG, opening a new window for secure healthcare.

ICD-10 Neoplasm Location Using Text Classification Models in Spanish Electronic Health Records

The majority of clinical information stored in Spanish healthcare systems is found as unstructured text in electronic health records (EHRs). The automatic extraction of valuable information contained in these documents is a critical task. Valuable information for oncology clinical analysis units and Real-World Evidence studies includes the location of the neoplasm presented by a patient. This location, included in the ICD-10 coding category, can be extracted from the texts by natural language processing (NLP). This study set out to explore the classification of medical documents in Spanish for the purpose of extracting the location of the patient’s primary neoplasm. A private corpus composed of 23,704 real clinical EHRs was utilised. The prediction problem was approached through a classification of 12 primary organ groupings and 29 specific locations. In order to achieve this, four NLP methodologies were developed: traditional machine learning (ML); ensemble ML; recurrent neural networks (RNN); and Transformers-based models. Our findings demonstrate that traditional ML models exhibit superior performance when compared to RNNs and Transformers. Models such as XGBoost and SVMs demonstrate remarkable efficacy, attaining an F1-score of 0.938 for the 12-class classification and an F1-score of 0.838 for the 29-class specific classification, respectively. The pre-trained RoBERTa-Base-Biomed model, which incorporates medical and clinical corpora in Spanish, demonstrates an F1-score of 0.808 for the 29-location problem. However, the Transformer model exhibits superior performance than the rest of the approaches when dealing with an external corpus, indicating a higher generalisation capacity.

Rethinking Propagation Methods for Interactive Medical Image Segmentation

Propagation-based methods have drawn increasing research attention in interactive medical image segmentation. However, existing propagation-based methods face two significant challenges: 1) Due tothe continuous nature of anatomical structures within the organs and tumors throughout the volume, over-propagation is likely to occur as the propagation process reaches the end of structures, leadingto a degradation in segmentation performance. 2) During the multi-round refinement process, selecting the worst-segmented slice for refinement tends to hinder the optimization of segmentation results. To overcome these challenges, we propose the Discrepancy Aware Network (DANet), which includes a Discrepancy Learning Module (DLM) and employs a confidence loss to achieve accurate segmentation. Specifically, DLM captures the temporal-contextual discrepancy between previous and current slices, enabling the model to perceive the variations of the target. Furthermore, the confidence loss is responsible for regularizing the over-confident segmentation at the image level by estimating the target foreground. Additionally, we design a straightforward slice selection strategy to optimize the refinement process. Extensive experimental results on five public medical datasets demonstrate significant improvements over state-of-the-art methods (e.g., with +1.07% improvement on the MSD-Spleen dataset).

Utilizing Multi-PPG-Sensor Site Information in a Localized Wrist Area for Improving Cuffless Blood Pressure Estimation

Blood pressure (BP) measurement accuracy is highly sensitive to sensor placement. To address this, we investigated the effect of multi-sensor sites on BP estimation using a 9-channel single-wavelength photoplethysmography (PPG) sensor array placed within a localized wrist area. As a starting point, we analyzed the variability of 9 PPG features across channels, revealing notable site-specific variations, with amplitude-based features showing greater sensitivity. Leveragingthese these findings, we developed 3M-BPNet, which processes PPG signals through varying channel/site counts. The network incorporates signal trimming and Bayesian optimization for channel-weight allocation, followed by a Random Forest Regression (RFR) model with personalized calibration. Tested on 121 subjects, the 3M-BPNet achieved mean absolute errors (MAEs) of 4.84 mmHg for systolic BP (SBP) and 3.28 mmHg for diastolic BP (DBP), outperforming a standard RFR model. Notably, BP estimation accuracy improved as the channel counts increased from 1 to 5, then declined beyond 5. The optimal 5-channel combination (C5-C6-C7-C8-C9), located near the radial artery, yielded MAEs of 1.33 mmHg for SBP and 1.16 mmHg for DBP, corresponding to accuracy gains of 72.6% for SBP and 73.2% for DBP over the single-channel setup (P

ACGM: Attribute-Centric Graph Modeling Network for Concurrent Missing Tabular Data Imputation and COVID-19 Prognosis

COVID-19 prognosis using clinical tabular data faces significant challenges due to missing values and class imbalance issues. Existing methods often overlook the complex high-order interrelationship among clinicalattributes and struggle with training stability on imbalanced datasets. We propose ACGM, an attribute-centric graph modeling network that simultaneously addresses missing data imputation and COVID-19 prognosis. ACGM consists of three key modules: an attributes preprocessing module (APM) for coarse-grained imputation initialization, a graph-enhanced attributes imputation module (GEAIM) that models high-order inter-attribute relationships through graph structures, and a graph-enhanced disease prognosis module (GEDPM) that leverages these complex attribute interactions for final prediction. GEAIM and GEDPM employ a mean-teacher strategy with attributes graph matching to preserve high-order relationships, enhance training stability, and maintain structural integrity of attribute interactions. Extensive experiments are conducted on four public COVID-19 tabular datasets, demonstrating the superiority of our ACGM over existing methods. Through comprehensive interpretability analysis, we identify that attributes such as LDH, Difficulty In Breathing, and SaO2 significantly impact COVID-19 prognosis, aligning well with clinical insights and radiologist assessments.

Hierarchical Deep Decision Tree-Based Network for Odontogenic Cystic Lesion Classification in CBCT Images

Odontogenic cystic lesions (OCLs) are complex jaw abnormalities that require a precise diagnosis of the disease for treatment. Visual OCL diagnosis is commonly based on reviewing cone-beam computed tomography (CBCT) to identify morpho-pathological features associated with specific lesion types in a hierarchical manner. Current state-of-the-art methods focus on extracting features from the image without any guidance beyond the lesion diagnosis, and do not fully leverage the hierarchical relationship between the lesion diagnosis and morphological features. In this study, we propose a hierarchical deep decision tree network (H2DT-Net) with three modules: a deep decision tree-based hierarchical learning module (DHLM) to leverage inter-categorical relationships; a feature category embedding module (FCEM) to capture representations from both diagnostic and morpho-pathological domains and support the DHLM; and a lesion localised attention module (LLAM) to facilitate the feature extraction process by generating lesion-focused attention maps. Evaluated on 289 CBCT images, H2DT-Net achieved state-of-the-art performance in OCL classification. We further demonstrate that our method is effective in clinical settings, where it outperformed six maxillofacial clinicians in diagnostic assessment.

Joint Dynamic Brain Network Estimation and Graph Representation Learning for the Recognition of Neurological Disorders

Recently, Graph Neural Networks (GNNs) have shown significant improvements in the recognition of neurological disorders by incorporating brain networks/graphs. However, most existing approaches have three main limitations. First, these methodologies rely on precomputed brain networks as input, typically derived from statistical metrics (e.g., Pearson correlation), which are inherently not learnable. Second, methods often assume that the magnitude of the brain interactions remains constant across the whole scan duration. Third, representations produced by models often lack interpretability and robustness when applied across brain disorders. To address these limitations, we propose a novel model called the Effective Brain Inference Graph Neural Network (EBIGNN), which infers dynamic Effective Connectivity (dEC) to characterize brain networks trained with direct feedback from downstream tasks within a unified end-to-end framework. EBIGNN is highly flexible in learning the most relevant graph structures customized to the specific underlying brain condition. The proposed model offers strong interpretability, providing valuable insights into the temporal evolution and altered connectivity patterns essential for understanding brain disorders. The model is validated on three publicly available datasets, demonstrating superior performance compared to other state-of-the-art methods. Moreover, the findings are consistent with previous neuroimaging-derived evidence of biomarkers, underscoring the model’s robustness in clinical settings.

Enhancing Deep Learning Inference of Gene Regulatory Networks via Construction of Image Representation of Cell-Cell Interactions From scRNA-Seq Data

Understanding gene regulatory networks (GRNs) holds paramount importance for deciphering the intricate interplay among genes and their influence on biological processes and disease pathogenesis. The emergence of single-cell RNA sequencing (scRNA-seq) techniques has heralded a new era in GRN inference by capturing the nuanced heterogeneity and dynamic nature of gene expression at the single-cell level. However, extracting meaningful patterns from scRNA-seq measurements to infer GRNs poses significant challenges to existing methodologies due to the sheer scale and inherent complexity of the data. Here we propose a highly accurate and computationally efficient strategy for scRNA-seq-based GRN inference. Our approach leverages the underlying interactive relationships among the cells using state-of-the-art deep learning strategy. Specifically, a spatially semantic image representation, termed CelloGraph, is first introduced to portray the expressions of each gene across cells. The allocation of a cell to a spatial grid point of the CelloGraph is dictated by its interactions with other cells within the system, as determined by the maximization of system entropy of cell-cell interactions. Subsequently, the CelloGraphs of all pertinent genes are analyzed by using a customarily designed convolutional neural network (CNN) to discern discriminant patterns in the data and infer GRNs. The efficacy of the proposed approach is demonstrated through diverse real-world biomedical datasets. By harnessing the distinctive attributes of spatially semantic CelloGraphs and leveraging the unique pattern discovery capabilities of CNNs, our methodology paves the way for a deeper comprehension of the underlying mechanisms that govern gene expression and regulation. The proposed strategy not only overcomes challenges in scRNA-seq-based GRN inference but also promises to provide a more comprehensive understanding of intricate biological processes.

A 3D Edge-Attention Denoising Diffusion Network for Prostate Segmentation in Puncture Biopsy

Prostate cancer is the second most common cancer in men, and transrectal ultrasound (TRUS) guided biopsy is the standard method to diagnose prostate cancer. Accurate prostate segmentation in TRUS images is crucial for precise biopsy. Manual segmentation is laborious, while automated segmentation faces significant challenges due to the low signal-to-noise ratio, blurred boundaries, and presence of noise and artifacts. To address these issues, this paper proposes a 3D edge-attention denoising diffusion network, aiming to achieve high accuracy and generalizability for prostate segmentation in TRUS-guided biopsy. The proposed network incorporates an edge attention denoising U-Net (EAD U-Net) to extract and utilize desired edge information in TRUS images, improving the segmentation accuracy in challenging regions of the prostate. To reduce uncertainty and enhance network accuracy, we incorporate a Kalman fusion module, which utilizes the Kalman filter and all estimations from the EAD U-Net in reverse process to obtain the optimal segmentation estimation. The proposed network was evaluated using 1834 3D ultrasound images from two open-source datasets. Comparative experiments with existing methods demonstrate that our method surpasses state-of-the-art techniques, proving its effectiveness in prostate segmentation from TRUS images. The proposed method achieved an average Dice similarity coefficient of 92.92% and 94.0%, and the 95th percentile of Hausdorff distance of 1.07 mm and 0.77 mm on two datasets, demonstrating the potential to facilitate accurate MRI-TRUS fusion guided prostate biopsy.

Temporal Cardiovascular Dynamics for Improved PPG-Based Heart Rate Estimation

The oscillations of the human heart rate are inherently complex and non-linearβ€”they are best described by mathematical chaos, and they present a challenge when applied to the practical domain of cardiovascular health monitoring in everyday life. In this work, we study the non-linear chaotic behavior of heart rate through mutual information and introduce a novel approach for enhancing heart rate estimation in real-life conditions. Our proposed approach not only explains and handles the non-linear temporal complexity from a mathematical perspective but also improves the deep learning solutions when combined with them. We validate our proposed method on four established datasets from real-life scenarios and compare its performance with existing algorithms thoroughly with extensive ablation experiments. Our results demonstrate a substantial improvement, up to 40%, of the proposed approach in estimating heart rate compared to traditional methods and existing machine-learning techniques while reducing the reliance on multiple sensing modalities and eliminating the need for post-processing steps.

LRCOMF: A Learn-Review-Challenge Online Meta Learning Framework for EEG Emotion Recognition With Unlabeled Online Samples

Emotion recognition based on Electroencephalogram (EEG) is of great importance for cognitive psychology, disease therapy, etc. However, most mainstream recognition models are trained using batch learning, which fails to adapt to the dynamic and non-stationary nature of real-time EEG streams. In contrast, online learning can adjust model parameters continuously, but it typically relies on labeled data and a fully trained initial model, which are unavailable in practical scenarios. To tackle this challenge, we propose a novel learn-review-challenge online meta-learning framework (LRCOMF) for unlabeled online EEG learning. This framework incorporates a meta updating module through a multi-task cache and a customized sampling strategy to improve the model’s generalization during online learning. A sample judgement module is implemented based on a prototype weight being designed to estimate the confidence of predicted labels during the β€œlearn-review” step. Additionally, a challenge module using a clustering quality metric determines whether low-confidence samples can be reconsidered during the β€œchallenge” phase. The validation employed the DEAP and DREAMER datasets. In comparison to the initial baseline condition, the recognition model exhibited substantial enhancements 3.59%/3.87%/1.98% (p $

Robust Multimodal Cough Detection With Optimized Out-of-Distribution Detection for Wearables

Longitudinal and continuous monitoring of cough is crucial for early and accurate diagnosis of respiratory diseases. Recent developments in wearables provides at-home remote symptom monitoring with respect to more accurate and less frequent assessment in the clinics, but face practical challenges such as speech privacy, poor audio quality and background noise in uncontrolled real-world settings. This study addresses these challenges by developing and optimizing a multimodal cough detection system, enhanced with an Out-of-Distribution (OOD) detection algorithm. The cough sensing modalities include audio and Inertial Measurement Unit (IMU) signals. The system is optimized through training with an enhanced dataset and a weighted multi-loss approach for in-distribution classification, while OOD detection is improved by reconstructing training data components. Experiments demonstrate robustness across window sizes from 1–5 seconds and effectiveness at low audio sampling rates, where privacy is preserved. The optimized system achieves 90.1% accuracy with a cough F1 score of 0.75 at 16 kHz and 87.3% accuracy with an F1 score of 0.70 at 750 Hz, even with half the inference data being OOD. Most misclassifications arise from nonverbal sounds (e.g., sneezes, groans). Overall, the proposed Audio-IMU multimodal model with OOD detection significantly improves cough detection performance and offers a practical solution for real-world wearable applications.

Enhancing Fairness and Accuracy in Diagnosing Type 2 Diabetes in Young Adult Population

While type 2 diabetes is predominantly found in the elderly population, recent publications indicate an increasing prevalence in the young adult population. Failing to diagnose it in the minority younger age group could have significant adverse effects on their health. Several previous works acknowledge the bias of machine learning models towards different gender and race groups and propose various approaches to mitigate it. However, those works failed to propose any effective methodologies to diagnose diabetes in the young population, which is the minority group in the diabetic population. This is the first paper where we mention digital ageism towards the young adult population diagnosing diabetes. In this paper, we identify this deficiency in traditional machine learning models and propose an algorithm to mitigate the bias towards the young population when predicting diabetes. Deviating from the traditional concept of one-model-fits-all, we train customized machine-learning models for each age group. Our pipeline trains a separate machine learning model for every 5-year age band (i.e., age groups 30-34, 35-39, and 40-44). The proposed solution consistently improves recall of diabetes class by 26% to 40% in the young age group (30-44). Moreover, our technique outperforms 7 commonly used whole-group resampling techniques (i.e., random oversampling, random undersampling, SMOTE, ADASYN, Tomek-links, ENN, and Near Miss) by at least 36% in terms of diabetes recall in the young age group. Feature important analysis shows that the age attribute has a significant contribution to the decision of the original model, which was marginalized in the age-personalized model. Our method shows improved performance (e.g., balanced accuracy improved 7-12%) over multiple machine learning models and multiple sampling algorithms.

Morphology Prior Enhanced Teeth Segmentation for High-Resolution Oral Scans

Deep learning methods have been proposed for tooth segmentation on high-resolution intra-oral scans (IOS) that plays a crucial role in clinical dental practice. However, they generally segment teeth in a low-resolution data with a fixed receptive field and generate final segmentation by up-sampling interpolation, and neglect teeth’s morphology priors: their similar dental arch structures and significantly different curvatures in different parts of each tooth. They thus lack adaptability to different parts of each tooth, and show less accurate segmentation of boundary points between teeth and gums due to the up-sampling computation. Further, cluttered poses of IOS limit their generalization and usability of teeth location and geometric information. To address these limitations, a morphology prior enhanced teeth segmentation framework is proposed in this paper. Firstly, a robust preprocessing is introduced to align poses of different IOS by computing their dental arch orientations, thereby improving segmentation generalization and usability of IOS geometric information. Secondly, a decomposition-merging strategy is designed to avoid the up-sampling limitation, which decomposes an IOS into multiple low-resolution data and merges their segmentation outcomes into a high-resolution result. Thirdly, an innovative module integrating semantic and geometric features is proposed to adaptively select deformable receptive fields. It geometrically samples within a variable probability space to construct receptive fields with varied graph relationships for different points, facilitating adaptive segmentation of different parts of each tooth. Experimental results on 6238 IOS from four centers demonstrate that our method significantly outperforms 11 state-of-the-art methods, achieving a 6.93% enhancement for cross-center testing.
❌