❌

Reading view

MoChat: Joints-Grouped Spatio-Temporal Grounding Multimodal Large Language Model for Multi-Turn Motion Comprehension and Description

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. This limitation is particularly pronounced in home exercise monitoring, neurological disorder assessment, and rehabilitation, where precise motion analysis is crucial for ensuring exercise efficacy, detecting early signs of neurological conditions, and guiding personalized recovery programs. In this paper, we propose MoChat, a multimodal large language model capable of spatio-temporal grounding of human motion and multi-turn dialogue understanding. To achieve this, we first group spatial features in skeleton frames according to human anatomical structures and process them through a Joints-Grouped Skeleton Encoder. The encoder’s outputs are fused with large language model embeddings to generate spatio-aware representations. A cross-attention-based Regression Head module is then designed to align hidden-layer embeddings and skeletal sequence embeddings, enabling precise temporal grounding. Furthermore, we develop a pipeline for temporal grounding task to extract timestamps from skeleton-text pairs and construct a multi-turn instruction dialogues for spatial grounding task. Finally, various task instructions are generated for jointly training. Experimental results demonstrate that MoChat achieves state-of-the-art performance across multiple metrics in motion understanding tasks, making it as the first model capable of fine-grained spatio-temporal grounding of human motion.
  •  

4PM: Privacy-Preserving Patient-Provider Matching Service in Digital Healthcare System

For digital health platforms, the challenge is balancing patient privacy with the ability to match patients to the right providers quickly and accurately. Existing systems often suffer from privacy leakage, insufficient matching precision, and degraded performance when dealing with large-scale data. In this paper, we propose 4PM, a novel privacy-preserving patient–provider matching scheme that leverages secure computation to deliver strong privacy guarantees while ensuring efficient and accurate matching. Our method partitions patient data between two non-colluding servers via secret sharing, employing the optimized Millionaires' Protocol for secure ranking and leveraging oblivious retrieval techniques for privacy-preserving matching. 4PM significantly reduces the computational complexity of high-dimensional data, achieving end-to-end latency within 0.5 seconds in scenarios with 200 doctors and 200-dimensional symptom vectors. Our work contributes to fostering secure and trustworthy healthcare in the digital era.
  •  

DGAN-MPCC: A Novel Dual-GAN Enhanced Multi-Positive Contrastive Clustering Method for Omics Data

AI-driven clustering methods have significantly enhanced the capacity of researchers to explore the heterogeneity inherent in single-cell omics data, which is a crucial aspect of understanding complex biological systems in healthcare. Despite advancements, most existing methods still face challenges, such as (1) inherent sparsity and noise in cell data, which frequently lead to overfitting in networks. To address this, some researchers have proposed using Generative Adversarial Networks (GANs), however, the conventional single GAN architecture primarily focuses on simple data enhancement and lacks the capacity to infer complex biological data, thus leading to suboptimal clustering performance. (2) Contrastive learning has been proposed to obtain high-quality clustering structures; however, existing methods predominantly rely on a single positive pair, which prevents them from modeling and learning continuous transitions in cell states and thus hinders the establishment of feature representations sensitive to cell types. To address these issues, we propose a novel Dual-GAN Enhanced Multi-Positive Contrastive Clustering Method, DGAN-MPCC, tailored for low-quality single-cell data. Specifically, we propose using two independent GANs to simultaneously enhance the quality of both the input and bottleneck layers, thereby refining the generated cell embedding. Additionally, we have developed a multi-positive contrastive clustering framework that adaptively defines a multi-positive set from clustering structures, enabling each sample to establish positive relationships with all samples within the same cluster, thereby diversifying supervisory signals within the same class. Extensive experiments on several real-world single-cell datasets demonstrate that DGAN-MPCC surpasses current methods across multiple scenarios, providing a more robust and efficient tool for AI-driven decision-making in healthcare.
  •  

Extraction of Seafarers’ Occupational Plasticity Brain Network Based on Effective Connectivity Lateralization

Lateralization is an effective model for exploring changes in brain activity and is widely used to assess brain function. Seafarers, as an occupation working in marine environments, are subjected to long-term specialized occupational demands and experiences, which inevitably impact brain function. By utilizing lateralization, the influence of occupational experience on brain activity can be further explored. A novel Effective Connectivity Lateralization Analysis (ECLA) framework is proposed, which incorporates a Transformer-based Granger causality model (Transformer-GC) to analyze the effects of seafaring on brain plasticity. The Transformer-GC model constructs effective connectivity (EC) matrices, and lateralization indices are derived to investigate occupational influences on brain activity. Two control groups of non-seafarers are included to identify seafarers’ unique occupational plasticity brain networks. Results show that Transformer-GC achieves an accuracy improvement of nearly 16% and 19.4% over the GRU-based and MVGC model, respectively, and a 5% gain over Pearson-based functional connectivity, confirming its superior performance. Moreover, the results of the ECLA showed significant differences in VentralAttention, Somatomotor, DorsalAttention in the seafarer, demonstrating that these brain networks are affected by the long-term work of seafarers. The findings demonstrate the effectiveness of ECLA in revealing the impact of long-term maritime work on brain plasticity, particularly in identifying the brain network of seafarers’ occupational plasticity. It is shown that occupational experience can reshape the lateralization of brain functional activity, offering new insights into neural plasticity across different professions.
  •  

Joint Learning of Confidence Fusion, Semantic Alignment and Group-Guided Reliability: A Novel Semi-Supervised Learning Framework for 3D Medical Image Segmentation

Semi-supervised learning (SSL) has shown strong potential in reducing the reliance on large-scale voxel-level annotations for 3D medical image segmentation. However, existing SSL methods often suffer from unstable training and limited generalization due to unreliable pseudo-labels and insufficient structural modeling in unlabeled data. These challenges are especially evident in volumetric contexts, where anatomical structures exhibit high inter-class imbalance and complex spatial dependencies. To address these issues, we propose a semi-supervised framework built upon a single-network architecture that integrates feature learning, consistency regularization, and pseudo-label reliability modeling in a unified manner. The framework comprises three key components: 1) a Confidence-aware Multi-level Fusion Network (CMFN) for capturing robust multi-scale semantic representations; 2) a Semantic-Enhanced Center Alignment (SECA) module to align feature distributions of group-level anatomical structures and mitigate semantic drift in pseudo-labels; and 3) a Group-Guided Reliability Assessment (GGRA) module that enhances pseudo-label reliability by modeling confidence errors in a group-aware structural context. Together, these modules enhance both feature discriminability and the reliability of pseudo-labels.We evaluate our framework on three public 3D medical image segmentation benchmarks: LA, BTCV, and BraTS19. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches under limited annotation, achieving superior accuracy and generalization across diverse anatomical structures and segmentation tasks.
  •  

A Review of Methods for Trustworthy AI in Medical Imaging: The FUTURE-AI Guidelines

Recent advancements in artificial intelligence (AI) and the vast data generated by modern clinical systems have driven the development of AI solutions in medical imaging, encompassing image reconstruction, segmentation, diagnosis, and treatment planning. Despite these successes and potential, many stakeholders worry about the risks and ethical implications of imaging AI, viewing it as complex, opaque, and challenging to understand, use, and trust in critical clinical applications. The FUTURE-AI guideline for trustworthy AI in healthcare was established based on six guiding principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Through international consensus, a set of recommendations was defined, covering the entire lifecycle of medical AI tools, from design, development, and validation to regulation, deployment, and monitoring. In this paper, we describe how these specific recommendations can be instantiated in the domain of medical imaging, providing an overview of current best practices along with guidelines and concrete metrics on how those recommendations could be met, offering a valuable resource to the international medical imaging community.
  •  

FEI-Hi: Federated Edge Intelligence for Healthcare Informatics

As the Internet of Things (IoT) and artificial intelligence (AI) technologies are rapidly evolving, smart healthcare has emerged as a transformative solution to enhance healthcare quality and optimize resource allocation. This study introduces FEI-Hi, a federated edge intelligence paradigm that integrates edge computing with federated learning (FL) to enable secure and efficient medical data processing. FEI-Hi comprises three principal layers: FL layer, which facilitates cross-device collaborative training through encrypted model updates; aggregation layer, which refines the global model by consolidating updates; and edge layer, which performs local data processing and model inference. FEI-Hi leverages distributed intelligent computation, model parameter compression, and efficient node clustering to enhance the accuracy and efficiency of medical data processing significantly. By employing Wasserstein distance for clustering and parameter selection, FEI-Hi ensures model convergence and stability. Experimental results on multiple medical datasets demonstrate a 30% improvement in the model training speed and an F1-score exceeding 90%, surpassing the state-of-the-art (SOTA) benchmarks in model parameter transfer efficiency, training speed, and accuracy.
  •  

Learning Across the Divide: Personalised Federated Learning for Robust Clinical Modelling Under Data-View Heterogeneity

Federated Learning (FL) enables collaborative clinical modelling across distributed electronic health records (EHRs) without sharing sensitive patient data. However, variations in medical practice, documentation standards, and data collection across institutions create data-view heterogeneity, where clients possess different or only partially overlapping clinical feature sets. This misalignment hinders the use of standard FL methods. Existing approaches rely on complex preprocessing and manual harmonisation, which can cause information loss, reduce data utility, limit scalability, and restrict client-specific personalisation. To address these limitations, we propose Personalised Attention-based Federated Graph Network (PAFNet), a scalable FL framework that enables meaningful parameter exchange across heterogeneous clients by mapping their distinct data-views into a shared latent space through client-specific projection layers. It then applies a personalised adaptation mechanism using trainable parameter masks, allowing each client to selectively incorporate global model parameters relevant to its own feature set. This design preserves local specificity, improves generalisation, and removes the need for heavy manual preprocessing common in existing approaches. Across CURIAL, eICU, and MIMIC-III datasets, PAFNet consistently outperformed state-of-the-art data-view heterogeneity FL baselines, demonstrating strong generalisation under substantial differences in client feature sets. By enabling effective personalisation and cross-institutional knowledge sharing without extensive harmonisation, PAFNet offers a robust and scalable solution for the federated training of clinical models in data-view heterogeneous environments.
  •  

A Novel Multi-Task Teacher-Student Architecture With Self-Supervised Pretraining for 48-Hour Vasoactive-Inotropic Trend Analysis in Sepsis Mortality Prediction

Sepsis is a major cause of ICU mortality, where early recognition and effective interventions are essential for improving patient outcomes. However, the vasoactive-inotropic score (VIS) varies dynamically with a patient’s hemodynamic status, complicated by irregular medication patterns, missing data, and confounders, making sepsis prediction challenging. To address this, we propose a novel Teacher–Student multitask framework with self-supervised VIS pretraining via a Masked Autoencoder (MAE). The teacher model performs mortality classification and severity-score regression, while the student distills robust time-series representations, enhancing adaptation to heterogeneous VIS data. Compared to LSTM-based methods, our approach achieves an AUROC of 0.829 on MIMIC-IV 3.0 (9,476 patients), outperforming the baseline (0.74). SHAP analysis revealed that SOFA score (0.147) had the greatest impact on ICU mortality, followed by LODS (0.033), single marital status (0.031), and Medicaid insurance (0.023), highlighting the role of sociodemographic factors. SAPSII (0.020) also contributed significantly. These findings suggest that clinical and social factors should be considered in ICU decision-making. Our multitask and distillation strategies enable earlier identification of high-risk patients, improving prediction accuracy and disease management, offering new tools for ICU decision support.
  •  

Beyond Contact: An Open-Set Biometric Identification System Using Radar-Extracted Heart Signals

This paper proposes a novel radar-based framework for non-contact biometric identification through heart signal extraction, targeting secure and privacy-conscious identification scenarios. Traditional biometric methods, such as fingerprint and facial recognition, face challenges including privacy concerns, vulnerability to spoofing, and the requirement for close proximity or direct line-of-sight. Our framework addresses these issues by reconstructing electrocardiogram (ECG) signals from radar-extracted cardiac motion data and implementing an open-set person identification system. Specifically, the framework integrates ECGReconNet, a specialized deep learning model for reconstructing ECG signals from human chest wall displacement, the InceptionTime model enhanced with fixed-Class Anchor Clustering (fixed-CAC) loss for robust feature anchoring, and a hypersphere-based delineation method to differentiate known from unknown individuals. Experimental results on a public dataset demonstrate state-of-the-art performance, achieving 99.61% accuracy in closed-set identification (27 subjects) and 93.97% accuracy under challenging open-set conditions (14 known and 13 unknown subjects). However, the proposed approach exhibits limitations, including sensitivity to abrupt body movements and environmental noise, potential performance degradation under severe cardiac irregularities, and reduced efficacy with increased numbers of unknown identities.
  •  

ViG3D-UNet: Volumetric Vascular Connectivity-Aware Segmentation via 3D Vision Graph Representation

Accurate vascular segmentation is essential for coronary visualization and the diagnosis of coronary heart disease. This task involves the extraction of sparse tree-like vascular branches from volumetric space. However, existing methods have faced significant challenges due to discontinuous vascular segmentation and missing endpoints. To address this issue, a 3D vision graph neural network framework, named ViG3D-UNet, was introduced. This method integrates 3D graph representation and aggregation within a U-shaped architecture to facilitate continuous vascular segmentation. The ViG3D module captures volumetric vascular connectivity and topology, while the convolutional module extracts fine vascular details. These two branches are combined through channel attention to form the encoder feature. Subsequently, a paperclip-shaped offset decoder minimizes redundant computations in the sparse feature space and restores the feature map size to match the original input dimensions. To evaluate the effectiveness of the proposed approach for continuous vascular segmentation, evaluations were performed on two public datasets, ASOCA and ImageCAS. The segmentation results show that the ViG3D-UNet surpassed competing methods in maintaining vascular segmentation connectivity while achieving high segmentation accuracy.
  •  

MediTEDNet: Visual State Space Model for Thyroid Eye Disease Classification

Thyroid eye disease (TED) is a prevalent autoimmune orbital disorder that can severely impair visual function and significantly diminish patients’ quality of life. In recent years, several studies have attempted to automate TED diagnosis using optical coherence tomography (OCT) images. However, existing approaches primarily rely on convolutional neural networks (CNNs) combined with attention mechanisms and are mostly trained using traditional cross-entropy loss. Although Transformers excel at modeling long-range dependencies, their quadratic computational complexity when processing high-resolution medical images, along with subpar classification accuracy in challenging scenarios such as highly similar pathological regions and blurred image boundaries, limit their clinical applicability. To tackle these challenges, we propose a hybrid architecture that integrates CNNs, attention mechanisms, and visual state space models (VSSMs) to enhance the robustness and discriminability of image features. In addition, to achieve intra-class compactness and inter-class separation, we design a contrastive loss based on positive and negative sample prototypes. Specifically, we introduce proximal inter-class mean sampling (PICMS) and incorporate a normalized distance metric guided by a distinguishable–indistinguishable triplet partitioning mechanism. We also introduce a hierarchical noise-resilient training strategy to reduce the effects of noise frequently present in clinical images. To assess the effectiveness of our proposed model, we conduct experiments on two public datasets (OCT-2017 and OCT-C8) and a clinical dataset of TED images. The results reveal that our model outperforms existing methods across multiple evaluation metrics, including accuracy and F1-score while demonstrating superior diagnostic stability and generalization capability.
  •  

Enhancing the Interpretation of Skin Lesion Diagnosis: Concept Adaptive Fine-Tuning of Vision-Language Models

Significant progress has been made in applying deep learning for the automatic diagnosis of skin lesions. However, most models remain unexplainable, which severely hinders their application in clinical settings. Concept-based ante-hoc interpretable models have the potential to clarify the decision-making process of diagnosis by learning high-level, human-understandable concepts, while they can only provide numerical values of conceptual contributions. Pre-trained Vision-Language Models (VLMs) can learn rich vision-language correlations from large-scale image-text pairs. Fine-tuning pre-trained VLMs for specific downstream tasks is an effective way to reduce data requirements. Nevertheless, when there is a substantial disparity between the pre-trained model and the target task, existing tuning methods frequently struggle to generalize, necessitating substantial training data to fully adapt VLMs to specialized medical tasks. In this work, we propose a concept adaptive fine-tuning (CptAFT) method based on the pre-trained VLM, BiomedCLIP, to develop a concept-based multi-modal interpretable skin lesion diagnosis model. By incorporating medical texts, such as reports and conceptual terms, our model can recognize fine-grained features and provide robust, natural language-driven interpretability. Moreover, our concept-adaptive method that reconstructs images using concept logits and imposes a consistency loss with the original image, enabling the VLM to quickly adapt to the task with a small amount of training data. Extensive experimental results demonstrate that our approach outperforms state-of-the-art closed box and interpretable models in both classification performance and medically relevant interpretability. In particular, after fine-tuning with a small amount of data, our model outperforms MONET, a model trained on the large Skin Disease Image-Report dataset, by 8.28% in concept recognition ability, demonstrating the interpretability of our model.
  •  

Graph Clustering-Guided Multi-View Neighborhood-Enhanced Graph Contrastive Learning for Drug-Target Interaction Prediction

Drug-target interaction (DTI) identification is of great significance in drug development in various areas, such as drug repositioning and potential drug side effects. Although a great variety of computational methods have been proposed for DTI prediction, it is still a challenge in the face of sparsely correlated drugs or targets. To address the impact of data sparsity on the model, we propose a multi-view neighborhood-enhanced graph contrastive learning approach (MneGCL), which is based on graph clustering according to the adjacency relationship in various similarity networks between drugs or targets, to fully exploit the information of drugs and targets with few corrections. MneGCL first performs semantic clustering of drugs and targets by identifying strongly correlated nodes in the semantic similarity network to construct semantic contrastive prototypes, while simultaneously establishing phenotypic prototypes based on the Gaussian interaction profile kernel similarity. These complementary views are then combined through neighborhood-enhanced contrastive learning to effectively capture latent homogeneous features and enhance representation learning for sparse nodes in heterogeneous graphs, with final predictions generated through a graph autoencoders framework. Comparative experimental results demonstrate that MneGCL achieves superior performance across three benchmark datasets, with particularly notable improvements on the highly sparse DrugBank dataset, showing an average $2.5 \%$ increase to baseline models. Additional experiments further validate the effectiveness of MneGCL in enriching feature representations for sparsely connected nodes.
  •  

Multicontrast MR-Guided Diffusion Model for Ultra-Low-Dose Brain PET Denoising in Temporal Lobe Epilepsy

Positron Emission Tomography (PET) is a critical imaging modality in nuclear medicine but requires radioactive tracer administration, which increases radiation exposure risks. While recent studies have investigated MR-guided low-dose PET denoising, they neglect two critical factors: the synergistic roles of multicontrast MR images and disease-specific denoising requirements. In this work, we propose a diffusion model that integrates T1-weighted, T2 fluid attenuated inversion recovery (T2 FLAIR), and hippocampal-optimized (T2 HIPPO) MR sequences to achieve ultra-low-dose PET denoising tailored for temporal lobe epilepsy (TLE). Our parallel cross-modal fusion (PCMF) module employs dedicated encoders to extract cross-modal features—which are dynamically integrated via attention mechanisms. Extensive experiments demonstrate that our method outperforms other approaches in preserving image quality. The PSNR and SSIM obtained were 37.0251 $\pm$ 1.5215 dB and 0.9760 $\pm$ 0.0057 (p
  •  

Leveraging Multi-Text Joint Prompts in SAM for Robust Medical Image Segmentation

The Segment Anything Model (SAM) has attracted considerable attention due to its impressive performance and demonstrates potential in medical image segmentation. Compared to SAM’s native point andbounding box prompts, text prompts offer a simpler and more efficient alternative in the medical field, yet this approach remains relatively underexplored. In this paper, we propose a SAM-based framework that integrates a pre-trained vision-language model to generate referring prompts, with SAM handling the segmentation task. The outputs from multimodal models such as CLIP serve as input to SAM’s prompt encoder. A critical challenge stems from the inherent complexity of medical text descriptions: they typically encompass anatomical characteristics, imaging modalities, and diagnostic priorities, resulting in information redundancy and semantic ambiguity. To address this, we propose a text decomposition-recomposition strategy. First, clinical narratives are parsed into atomic semantic units (appearance, location, pathology, and so on). These elements are then recombined into optimized text expressions. We employ a cross-attention module among multiple texts to interact with the joint features, ensuring that the model focuses on features corresponding to effective descriptions. To validate the effectiveness of our method, we conducted experiments on several datasets. Compared to the native SAM based on geometric prompts, our model shows improved performance and usability.
  •  

Epileptic Seizure Prediction Using Multi-Strategy Data Augmentation and Hierarchical Contrastive Learning

Accurate early prediction of epileptic seizures is crucial for improving patients’ quality of life. However, existing seizure prediction methods often rely on large-scale labeled datasets and face challenges in generalization and real-time performance. To address these issues, this study proposes an efficient seizure prediction framework that achieves high performance even with limited labeled data, significantly reducing dependence on extensive annotations. To better distinguish preictal states, contrastive learning is employed to enhance feature separation between interictal and preictal periods, leading to improved sensitivity in detecting early seizure patterns. First, a data augmentation strategy is designed, incorporating wavelet-based frequency mixing, temporal masking, and window-based masking to enhance model robustness and generalization. Second, a hierarchical contrastive loss function is introduced, integrating instance-level and temporal contrastive learning to improve the model’s ability to capture preictal patterns. Finally, a lightweight SE-EEGNet is developed and optimized as a feature extractor, strengthening critical feature extraction and enabling real-time seizure prediction. On the CHB-MIT dataset, the proposed method achieves 94.51% accuracy, 95.05% sensitivity, a 0.024/h false positive rate (FPR), and a 20.12-minute prediction time using only 30% labeled data. On the Siena dataset, it achieves 93.14% accuracy, 92.77% sensitivity, and a 0.030/h FPR. Moreover, performance improves further as the amount of labeled data increases, validating the effectiveness and practical applicability of the proposed approach in seizure prediction.
  •  

AI-Based QRS Onset Detection in the Early Ventricular Activation Site ECGs

Identifying the onset of the QRS complex is an important step for localizing the site of origin (SOO) of premature ventricular complexes (PVCs) and the exit site of Ventricular Tachycardia (VT). However, identifying the QRS onset is challenging due to signal noise, baseline wander, motion artifact, and muscle artifact. Furthermore, in VT, QRS onset detection is especially difficult due to the overlap with repolarization from the prior beat. In this study, 7706 captured bipolar pacing beats (Stim-QRS
  •  

Interictal Epileptiform Discharge Detection Using Dual-Domain Features and GAN

Interictal Epileptiform Discharge is essential for identifying epilepsy. However, the unpredictable and non-stationary nature of electroencephalogram (EEG) patterns poses considerable challenges for reliable identification. Manual interpretation of EEG is subjective and time-consuming. With advancements in machine learning and deep learning, computer-aided approaches for automated IED detection have been rapidly developed. The state-of-the-art convolutional neural network (CNN)-based methods have shown promising results but struggle to capture long-term dependencies in time-series data. In contrast, Transformer excels at modeling sequential information through self-attention mechanisms, overcoming the CNN limitations. This study proposes an IED Detector (IEDD) that integrates convolutional layers and a Transformer to detect IEDs. The IEDD initially employs convolutional layers to extract local features of IEDs, followed by a Transformer to model long-term dependencies. To further extract spatial features, EEG data are represented as a three-dimensional tensor with embedded channel topology, where a CNN captures spatial features at each sampling point and a Long Short-Term Memory (LSTM) network models their temporal evolution. Additionally, due to the scarcity of IED data, a novel Transformer-based Generative Adversarial Network (GAN) is developed to augment the IED dataset. Experimental results show the proposed approach achieves an average accuracy of 96.11% on the augmented Dataset 1 and 95.25% on Dataset 2 for binary classification, with an average sensitivity of 87.26% and precision of 89.96% for multi-label classification. These findings provide valuable insights into advancing deep learning and Transformer-based approaches for automated IED detection.
  •  

Multi-Channel Temporal Interference Retinal Stimulation Based on Reinforcement Learning

Retinal degenerative diseases such as age-related macular degeneration and retinitis pigmentosa cause severe vision impairment, while current electrical stimulation therapies are limited by poor spatial targeting precision. As a promising non-invasive alternative, the efficacy of temporal interference stimulation (TIS) for retinal targeting depends on optimized multi-electrode parameters. This study reconstructed a whole-head finite element model with detailed ocular structures and applied reinforcement learning (RL)-based multi-channel electrode parameter optimization to retinal stimulation. Systematic evaluation demonstrated that the focal precision of TIS improves with increasing channel numbers (consistent across all subject head models), with RL significantly outperforming conventional genetic algorithms (GA) and unsupervised neural networks (USNN) in focusing capability. Furthermore, by implementing the computationally intensive envelope calculation using the JAX framework, we achieved a nearly order-of-magnitude reduction in optimization time (to approx. 2 minutes per run on an RTX 4090D), significantly enhancing the practical feasibility of the proposed RL framework. This work provides a novel and computationally efficient methodology for precise non-invasive neuromodulation parameter optimization, applicable not only to retinal diseases but potentially to broader neurological conditions.
  •  
❌