❌

Normal view

MedCondDiff: Lightweight, Robust, Semantically Guided Diffusion for Medical Image Segmentation

arXiv:2512.00350v1 Announce Type: cross Abstract: We introduce MedCondDiff, a diffusion-based framework for multi-organ medical image segmentation that is efficient and anatomically grounded. The model conditions the denoising process on semantic priors extracted by a Pyramid Vision Transformer (PVT) backbone, yielding a semantically guided and lightweight diffusion architecture. This design improves robustness while reducing both inference time and VRAM usage compared to conventional diffusion models. Experiments on multi-organ, multi-modality datasets demonstrate that MedCondDiff delivers competitive performance across anatomical regions and imaging modalities, underscoring the potential of semantically guided diffusion models as an effective class of architectures for medical imaging tasks.

SelfAI: Building a Self-Training AI System with LLM Agents

arXiv:2512.00403v1 Announce Type: cross Abstract: Recent work on autonomous scientific discovery has leveraged LLM-based agents to integrate problem specification, experiment planning, and execution into end-to-end systems. However, these frameworks are often confined to narrow application domains, offer limited real-time interaction with researchers, and lack principled mechanisms for determining when to halt exploration, resulting in inefficiencies, reproducibility challenges, and under-utilized human expertise. To address these gaps, we propose \textit{SelfAI}, a general multi-agent platform that combines a User Agent for translating high-level research objectives into standardized experimental configurations, a Cognitive Agent powered by LLMs with optimal stopping criteria to iteratively refine hyperparameter searches, and an Experiment Manager responsible for orchestrating parallel, fault-tolerant training workflows across heterogeneous hardware while maintaining a structured knowledge base for continuous feedback. We further introduce two novel evaluation metrics, Score and $\text{AUP}_D$, to quantify discovery efficiency and search diversity. Across regression, NLP, computer vision, scientific computing, medical imaging, and drug discovery benchmarks, SelfAI consistently achieves strong performance and reduces redundant trials compared to classical Bayesian optimization and LLM-based baselines, while enabling seamless interaction with human researchers.

Human Decision-making is Susceptible to AI-driven Manipulation

arXiv:2502.07663v3 Announce Type: replace Abstract: AI systems are increasingly intertwined with daily life, assisting users with various tasks and guiding decision-making. This integration introduces risks of AI-driven manipulation, where such systems may exploit users' cognitive biases and emotional vulnerabilities to steer them toward harmful outcomes. Through a randomized between-subjects experiment with 233 participants, we examined human susceptibility to such manipulation in financial (e.g., purchases) and emotional (e.g., conflict resolution) decision-making contexts. Participants interacted with one of three AI agents: a neutral agent (NA) optimizing for user benefit without explicit influence, a manipulative agent (MA) designed to covertly influence beliefs and behaviors, or a strategy-enhanced manipulative agent (SEMA) equipped with established psychological tactics, allowing it to select and apply them adaptively during interactions to reach its hidden objectives. By analyzing participants' preference ratings, we found significant susceptibility to AI-driven manipulation. Particularly across both decision-making domains, interacting with the manipulative agents significantly increased the odds of rating hidden incentives higher than optimal options (Financial, MA: OR=5.24, SEMA: OR=7.96; Emotional, MA: OR=5.52, SEMA: OR=5.71) compared to the NA group. Notably, we found no clear evidence that employing psychological strategies (SEMA) was overall more effective than simple manipulative objectives (MA) on our primary outcomes. Hence, AI-driven manipulation could become widespread even without requiring sophisticated tactics and expertise. While our findings are preliminary and derived from hypothetical, low-stakes scenarios, we highlight a critical vulnerability in human-AI interactions, emphasizing the need for ethical safeguards and regulatory frameworks to protect human autonomy.

Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models

arXiv:2509.24365v2 Announce Type: replace-cross Abstract: Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: when trained on multimodal inputs, modality-shared transformers suffer from severe gradient conflicts between vision and text, particularly in shallow and deep layers. We trace this issue to the fundamentally different low-level statistical properties of images and text, while noting that conflicts diminish in middle layers where representations become more abstract and semantically aligned. To overcome this challenge, we propose Uni-X, a two-end-separated, middle-shared architecture. Uni-X dedicates its initial and final layers to modality-specific processing, while maintaining shared parameters in the middle layers for high-level semantic fusion. This X-shaped design not only eliminates gradient conflicts at both ends but also further alleviates residual conflicts in the shared layers. Extensive experiments validate the effectiveness of Uni-X. Under identical training conditions, Uni-X achieves superior training efficiency compared to strong baselines. When scaled to 3B parameters with larger training data, Uni-X matches or surpasses 7B AR-based UMMs, achieving a GenEval score of 82 for image generation alongside strong performance in text and vision understanding tasks. These results establish Uni-X as a parameter-efficient and scalable foundation for future unified multimodal modeling. Our code is available at https://github.com/CURRENTF/Uni-X

The AI Productivity Index (APEX)

arXiv:2509.25721v3 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

25 November 2025 at 19:00

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | PMC:PMC12647489 | DOI:10.1007/s12672-025-04010-z

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | DOI:10.1007/s12672-025-04010-z

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. Hypothesis formulation involves knowledge synthesis and hypothesis generation. Hypothesis validation includes the verification of scientific claims, theorem proving, and experiment validation. Manuscript publication encompasses manuscript writing and the peer review process. Furthermore, we identify and discuss the current challenges faced in these areas, as well as potential future directions for research. Finally, we also offer a comprehensive overview of existing benchmarks and tools across various domains that support the integration of AI into the research process. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zkzhou126/AI-for-Research.

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer

Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval

arXiv:2511.13293v1 Announce Type: new Abstract: Accurate healthcare prediction is critical for improving patient outcomes and reducing operational costs. Bolstered by growing reasoning capabilities, large language models (LLMs) offer a promising path to enhance healthcare predictions by drawing on their rich parametric knowledge. However, LLMs are prone to factual inaccuracies due to limitations in the reliability and coverage of their embedded knowledge. While retrieval-augmented generation (RAG) frameworks, such as GraphRAG and its variants, have been proposed to mitigate these issues by incorporating external knowledge, they face two key challenges in the healthcare scenario: (1) identifying the clinical necessity to activate the retrieval mechanism, and (2) achieving synergy between the retriever and the generator to craft contextually appropriate retrievals. To address these challenges, we propose GHAR, a \underline{g}enerative \underline{h}ierarchical \underline{a}gentic \underline{R}AG framework that simultaneously resolves when to retrieve and how to optimize the collaboration between submodules in healthcare. Specifically, for the first challenge, we design a dual-agent architecture comprising Agent-Top and Agent-Low. Agent-Top acts as the primary physician, iteratively deciding whether to rely on parametric knowledge or to initiate retrieval, while Agent-Low acts as the consulting service, summarising all task-relevant knowledge once retrieval was triggered. To tackle the second challenge, we innovatively unify the optimization of both agents within a formal Markov Decision Process, designing diverse rewards to align their shared goal of accurate prediction while preserving their distinct roles. Extensive experiments on three benchmark datasets across three popular tasks demonstrate our superiority over state-of-the-art baselines, highlighting the potential of hierarchical agentic RAG in advancing healthcare systems.

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v2 Announce Type: replace Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

The dual immunomodulatory role of B cells in tumorigenesis: mechanisms, microenvironment crosstalk, and therapeutic implications

Front Immunol. 2025 Oct 30;16:1649812. doi: 10.3389/fimmu.2025.1649812. eCollection 2025.

ABSTRACT

B lymphocytes exhibit a multifaceted and context-dependent role in tumor biology, acting as both promoters and suppressors of malignancy through dynamic interactions within the tumor microenvironment (TME). This review synthesizes current evidence on the dual functions of B cells in tumor immunity, highlighting their capacity to orchestrate antitumor responses via antigen presentation, antibody-dependent cytotoxicity, and tertiary lymphoid structure (TLS)-mediated T cell activation, while paradoxically driving immunosuppression through regulatory B cells (Bregs), pro-angiogenic signaling, and immune checkpoint modulation. Key mechanisms include TLS formation, which enhances cytotoxic T cell priming and correlates with improved immunotherapy outcomes, and Breg-mediated secretion of IL-10/TGF-β, which fosters T cell exhaustion and myeloid-derived suppressor cell recruitment. Tumor-type specificity is evident: TLS-rich malignancies like melanoma and Non-Small Cell Lung Cancer (NSCLC) show B cell-driven immune activation, whereas pancreatic and hepatocellular carcinomas demonstrate B cell functional plasticity influenced by metabolic and epigenetic reprogramming. Therapeutically, B cell-targeted strategies-including CD20 antibodies, CAR-T cells, and B cell epitope vaccines-demonstrate efficacy in hematologic and solid tumors, yet face challenges due to subset heterogeneity and sex-specific response disparities. Emerging approaches combine immune checkpoint inhibitors (ICBs) with TLS-inducing agents or exploit B cell-derived biomarkers for personalized therapy. Future directions emphasize deciphering B cell metabolic-niche crosstalk, optimizing combinatorial regimens, and leveraging spatial multiomics to resolve functional heterogeneity. By bridging mechanistic insights with clinical translation, this work underscores B cells as pivotal regulators of tumor immunity and advocates for precision strategies to harness their antitumor potential while mitigating pro-tumor plasticity.

PMID:41246318 | PMC:PMC12611826 | DOI:10.3389/fimmu.2025.1649812

MedFuse: Multiplicative Embedding Fusion For Irregular Clinical Time Series

arXiv:2511.09247v1 Announce Type: new Abstract: Clinical time series derived from electronic health records (EHRs) are inherently irregular, with asynchronous sampling, missing values, and heterogeneous feature dynamics. While numerical laboratory measurements are highly informative, existing embedding strategies usually combine feature identity and value embeddings through additive operations, which constrains their ability to capture value-dependent feature interactions. We propose MedFuse, a framework for irregular clinical time series centered on the MuFuse (Multiplicative Embedding Fusion) module. MuFuse fuses value and feature embeddings through multiplicative modulation, preserving feature-specific information while modeling higher-order dependencies across features. Experiments on three real-world datasets covering both intensive and chronic care show that MedFuse consistently outperforms state-of-the-art baselines on key predictive tasks. Analysis of the learned representations further demonstrates that multiplicative fusion enhances expressiveness and supports cross-dataset pretraining. These results establish MedFuse as a generalizable approach for modeling irregular clinical time series.

Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution

Stereo-seq V2 facilitates single-cell-resolution spatial RNA mapping in FFPE samples through random primer capture, uncovering ncRNAs, host-pathogen transcriptome profiling, and spatial immune repertoires in situ.

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads

arXiv:2506.20535v2 Announce Type: replace-cross Abstract: The rapid advancement of AI, particularly large language models (LLMs), has raised significant concerns about the energy use and carbon emissions associated with model training and inference. However, existing tools for measuring and reporting such impacts are often fragmented, lacking systematic metric integration and offering limited support for correlation analysis among them. This paper presents AIMeter, a comprehensive software toolkit for the measurement, analysis, and visualization of energy use, power draw, hardware performance, and carbon emissions across AI workloads. By seamlessly integrating with existing AI frameworks, AIMeter offers standardized reports and exports fine-grained time-series data to support benchmarking and reproducibility in a lightweight manner. It further enables in-depth correlation analysis between hardware metrics and model performance and thus facilitates bottleneck identification and performance enhancement. By addressing critical limitations in existing tools, AIMeter encourages the research community to weigh environmental impact alongside raw performance of AI workloads and advances the shift toward more sustainable "Green AI" practices. The code is available at https://github.com/SusCom-Lab/AIMeter.

Nanomaterial-assisted immunodiagnostic profiling and therapeutic targeting of hepatocellular carcinoma: from molecular biomarkers to clinical applications

Front Immunol. 2025 Oct 14;16:1668630. doi: 10.3389/fimmu.2025.1668630. eCollection 2025.

ABSTRACT

AIMS AND OBJECTIVES: This study aimed to identify immunologically relevant transcriptomic and proteomic biomarkers in hepatocellular carcinoma (HCC) and to characterize their B-cell epitopes for potential integration into nanomaterial-based biosensors and immunomodulatory platforms for early diagnosis and targeted therapy.

METHODS: We conducted a comprehensive multi-omics analysis by integrating transcriptomic (TCGA-LIHC) and proteomic data to identify differentially expressed genes (DEGs) in HCC. Protein-protein interaction networks and pathway enrichment were used to prioritize hub genes. Five candidate biomarkers, RFC2, HSP90AB1, YWHAZ, CYP2E1, and ADH4, were selected for qRT-PCR and serum ELISA validation in clinical cohorts comprising 85 HCC patients and 50 healthy controls. B-cell epitope prediction was performed using BepiPred 2.0 and validated through synthetic peptide-based ELISA in the same cohort to assess immunoreactivity. Diagnostic performance was evaluated using ROC curve analysis.

RESULTS: RFC2, HSP90AB1, and YWHAZ were significantly upregulated (|log2FC|>0.2) and showed high serological expression, whereas CYP2E1 and ADH4 were consistently downregulated. Predicted B-cell epitopes from RFC2, HSP90AB1, and YWHAZ exhibited strong immunoreactivity (AUC>0.84), indicating their diagnostic potential. Enrichment analysis revealed that upregulated DEGs were involved in cell cycle and mitotic progression, while downregulated genes were linked to immune suppression and metabolic dysfunction. These validated immunogenic epitopes offer promising anchors for nanomaterial-functionalized biosensors, such as gold nanoparticle-conjugated ELISA, graphene-based electrochemical platforms, and peptide-coated quantum dots, for ultrasensitive and multiplexed HCC detection.

CONCLUSION: By integrating transcriptomic and proteomic screening with epitope-level validation, we identified a novel panel of immunogenic biomarkers suitable for nanomaterial-enabled diagnostics in HCC. These findings support the translational potential of peptide-nano scaffold conjugates in developing minimally invasive, immune-responsive biosensing and therapeutic tools tailored for early-stage liver cancer management.

PMID:41164201 | PMC:PMC12558944 | DOI:10.3389/fimmu.2025.1668630

Prospective proteomics for discovering biomarkers in lung adenocarcinoma: a literature review

Transl Cancer Res. 2025 Sep 30;14(9):6102-6117. doi: 10.21037/tcr-2025-1092. Epub 2025 Sep 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung adenocarcinoma (LUAD), as the main subtype of non-small cell lung cancer (NSCLC), faces clinical challenges including molecular heterogeneity, late diagnosis, and aggressive growth, leading to a low 5-year survival rate. Biomarkers are critical for early detection, accurate differentiation of benign/malignant lesions, and guiding personalized treatment strategies. Proteomic technologies using liquid biopsy show potential by analyzing protein changes and post-translational modifications (PTMs) to identify novel biomarkers and unravel cancer mechanisms. This review examines proteomic advances in LUAD, compares platform strengths, lists validated protein markers, and discusses challenges like specificity and regulations. It aims to develop a precision medicine framework by integrating multi-omics data for improved diagnosis and treatment.

METHODS: This study conducted a literature review by searching the PubMed and Web of Science databases for original articles written in English from 2002 to 2025, using the keywords "lung adenocarcinoma" OR "LUAD" AND "biomarkers" AND "proteomics" OR "SomaScan" OR "spatial proteomics" to identify the latest research findings in the field of proteomics technology and LUAD biomarkers. The included studies mainly focused on the current landscape of biomarkers in the diagnosis, treatment, and prognosis of LUAD.

KEY CONTENT AND FINDINGS: This review discusses high-throughput methods for comprehensive protein profiling in accessible biospecimens (tissues, blood, urine) to identify biomarkers for LUAD. We systematically evaluate emerging proteomic strategies, including mass spectrometry (MS), proximity extension assays (PEAs), spatial proteomics techniques, and SomaScan platforms-coupled with innovative computational frameworks have revolutionized biomarkers discovery and their translational potential in developing precision diagnostics and targeted therapies. Additionally, the review addresses challenges in integrating proteomics with genomics, transcriptomics, and metabolomics, offering new methodologies and expanding research in life sciences. As technological advancements continue, it is anticipated that more potential biomarkers will be conducted to validate the broader application in LUAD treatment, addressing early-stage disease complexities and aiding in selecting more effective treatment strategies.

CONCLUSIONS: By synthesizing cutting-edge evidence on proteome-driven LUAD biomarkers, this review elucidates actionable strategies to refine early detection protocols and mechanism-informed personalized treatment frameworks, directly advancing precision oncology initiatives for this prevalent malignancy through biomarker-guided clinical decision-making and multi-omics integration.

PMID:41158224 | PMC:PMC12554480 | DOI:10.21037/tcr-2025-1092

From Detection to Discovery: A Closed-Loop Approach for Simultaneous and Continuous Medical Knowledge Expansion and Depression Detection on Social Media

arXiv:2510.23626v1 Announce Type: cross Abstract: Social media user-generated content (UGC) provides real-time, self-reported indicators of mental health conditions such as depression, offering a valuable source for predictive analytics. While prior studies integrate medical knowledge to improve prediction accuracy, they overlook the opportunity to simultaneously expand such knowledge through predictive processes. We develop a Closed-Loop Large Language Model (LLM)-Knowledge Graph framework that integrates prediction and knowledge expansion in an iterative learning cycle. In the knowledge-aware depression detection phase, the LLM jointly performs depression detection and entity extraction, while the knowledge graph represents and weights these entities to refine prediction performance. In the knowledge refinement and expansion phase, new entities, relationships, and entity types extracted by the LLM are incorporated into the knowledge graph under expert supervision, enabling continual knowledge evolution. Using large-scale UGC, the framework enhances both predictive accuracy and medical understanding. Expert evaluations confirmed the discovery of clinically meaningful symptoms, comorbidities, and social triggers complementary to existing literature. We conceptualize and operationalize prediction-through-learning and learning-through-prediction as mutually reinforcing processes, advancing both methodological and theoretical understanding in predictive analytics. The framework demonstrates the co-evolution of computational models and domain knowledge, offering a foundation for adaptive, data-driven knowledge systems applicable to other dynamic risk monitoring contexts.
❌