❌

Normal view

From Detection to Discovery: A Closed-Loop Approach for Simultaneous and Continuous Medical Knowledge Expansion and Depression Detection on Social Media

arXiv:2510.23626v1 Announce Type: cross Abstract: Social media user-generated content (UGC) provides real-time, self-reported indicators of mental health conditions such as depression, offering a valuable source for predictive analytics. While prior studies integrate medical knowledge to improve prediction accuracy, they overlook the opportunity to simultaneously expand such knowledge through predictive processes. We develop a Closed-Loop Large Language Model (LLM)-Knowledge Graph framework that integrates prediction and knowledge expansion in an iterative learning cycle. In the knowledge-aware depression detection phase, the LLM jointly performs depression detection and entity extraction, while the knowledge graph represents and weights these entities to refine prediction performance. In the knowledge refinement and expansion phase, new entities, relationships, and entity types extracted by the LLM are incorporated into the knowledge graph under expert supervision, enabling continual knowledge evolution. Using large-scale UGC, the framework enhances both predictive accuracy and medical understanding. Expert evaluations confirmed the discovery of clinically meaningful symptoms, comorbidities, and social triggers complementary to existing literature. We conceptualize and operationalize prediction-through-learning and learning-through-prediction as mutually reinforcing processes, advancing both methodological and theoretical understanding in predictive analytics. The framework demonstrates the co-evolution of computational models and domain knowledge, offering a foundation for adaptive, data-driven knowledge systems applicable to other dynamic risk monitoring contexts.

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

arXiv:2510.21293v2 Announce Type: replace Abstract: Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. However, current research often adopts techno-centric approaches, focusing primarily on technical attributes such as reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.

Integrating deep learning and multi-omics features in radiation pneumonitis prediction for lung cancer patients using PET/CT

BMC Med Imaging. 2025 Oct 27;25(1):426. doi: 10.1186/s12880-025-01971-z.

ABSTRACT

BACKGROUND: To investigate the feasibility and accuracy of PET radiomics features, along with their combination with CT radiomics, dosiomics, and deep learning (DL) features, in predicting radiation pneumonitis (RP) in lung cancer patients treated with volumetric modulated arc therapy (VMAT).

METHODS: A total of 206 and 27 lung cancer patients who underwent VMAT with pre-treatment PET/CT imaging were enrolled from Hospital One and Hospital Two for model training and external validation, respectively. Four machine learning (ML) methods were applied to build radiomics models with features extracted from CT (R_CT), PET (R_PET), radiomics features fused PET/CT (R_fFU) and fused PET/CT images (R_ iFU), as well dosiomics features (D). Three DL models were built to extract features from PET (DL_PET), CT (DL_CT), and fused PET/CT images (DL_FU). The best-performing radiomics and DL models were combined with dosiomics to create the final joint model. ROC curves with AUC, accuracy, sensitivity, and specificity evaluated the performance. A nomogram was constructed using top-performing model features, parameters, and relevant clinical factors.

RESULTS: The extreme gradient boosting (XGBoost) and 18-layer residual neural network (Resnet-18) achieved the best performance. The R+D+DL model combined radiomics, dosiomics, and DL features achieved AUCs of 0.93, 0.92 and 0.89 in the training, internal validaiton and external validation cohorts, respectively. A nomogram constructed with gender, Adaptive RT, SUVp90, and XGBoost-score achieved an AUC of 0.94 for RP prediction in VMAT-treated lung cancer patients using PET/CT.

CONCLUSION: Integrating radiomics, DL, dosiomics features and SUVp90 is promising in the RP prediction for lung cancer patients underwent VMAT using PET/CT images.

PMID:41146084 | DOI:10.1186/s12880-025-01971-z

Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions

arXiv:2411.14571v2 Announce Type: replace-cross Abstract: Answering end user security questions is challenging. While large language models (LLMs) like GPT, LLAMA, and Gemini are far from error-free, they have shown promise in answering a variety of questions outside of security. We studied LLM performance in the area of end user security by qualitatively evaluating 3 popular LLMs on 900 systematically collected end user security questions. While LLMs demonstrate broad generalist ``knowledge'' of end user security information, there are patterns of errors and limitations across LLMs consisting of stale and inaccurate answers, and indirect or unresponsive communication styles, all of which impacts the quality of information received. Based on these patterns, we suggest directions for model improvement and recommend user strategies for interacting with LLMs when seeking assistance with security.

Multimodal 3D Genome Pre-training

arXiv:2504.09060v2 Announce Type: replace-cross Abstract: Deep learning techniques have driven significant progress in various analytical tasks within 3D genomics in computational biology. However, a holistic understanding of 3D genomics knowledge remains underexplored. Here, we propose MIX-HIC, the first multimodal foundation model of 3D genome that integrates both 3D genome structure and epigenomic tracks, which obtains unified and comprehensive semantics. For accurate heterogeneous semantic fusion, we design the cross-modal interaction and mapping blocks for robust unified representation, yielding the accurate aggregation of 3D genome knowledge. Besides, we introduce the first large-scale dataset comprising over 1 million pairwise samples of Hi-C contact maps and epigenomic tracks for high-quality pre-training, enabling the exploration of functional implications in 3D genomics. Extensive experiments show that MIX-HIC can significantly surpass existing state-of-the-art methods in diverse downstream tasks. This work provides a valuable resource for advancing 3D genomics research.
❌