❌

Reading view

The Relationship Between Physician Self-Disclosure and Patient Acquisition in Digital Health Markets: Cross-Sectional Study

Background: Online health communities have evolved into digital marketplaces where physicians have to compete for patients. Existing research examines physician-patient dynamics through a patient-centric lens, treating physicians as passive recipients of ratings and reviews, while the strategic role of physician self-disclosure remains unexamined. This gap constrains a comprehensive understanding of how physicians can actively shape patient decisions, making the investigation of strategic self-disclosure imperative. Objective: This study aims to investigate the relationship between physician self-disclosure breadth (scope of information) and depth (detailed expertise) and patient decision-making, as well as whether regional digital health care level (DHL) moderates these relationships. Methods: We conducted a cross-sectional analysis of observational data to test these relationships. Data were collected from China’s online health care platform Haodf from September to December 2024. Self-disclosure breadth (including clinical performance, academic experience, and social reputation), self-disclosure depth (including expertise coverage, richness, and granularity), and patient decision-making (total visits) were captured through manual content coding and quantitative measurement. We used structured content analysis to extract the disclosure components, informational scope, and descriptive details of each profile. Then, using validated operational formulas, we calculated the composite indices for disclosure breadth and depth based on the coded dimensions. The study generated 1798 final physician samples with complete data across 14 focal variables. The hypotheses were tested using an ordinary least squares regression model, and 4 robustness checks were conducted, including variable substitution and different resampling techniques. Results: In the primary ordinary least squares regression models, self-disclosure breadth was significantly and positively associated with patient visits (β=0.255, 95% CI 0.054-0.456; P=.01), as was self-disclosure depth (β=0.098, 95% CI 0.030-0.167; P=.005). The breadth×DHL interaction was positive and significant (β=0.261, 95% CI 0.061-0.461; P=.01). Similarly, the depth×DHL interaction was positive and significant (β=0.070, 95% CI 0.002-0.138; P=.045). It should be noted that the association for self-disclosure breadth was stronger than that of self-disclosure depth. DHL strengthened the relationship between the disclosure strategies with patient visits. This contextual amplification indicates that DHL serves as a critical boundary condition, determining the degree to which physician self-disclosure strategies translate into patient acquisition outcomes. Conclusions: This study reconceptualizes physicians as strategic agents shaping patient decision-making through purposeful self-disclosure. Different from existing studies treating physicians as passive recipients of ratings and reviews, our research demonstrates that physicians can strategically shape patient acquisition through self-disclosure breadth and depth. This study brings new insights to digital health markets by demonstrating that self-disclosure operates as a viable patient acquisition mechanism, wherein the DHL acts as a critical boundary condition. The findings have real-world implications: (1) physicians can leverage evidence-based disclosure strategies, (2) platforms should implement context-adaptive features, and (3) policymakers should prioritize digital infrastructure investments to enhance physicians' competitive capabilities and patient decision-making quality.
  •  

Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference

arXiv:2601.15333v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) possess strong representation and reasoning capabilities, but their application to structure-based drug design (SBDD) is limited by insufficient understanding of protein structures and unpredictable molecular generation. To address these challenges, we propose Exploration-Augmented Latent Inference for LLMs (ELILLM), a framework that reinterprets the LLM generation process as an encoding, latent space exploration, and decoding workflow. ELILLM explicitly explores portions of the design problem beyond the model's current knowledge while using a decoding module to handle familiar regions, generating chemically valid and synthetically reasonable molecules. In our implementation, Bayesian optimization guides the systematic exploration of latent embeddings, and a position-aware surrogate model efficiently predicts binding affinity distributions to inform the search. Knowledge-guided decoding further reduces randomness and effectively imposes chemical validity constraints. We demonstrate ELILLM on the CrossDocked2020 benchmark, showing strong controlled exploration and high binding affinity scores compared with seven baseline methods. These results demonstrate that ELILLM can effectively enhance LLMs capabilities for SBDD.
  •  

CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records

arXiv:2507.22533v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of clinical hallucination, as conventional grounding techniques such as Retrieval-Augmented Generation (RAG) do not adequately incorporate process-oriented clinical guidelines; and unreliable evaluation metrics that hinder the validation of AI systems in oncology. To address these issues, we propose CliCARE, a framework for Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records. The framework operates by transforming unstructured, longitudinal EHRs into patient-specific Temporal Knowledge Graphs (TKGs) to capture long-range dependencies, and then grounding the decision support process by aligning these real-world patient trajectories with a normative guideline knowledge graph. This approach provides oncologists with evidence-grounded decision support by generating a high-fidelity clinical summary and an actionable recommendation. We validated our framework using large-scale, longitudinal data from a private Chinese cancer dataset and the public English MIMIC-IV dataset. In these settings, CliCARE significantly outperforms baselines, including leading long-context LLMs and Knowledge Graph-enhanced RAG methods. The clinical validity of our results is supported by a robust evaluation protocol, which demonstrates a high correlation with assessments made by oncologists.
  •  

Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

arXiv:2601.04577v1 Announce Type: new Abstract: While AI innovation accelerates rapidly, the intellectual process behind breakthroughs -- how researchers identify gaps, synthesize prior work, and generate insights -- remains poorly understood. The lack of structured data on scientific reasoning hinders systematic analysis and development of AI research agents. We introduce Sci-Reasoning, the first dataset capturing the intellectual synthesis behind high-quality AI research. Using community-validated quality signals and an LLM-accelerated, human-verified pipeline, we trace Oral and Spotlight papers across NeurIPS, ICML, and ICLR (2023-2025) to its key predecessors, articulating specific reasoning links in a structured format. Our analysis identifies 15 distinct thinking patterns, with three dominant strategies accounting for 52.7%: Gap-Driven Reframing (24.2%), Cross-Domain Synthesis (18.0%), and Representation Shift (10.5%). The most powerful innovation recipes combine multiple patterns: Gap-Driven Reframing + Representation Shift, Cross-Domain Synthesis + Representation Shift, and Gap-Driven Reframing + Cross-Domain Synthesis. This dataset enables quantitative studies of scientific progress and provides structured reasoning trajectories for training the next generation AI research agents.
  •  

Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models

arXiv:2601.01321v1 Announce Type: new Abstract: Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.
  •  

Scaling Multimodal Search and Recommendation with Small Language Models via Upside-Down Reinforcement Learning

arXiv:2502.09854v2 Announce Type: replace-cross Abstract: In this work, we investigate how small language models (SLMs) can be scaled to support multimodal search and recommendation use cases while remaining efficient enough for real-time, resource-constrained deployments. We present a framework that combines upside-down reinforcement learning with synthetic data distillation from a large language model (Llama-3) to train a 100M-parameter GPT-2 model for multitask prompt generation. Despite being up to 80 times smaller than state-of-the-art large language models (LLMs), our SLM achieves relevance and diversity scores within 6% of competitive baselines such as Llama-3 8B, Qwen3 8B, and Ministral 8B. These results demonstrate that SLMs can effectively handle multimodal search and recommendation tasks, while dramatically reducing inference latency and memory overhead. Our study highlights the potential of lightweight models as practical engines for scalable multimodal discovery, bridging the gap between cutting-edge research and real-world multimodal applications such as media recommendations and creative content generation.
  •  

SMMILe enables accurate spatial quantification in digital pathology using multiple-instance learning

Nature Cancer, Published online: 19 November 2025; doi:10.1038/s43018-025-01060-8

Gao et al. present SMMILe, a multiple-instance learning-based tool that leverages whole-slide images for accurate spatial quantification without compromising on classification performance, and show it outperforms state-of-the-art methods.
  •  

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v2 Announce Type: replace Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.
  •  

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.
  •  

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
  •  

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

arXiv:2410.14879v4 Announce Type: replace-cross Abstract: Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate periods of raw sensor data to model momentary behaviors, (e.g., physical activity recognition), there still remains a significant gap in the translation of these sensing streams into meaningful, high-level, context-aware insights that are required for various applications (e.g., summarizing an individual's daily routine). To bridge this gap, experts often need to employ a context-driven sensemaking process in real-world studies to derive insights. This process often requires manual effort and can be challenging even for experienced researchers due to the complexity of human behaviors. We conducted three rounds of user studies with 21 experts to explore solutions to address challenges with sensemaking. We follow a human-centered design process to identify needs and design, iterate, build, and evaluate Vital Insight (VI), a novel, LLM-assisted, prototype system to enable human-in-the-loop inference (sensemaking) and visualizations of multi-modal passive sensing data from smartphones and wearables. Using the prototype as a technology probe, we observe experts' interactions with it and develop an expert sensemaking model that explains how experts move between direct data representations and AI-supported inferences to explore, question, and validate insights. Through this iterative process, we also synthesize and discuss a list of design implications for the design of future AI-augmented visualization systems to better assist experts' sensemaking processes in multi-modal health sensing data.
  •  

Language Ranker: A Lightweight Ranking framework for LLM Decoding

arXiv:2510.21883v1 Announce Type: cross Abstract: Conventional research on large language models (LLMs) has primarily focused on refining output distributions, while paying less attention to the decoding process that transforms these distributions into final responses. Recent advances, such as scaling the computation of inference time with reward models, have underscored the importance of decoding, but these methods often suffer from high computational costs and limited applicability. In this paper, we revisit LLM generation through the lens of recommender systems, conceptualizing the decoding process as analogous to the ranking stage in recommendation pipelines. From this perspective, we observe that both traditional decoding methods and reward models exhibit clear limitations such as redundancy. Motivated by this insight, we propose Language Ranker, a novel framework that introduces a lightweight module to rerank candidate responses using features extracted by the base model. Experiments across a wide range of tasks show that Language Ranker achieves performance comparable to large-scale reward models, while requiring only
  •  

Cell-free epigenomes enhanced fragmentomics-based model for early detection of lung cancer

Clin Transl Med. 2025 Feb;15(2):e70225. doi: 10.1002/ctm2.70225.

ABSTRACT

BACKGROUND: Lung cancer is a leading cause of cancer mortality, highlighting the need for innovative non-invasive early detection methods. Although cell-free DNA (cfDNA) analysis shows promise, its sensitivity in early-stage lung cancer patients remains a challenge. This study aimed to integrate insights from epigenetic modifications and fragmentomic features of cfDNA using machine learning to develop a more accurate lung cancer detection model.

METHODS: To address this issue, a multi-centre prospective cohort study was conducted, with participants harbouring suspicious malignant lung nodules and healthy volunteers recruited from two clinical centres. Plasma cfDNA was analysed for its epigenetic and fragmentomic profiles using chromatin immunoprecipitation sequencing, reduced representation bisulphite sequencing and low-pass whole-genome sequencing. Machine learning algorithms were then employed to integrate the multi-omics data, aiding in the development of a precise lung cancer detection model.

RESULTS: Cancer-related changes in cfDNA fragmentomics were significantly enriched in specific genes marked by cell-free epigenomes. A total of 609 genes were identified, and the corresponding cfDNA fragmentomic features were utilised to construct the ensemble model. This model achieved a sensitivity of 90.4% and a specificity of 83.1%, with an AUC of 0.94 in the independent validation set. Notably, the model demonstrated exceptional sensitivity for stage I lung cancer cases, achieving 95.1%. It also showed remarkable performance in detecting minimally invasive adenocarcinoma, with a sensitivity of 96.2%, highlighting its potential for early detection in clinical settings.

CONCLUSIONS: With feature selection guided by multiple epigenetic sequencing approaches, the cfDNA fragmentomics-based machine learning model demonstrated outstanding performance in the independent validation cohort. These findings highlight its potential as an effective non-invasive strategy for the early detection of lung cancer.

KEYPOINTS: Our study elucidated the regulatory relationships between epigenetic modifications and their effects on fragmentomic features. Identifying epigenetically regulated genes provided a critical foundation for developing the cfDNA fragmentomics-based machine learning model. The model demonstrated exceptional clinical performance, highlighting its substantial potential for translational application in clinical practice.

PMID:39909829 | PMC:PMC11798665 | DOI:10.1002/ctm2.70225

  •  
❌