❌

Normal view

FLORA: Unsupervised Knowledge Graph Alignment by Fuzzy Logic

arXiv:2510.20467v1 Announce Type: new Abstract: Knowledge graph alignment is the task of matching equivalent entities (that is, instances and classes) and relations across two knowledge graphs. Most existing methods focus on pure entity-level alignment, computing the similarity of entities in some embedding space. They lack interpretable reasoning and need training data to work. In this paper, we propose FLORA, a simple yet effective method that (1) is unsupervised, i.e., does not require training data, (2) provides a holistic alignment for entities and relations iteratively, (3) is based on fuzzy logic and thus delivers interpretable results, (4) provably converges, (5) allows dangling entities, i.e., entities without a counterpart in the other KG, and (6) achieves state-of-the-art results on major benchmarks.

Lost in Translation: Policymakers are not really listening to Citizen Concerns about AI

arXiv:2510.20568v1 Announce Type: new Abstract: The worlds people have strong opinions about artificial intelligence (AI), and they want policymakers to listen. Governments are inviting public comment on AI, but as they translate input into policy, much of what citizens say is lost. Policymakers are missing a critical opportunity to build trust in AI and its governance. This paper compares three countries, Australia, Colombia, and the United States, that invited citizens to comment on AI risks and policies. Using a landscape analysis, the authors examined how each government solicited feedback and whether that input shaped governance. Yet in none of the three cases did citizens and policymakers establish a meaningful dialogue. Governments did little to attract diverse voices or publicize calls for comment, leaving most citizens unaware or unprepared to respond. In each nation, fewer than one percent of the population participated. Moreover, officials showed limited responsiveness to the feedback they received, failing to create an effective feedback loop. The study finds a persistent gap between the promise and practice of participatory AI governance. The authors conclude that current approaches are unlikely to build trust or legitimacy in AI because policymakers are not adequately listening or responding to public concerns. They offer eight recommendations: promote AI literacy; monitor public feedback; broaden outreach; hold regular online forums; use innovative engagement methods; include underrepresented groups; respond publicly to input; and make participation easier.

MolBridge: Atom-Level Joint Graph Refinement for Robust Drug-Drug Interaction Event Prediction

arXiv:2510.20448v1 Announce Type: cross Abstract: Drug combinations offer therapeutic benefits but also carry the risk of adverse drug-drug interactions (DDIs), especially under complex molecular structures. Accurate DDI event prediction requires capturing fine-grained inter-drug relationships, which are critical for modeling metabolic mechanisms such as enzyme-mediated competition. However, existing approaches typically rely on isolated drug representations and fail to explicitly model atom-level cross-molecular interactions, limiting their effectiveness across diverse molecular complexities and DDI type distributions. To address these limitations, we propose MolBridge, a novel atom-level joint graph refinement framework for robust DDI event prediction. MolBridge constructs a joint graph that integrates atomic structures of drug pairs, enabling direct modeling of inter-drug associations. A central challenge in such joint graph settings is the potential loss of information caused by over-smoothing when modeling long-range atomic dependencies. To overcome this, we introduce a structure consistency module that iteratively refines node features while preserving the global structural context. This joint design allows MolBridge to effectively learn both local and global interaction outperforms state-of-the-art baselines, achieving superior performance across long-tail and inductive scenarios. patterns, yielding robust representations across both frequent and rare DDI types. Extensive experiments on two benchmark datasets show that MolBridge consistently. These results demonstrate the advantages of fine-grained graph refinement in improving the accuracy, robustness, and mechanistic interpretability of DDI event prediction.This work contributes to Web Mining and Content Analysis by developing graph-based methods for mining and analyzing drug-drug interaction networks.

User Perceptions of Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

arXiv:2510.20721v1 Announce Type: cross Abstract: Large language models (LLMs) have seen rapid adoption for tasks such as drafting emails, summarizing meetings, and answering health questions. In such uses, users may need to share private information (e.g., health records, contact details). To evaluate LLMs' ability to identify and redact such private information, prior work developed benchmarks (e.g., ConfAIde, PrivacyLens) with real-life scenarios. Using these benchmarks, researchers have found that LLMs sometimes fail to keep secrets private when responding to complex tasks (e.g., leaking employee salaries in meeting summaries). However, these evaluations rely on LLMs (proxy LLMs) to gauge compliance with privacy norms, overlooking real users' perceptions. Moreover, prior work primarily focused on the privacy-preservation quality of responses, without investigating nuanced differences in helpfulness. To understand how users perceive the privacy-preservation quality and helpfulness of LLM responses to privacy-sensitive scenarios, we conducted a user study with 94 participants using 90 scenarios from PrivacyLens. We found that, when evaluating identical responses to the same scenario, users showed low agreement with each other on the privacy-preservation quality and helpfulness of the LLM response. Further, we found high agreement among five proxy LLMs, while each individual LLM had low correlation with users' evaluations. These results indicate that the privacy and helpfulness of LLM responses are often specific to individuals, and proxy LLMs are poor estimates of how real users would perceive these responses in privacy-sensitive scenarios. Our results suggest the need to conduct user-centered studies on measuring LLMs' ability to help users while preserving privacy. Additionally, future research could investigate ways to improve the alignment between proxy LLMs and users for better estimation of users' perceived privacy and utility.

Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing

arXiv:2510.20727v1 Announce Type: cross Abstract: Objective: Fluoropyrimidines are widely prescribed for colorectal and breast cancers, but are associated with toxicities such as hand-foot syndrome and cardiotoxicity. Since toxicity documentation is often embedded in clinical notes, we aimed to develop and evaluate natural language processing (NLP) methods to extract treatment and toxicity information. Materials and Methods: We constructed a gold-standard dataset of 236 clinical notes from 204,165 adult oncology patients. Domain experts annotated categories related to treatment regimens and toxicities. We developed rule-based, machine learning-based (Random Forest, Support Vector Machine [SVM], Logistic Regression [LR]), deep learning-based (BERT, ClinicalBERT), and large language models (LLM)-based NLP approaches (zero-shot and error-analysis prompting). Models used an 80:20 train-test split. Results: Sufficient data existed to train and evaluate 5 annotated categories. Error-analysis prompting achieved optimal precision, recall, and F1 scores (F1=1.000) for treatment and toxicities extraction, whereas zero-shot prompting reached F1=1.000 for treatment and F1=0.876 for toxicities extraction.LR and SVM ranked second for toxicities (F1=0.937). Deep learning underperformed, with BERT (F1=0.873 treatment; F1= 0.839 toxicities) and ClinicalBERT (F1=0.873 treatment; F1 = 0.886 toxicities). Rule-based methods served as our baseline with F1 scores of 0.857 in treatment and 0.858 in toxicities. Discussion: LMM-based approaches outperformed all others, followed by machine learning methods. Machine and deep learning approaches were limited by small training data and showed limited generalizability, particularly for rare categories. Conclusion: LLM-based NLP most effectively extracted fluoropyrimidine treatment and toxicity information from clinical notes, and has strong potential to support oncology research and pharmacovigilance.

FieldGen: From Teleoperated Pre-Manipulation Trajectories to Field-Guided Data Generation

arXiv:2510.20774v1 Announce Type: cross Abstract: Large-scale and diverse datasets are vital for training robust robotic manipulation policies, yet existing data collection methods struggle to balance scale, diversity, and quality. Simulation offers scalability but suffers from sim-to-real gaps, while teleoperation yields high-quality demonstrations with limited diversity and high labor cost. We introduce FieldGen, a field-guided data generation framework that enables scalable, diverse, and high-quality real-world data collection with minimal human supervision. FieldGen decomposes manipulation into two stages: a pre-manipulation phase, allowing trajectory diversity, and a fine manipulation phase requiring expert precision. Human demonstrations capture key contact and pose information, after which an attraction field automatically generates diverse trajectories converging to successful configurations. This decoupled design combines scalable trajectory diversity with precise supervision. Moreover, FieldGen-Reward augments generated data with reward annotations to further enhance policy learning. Experiments demonstrate that policies trained with FieldGen achieve higher success rates and improved stability compared to teleoperation-based baselines, while significantly reducing human effort in long-term real-world data collection. Webpage is available at https://fieldgen.github.io/.

Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference

arXiv:2508.04586v4 Announce Type: replace-cross Abstract: Artificial Intelligence (AI) conferences are essential for advancing research, sharing knowledge, and fostering academic community. However, their rapid expansion has rendered the centralized conference model increasingly unsustainable. This paper offers a data-driven diagnosis of a structural crisis that threatens the foundational goals of scientific dissemination, equity, and community well-being. We identify four key areas of strain: (1) scientifically, with per-author publication rates more than doubling over the past decade to over 4.5 papers annually; (2) environmentally, with the carbon footprint of a single conference exceeding the daily emissions of its host city; (3) psychologically, with 71% of online community discourse reflecting negative sentiment and 35% referencing mental health concerns; and (4) logistically, with attendance at top conferences such as NeurIPS 2024 beginning to outpace venue capacity. These pressures point to a system that is misaligned with its core mission. In response, we propose the Community-Federated Conference (CFC) model, which separates peer review, presentation, and networking into globally coordinated but locally organized components, offering a more sustainable, inclusive, and resilient path forward for AI research.

VaultGemma: A Differentially Private Gemma Model

arXiv:2510.15001v2 Announce Type: replace-cross Abstract: We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemma 2 series, VaultGemma 1B represents a significant step forward in privacy-preserving large language models. We openly release this model to the community

A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist

arXiv:2510.19139v1 Announce Type: new Abstract: Despite the rapid expansion of Large Language Models (LLMs) in healthcare, the ability of these systems to assess clinical trial reporting according to CONSORT standards remains unclear, particularly with respect to their cognitive and reasoning strategies. This study applies a behavioral and metacognitive analytic approach with expert-validated data, systematically comparing two representative LLMs under three prompt conditions. Clear differences emerged in how the models approached various CONSORT items, and prompt types, including shifts in reasoning style, explicit uncertainty, and alternative interpretations shaped response patterns. Our results highlight the current limitations of these systems in clinical compliance automation and underscore the importance of understanding their cognitive adaptations and strategic behavior in developing more explainable and reliable medical AI.

MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration

arXiv:2510.19423v1 Announce Type: new Abstract: We introduce MSC-Bench, a large-scale benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents in a hierarchical Model-Context Protocol (MCP) ecosystem. Existing benchmarks often evaluate tools in isolation, ignoring challenges such as functional overlap and cross-server orchestration, leading to overly optimistic assessments. MSC-Bench addresses these gaps by constructing ground truth through 'equal function sets', allowing objective metrics such as F1 score and reducing the dependency on LLM-as-a-judge evaluation. Organized as a five-level curriculum, it systematically tests agent capabilities from single-tool orchestration to complex cross-server planning, and robustness to out-of-scope requests. Experiments reveal that rigid hierarchies can hinder performance without co-designed strategies, and even state-of-the-art agents exhibit systemic weaknesses in robustness. MSC-Bench provides a diagnostic framework to expose these limitations and guide the development of more capable and efficient tool-using agents. The benchmark and resources are publicly available at https://github.com/snooow1029/MSC_Bench.

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

arXiv:2510.19484v1 Announce Type: cross Abstract: The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular representation strategies during pretraining. To address these challenges, we introduce KnowMol-100K, a large-scale dataset with 100K fine-grained molecular annotations across multiple levels, bridging the gap between molecules and textual descriptions. Additionally, we propose chemically-informative molecular representation, effectively addressing limitations in existing molecular representation strategies. Building upon these innovations, we develop KnowMol, a state-of-the-art multi-modal molecular large language model. Extensive experiments demonstrate that KnowMol achieves superior performance across molecular understanding and generation tasks. GitHub: https://github.com/yzf-code/KnowMol Huggingface: https://hf.co/datasets/yzf1102/KnowMol-100K

RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning

arXiv:2510.14828v2 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based on Supervised Fine-Tuning (SFT) in planning tasks, they continue facing challenges in performing long-horizon manipulation tasks in complex real-world environments, owing to their restricted common sense and reasoning capabilities. Considering that aligning general-purpose vision language models to robotic planning tasks via supervised fine-tuning suffers from poor generalization and insufficient physical understanding, we propose RoboGPT-R1, a two-stage fine-tuning framework for embodied planning. In this framework, supervised training acquires foundational knowledge through expert sequences, followed by RL to address the model's shortcomings in visual-spatial understanding and reasoning. To achieve physical understanding and action sequence consistency in multi-step reasoning tasks, we design a rule-based reward function that simultaneously considers long-horizon performance and action constraint in the environment. The reasoning model, trained on Qwen2.5-VL-3B, significantly outperforms the larger-scale model, GPT-4o-mini, by 21.33% and surpasses other work trained on Qwen2.5-VL-7B by 20.33% on the EmbodiedBench benchmark.

ScholaWrite: A Dataset of End-to-End Scholarly Writing Process

arXiv:2502.02904v4 Announce Type: replace-cross Abstract: Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build writing assistants that truly align with writers' cognition, we must capture and decode the complete thought process behind how writers transform ideas into final texts. We present ScholaWrite, the first dataset of end-to-end scholarly writing, tracing the multi-month journey from initial drafts to final manuscripts. We contribute three key advances: (1) a Chrome extension that unobtrusively records keystrokes on Overleaf, enabling the collection of realistic, in-situ writing data; (2) a novel corpus of full scholarly manuscripts, enriched with fine-grained annotations of cognitive writing intentions. The dataset includes \LaTeX-based edits from five computer science preprints, capturing nearly 62K text changes over four months; and (3) analyses and insights into the micro-dynamics of scholarly writing, highlighting gaps between human writing processes and the current capabilities of large language models (LLMs) in providing meaningful assistance. ScholaWrite underscores the value of capturing end-to-end writing data to develop future writing assistants that support, not replace, the cognitive work of scientists.

Landscape of T-cell exhaustion heterogeneity and HBV integration in virus-related HCC revealed by whole-exome, transcriptome, and single-cell sequencing

JHEP Rep. 2025 Jul 10;7(11):101518. doi: 10.1016/j.jhepr.2025.101518. eCollection 2025 Nov.

ABSTRACT

BACKGROUND & AIMS: To enhance our understanding of the tumor immune microenvironment (TIME) in hepatocellular carcinoma (HCC), we investigated the heterogeneity of T-cell exhaustion and its association with HBV integrations and direct oncogenic potential in HCC.

METHODS: We conducted a multi-omics analysis, including single-cell RNA sequencing, whole-exome sequencing, whole-transcriptome sequencing, and next-generation sequencing (NGS)-based HBV integration analysis, in eight patients with virus-related HCC. For validation, bulk RNA sequencing and NGS-based HBV integration analysis were performed in an independent cohort (n = 106).

RESULTS: Based on the expression scores of exhaustion markers in effector CD8+ T cells, patients were classified into high (n = 2) and low (n = 6) exhaustion groups (p <0.001). The high-exhaustion group exhibited higher clonal expansion (Gini index: 0.83 vs. 0.48, p = 0.006) and sharing of CD8+ T effector memory and cycling T cells with elevated exhaustion markers. This group also showed increased clonal expansion of CD4+ regulatory T cells and follicular helper T cells (p <0.001) with higher PDCD1 expression. In addition, the high-exhaustion group had higher TP53 mutation rates and signature scores for proliferation subtypes compared with the low-exhaustion group, who predominantly harbored TERT mutations. Moreover, the high-exhaustion group demonstrated more pronounced HBV integrations with elevated intrahepatic covalently closed circular DNA (cccDNA) and pregenomic (pg)RNA levels. Similarly, in the validation cohort, the high-exhaustion group (n = 28) demonstrated stronger proliferation subtype signatures (p <0.001), along with higher HBV integrations, S-fusion transcripts, and an increased intrahepatic viral reservoir (cccDNA/pgRNA) (p <0.05) compared with the low-exhaustion group (n = 78).

CONCLUSIONS: Our study revealed the heterogeneity in T-cell exhaustion in the TIME of HCC, along with differences in HBV integrations and molecular subtypes. These findings provide insight into the intricate relationship between high exhaustion, proliferation subtype, increased HBV integrations, and enhanced HBV-induced oncogenic potential in virus-related HCC.

IMPACT AND IMPLICATIONS: This study provides a comprehensive immune landscape of T-cell exhaustion using multi-omics analysis, offering critical insights into T cell heterogeneity in virus-related HCC. It establishes a strong association between higher HBV integration, enhanced oncogenic potential, T-cell exhaustion, and proliferation subtypes in HCC. Our results also establish a basis for personalized therapies tailored to the immune-exhaustion status within the TIME of each patient with HCC.

PMID:41113120 | PMC:PMC12529496 | DOI:10.1016/j.jhepr.2025.101518

Single-cell multi-omics analysis reveals cancer regulatory elements of transcriptional programs and clinical implications

Cell Death Dis. 2025 Oct 21;16(1):746. doi: 10.1038/s41419-025-08060-7.

ABSTRACT

The regulatory mechanisms governing transcriptional programs in the cancer genome remain elusive, particularly those concerning cell-type specificity. We carefully curated single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data from eight distinct carcinoma tissues, including breast, skin, colon, endometrium, lung, ovary, liver, and kidney. Using single-cell multi-omics analysis, we identified extensive open chromatin regions and constructed peak-gene link networks, which can reveal distinct cancer gene regulation and genetic risks. We further explored conserved epigenetic regulation across cell types within cancer and elucidated their functional implications. Moreover, we identified cell-type-associated transcription factors (TFs) that regulate key cellular functions, such as the TEAD family of TFs, which widely control cancer-related signaling pathways in tumor cells. In colon cancer, we further identified tumor-specific TFs that are more highly activated in tumor cells than in normal epithelial cells, including CEBPG, LEF1, SOX4, TCF7, and TEAD4, which are pivotal in driving malignant transcriptional programs and represent potential therapeutic targets, as corroborated by single-cell sequencing data from multiple sources and in vitro experiments. Our findings provide a comprehensive understanding of the regulatory dynamics underlying carcinomas and offer valuable insights into potential therapeutic interventions.

PMID:41120274 | PMC:PMC12541060 | DOI:10.1038/s41419-025-08060-7

Exploring Patient Perspectives, Engagement, and Output Quality in Doctor-Supervised Use of Artificial Intelligence During Informed Consent Consultation With ChatGPT and Retrieval Augmented Generation (RAG): Quantitative Exploratory Study

Background: Comprehensive preoperative education is essential for optimizing outcomes and ensuring informed consent in patients undergoing total hip arthroplasty (THA). Emerging artificial intelligence (AI) tools, such as ChatGPT, offer scalable support for patient education, but their clinical application requires rigorous evaluation to ensure accuracy, safety, and trust. Objective: This study assessed patients’ preferences and satisfaction with AI-assisted informed consent in THA, comparing traditional physician consultations to those supported by native ChatGPT and a customized version enhanced with retrieval-augmented generation (RAG). It also examined how state anxiety and general attitudes toward AI affect preferences for AI-supported consent and whether RAG integration improves ChatGPT response quality. Methods: A total of 36 patients scheduled for elective THA were assigned to one of three groups (12 each): (1) standard physician-only consultations (control), (2) physician-assisted consultations supported by native ChatGPT, and (3) supported by ChatGPT enhanced through RAG. Data collection involved standardized Likert scale questionnaires assessing patient satisfaction with the consent process, perceived informedness, anxiety levels, and attitudes toward AI. The ChatGPT responses were independently evaluated by physicians for relevance, accuracy, clarity, completeness, adherence to evidence-based guidelines, and appropriate length. Instances of hallucinations, factually incorrect or misleading outputs, were identified and rated by severity. Statistical analyses compared outcomes across groups and explored associations. Results: Patients interacting with the ChatGPT+RAG model reported significantly higher satisfaction levels with information delivery (P=.01) and perceived level of informedness (P=.01) than those using the native ChatGPT model. The mean number of patient questions in the control group was 20, compared with 39 in the native ChatGPT group (P=.06) and 52 in the ChatGPT+RAG group (P=.002). The majority of participants across all groups preferred a human clinician providing less accurate information over a more accurate AI-only assistant. These preferences were not influenced by sociodemographic variables (age, gender, and education), health literacy, state anxiety, or general attitudes toward AI. The ChatGPT+RAG model outperformed the native ChatGPT model across all evaluated response quality dimensions (all P<.01) and exhibited a significantly lower hallucination rate (5/52, 10% versus 15/39, 38%; P=.002). Conclusions: Integrating RAG with ChatGPT significantly improves the quality, clarity, and reliability of preoperative information, enhancing patient satisfaction and engagement beyond native ChatGPT. However, patients maintain a strong preference for physician-led informed consent, underscoring the role of AI chatbots as complementary tools rather than replacements. These findings support the cautious adoption of customized AI assistants to augment, not substitute, human interaction in surgical consent processes. Trial Registration:

Assessing Large Language Models in Building a Structured Dataset From AskDocs Subreddit Data: Methodological Study

Background: In an era marked by the blooming reliance on digital platforms for healthcare consultation, the subreddit r/AskDocs has emerged as a pivotal forum. However, the vast, unstructured nature of forum data presents a formidable challenge; the extraction and meaningful analysis of such data require advanced tools that can navigate the complexities of language and context inherent in user-generated content. Objective: Our objective was to evaluate employing Large Language Models (LLMs) to systematically transform the rich, unstructured textual data from AskDocs into a structured dataset, an approach that aligns more closely with human cognitive processes compared to traditional data extraction methods. Methods: We developed a dataset of Reddit posts from r/AskDocs by extracting key information via human annotators. Then using specially engineered prompts we used state-of-the-art Large Language Models (LLMs) to extract data from posts and compared the results. The variation in the LLMs were further compared to the humans to show similarity. Results: Our findings indicate that LLMs not only match but, in several aspects, surpass even highly educated humans in extracting information, including both demographic and context details, from unstructured texts. Conclusions: This study not only validates the use of LLMs for analyzing digital healthcare communications but also opens new avenues for understanding online behaviors and interactions, signaling a shift towards more sophisticated methodologies in digital research and practice.
❌