Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration
arXiv:2510.19423v1 Announce Type: new Abstract: We introduce MSC-Bench, a large-scale benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents in a hierarchical Model-Context Protocol (MCP) ecosystem. Existing benchmarks often evaluate tools in isolation, ignoring challenges such as functional overlap and cross-server orchestration, leading to overly optimistic assessments. MSC-Bench addresses these gaps by constructing ground truth through 'equal function sets', allowing
-
cs.AI, q-bio.NC updates on arXiv.org
-
Small Language Models Offer Significant Potential for Science Community
arXiv:2510.18890v1 Announce Type: cross Abstract: Recent advancements in natural language processing, particularly with large language models (LLMs), are transforming how scientists engage with the literature. While the adoption of LLMs is increasing, concerns remain regarding potential information biases and computational costs. Rather than LLMs, I developed a framework to evaluate the feasibility of precise, rapid, and cost-effective information retrieval from extensive geoscience literature
Small Language Models Offer Significant Potential for Science Community
-
cs.AI, q-bio.NC updates on arXiv.org
-
"Over-the-Hood" AI Inclusivity Bugs and How 3 AI Product Teams Found and Fixed Them
arXiv:2510.19033v1 Announce Type: cross Abstract: While much research has shown the presence of AI's "under-the-hood" biases (e.g., algorithmic, training data, etc.), what about "over-the-hood" inclusivity biases: barriers in user-facing AI products that disproportionately exclude users with certain problem-solving approaches? Recent research has begun to report the existence of such biases -- but what do they look like, how prevalent are they, and how can developers find and fix them? To find
"Over-the-Hood" AI Inclusivity Bugs and How 3 AI Product Teams Found and Fixed Them
-
cs.AI, q-bio.NC updates on arXiv.org
-
Interpretable Question Answering with Knowledge Graphs
arXiv:2510.19181v1 Announce Type: cross Abstract: This paper presents a question answering system that operates exclusively on a knowledge graph retrieval without relying on retrieval augmented generation (RAG) with large language models (LLMs). Instead, a small paraphraser model is used to paraphrase the entity relationship edges retrieved from querying the knowledge graph. The proposed pipeline is divided into two main stages. The first stage involves pre-processing a document to generate set
Interpretable Question Answering with Knowledge Graphs
-
cs.AI, q-bio.NC updates on arXiv.org
-
KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
arXiv:2510.19484v1 Announce Type: cross Abstract: The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular representation strategies during pretraining. To address these challenges, we introduce KnowMol-100K, a large-scale dataset with 100K fine-grained mo
KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
-
cs.AI, q-bio.NC updates on arXiv.org
-
Insights into the Unknown: Federated Data Diversity Analysis on Molecular Data
arXiv:2510.19535v1 Announce Type: cross Abstract: AI methods are increasingly shaping pharmaceutical drug discovery. However, their translation to industrial applications remains limited due to their reliance on public datasets, lacking scale and diversity of proprietary pharmaceutical data. Federated learning (FL) offers a promising approach to integrate private data into privacy-preserving, collaborative model training across data silos. This federated data access complicates important data-c
Insights into the Unknown: Federated Data Diversity Analysis on Molecular Data
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Goal-Driven Survey on Root Cause Analysis
arXiv:2510.19593v1 Announce Type: cross Abstract: Root Cause Analysis (RCA) is a crucial aspect of incident management in large-scale cloud services. While the term root cause analysis or RCA has been widely used, different studies formulate the task differently. This is because the term "RCA" implicitly covers tasks with distinct underlying goals. For instance, the goal of localizing a faulty service for rapid triage is fundamentally different from identifying a specific functional bug for a d
A Goal-Driven Survey on Root Cause Analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation
arXiv:2510.19799v1 Announce Type: cross Abstract: Public and nonprofit organizations often hesitate to adopt AI tools because most models are opaque even though standard approaches typically analyze aggregate patterns rather than offering actionable, case-level guidance. This study tests a practitioner-in-the-loop workflow that pairs transparent decision-tree models with large language models (LLMs) to improve predictive accuracy, interpretability, and the generation of practical insights. Usin
Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation
-
cs.AI, q-bio.NC updates on arXiv.org
-
RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning
arXiv:2510.14828v2 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based on Supervised Fine-Tuning (SFT) in planning tasks, they continue facing challenges in performing long-horizon manipulation tasks in complex real-world environments, owing to their restricted common sense an
RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI
arXiv:2510.16206v2 Announce Type: replace Abstract: Since the rapid expansion of large language models (LLMs), people have begun to rely on them for information retrieval. While traditional search engines display ranked lists of sources shaped by search engine optimization (SEO), advertising, and personalization, LLMs typically provide a synthesized response that feels singular and authoritative. While both approaches carry risks of bias and omission, LLMs may amplify the effect by collapsing m
The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI
-
cs.AI, q-bio.NC updates on arXiv.org
-
LICO: Large Language Models for In-Context Molecular Optimization
arXiv:2406.18851v2 Announce Type: replace-cross Abstract: Optimizing black-box functions is a fundamental problem in science and engineering. To solve this problem, many approaches learn a surrogate function that estimates the underlying objective from limited historical evaluations. Large Language Models (LLMs), with their strong pattern-matching capabilities via pretraining on vast amounts of data, stand out as a potential candidate for surrogate modeling. However, directly prompting a pretra
LICO: Large Language Models for In-Context Molecular Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
ScholaWrite: A Dataset of End-to-End Scholarly Writing Process
arXiv:2502.02904v4 Announce Type: replace-cross Abstract: Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build writing assistants that truly align with writers' cognition, we must capture and decode the complete thought process behind how writers transform ideas into final texts. We present ScholaWrite, the first dataset of end-to-end scholarly writing, tracing the mul
ScholaWrite: A Dataset of End-to-End Scholarly Writing Process
-
cs.AI, q-bio.NC updates on arXiv.org
-
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
arXiv:2503.12730v5 Announce Type: replace-cross Abstract: Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQL generation as an ideal task to study, as it combines the formal structure of toy tasks with real-world complexity. We introduce TinySQL, a synthetic dataset, progressing from basic to advanced SQL operations, and train models ranging from 33M to 1B parameters
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Quantum Natural Language Processing: A Comprehensive Review of Models, Methods, and Applications
arXiv:2504.09909v2 Announce Type: replace-cross Abstract: In recent developments, deep learning methodologies applied to Natural Language Processing (NLP) have revealed a paradox: They improve performance but demand considerable data and resources for their training. Alternatively, quantum computing exploits the principles of quantum mechanics to overcome the computational limitations of current methodologies, thereby establishing an emerging field known as quantum natural language processing (
Quantum Natural Language Processing: A Comprehensive Review of Models, Methods, and Applications
-
cs.AI, q-bio.NC updates on arXiv.org
-
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
arXiv:2505.07897v3 Announce Type: replace-cross Abstract: Context lengths for models have grown rapidly, from thousands to millions of tokens in just a few years. The extreme context sizes of modern long-context models have made it difficult to construct realistic long-context benchmarks -- not only due to the cost of collecting million-context tasks but also in identifying realistic scenarios that require significant contexts. We identify code comprehension and repair as a natural testbed and
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
-
cs.AI, q-bio.NC updates on arXiv.org
-
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
arXiv:2506.16895v2 Announce Type: replace-cross Abstract: Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, existing models typically rely on millions of paired multimodal samples, which are prohibitively expensive or infeasible to obtain in many domains. In this work, we explore the feasibility of building multimodal models with limited amount of paired data by aligni
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
-
cs.AI, q-bio.NC updates on arXiv.org
-
ACT: Agentic Classification Tree
arXiv:2509.26433v2 Announce Type: replace-cross Abstract: When used in high-stakes settings, AI systems are expected to produce decisions that are transparent, interpretable, and auditable, a requirement increasingly expected by regulations. Decision trees such as CART provide clear and verifiable rules, but they are restricted to structured tabular data and cannot operate directly on unstructured inputs such as text. In practice, large language models (LLMs) are widely used for such data, yet
ACT: Agentic Classification Tree
-
Nature Cancer
-
Radiopharma pipeline builds ahead of key data
Nature Cancer, Published online: 01 October 2025; doi:10.1038/s43018-025-01044-8Investment continues as the field awaits phase 3 readouts for next-generation radioligand therapies in 2026.
Radiopharma pipeline builds ahead of key data
Nature Cancer, Published online: 01 October 2025; doi:10.1038/s43018-025-01044-8
Investment continues as the field awaits phase 3 readouts for next-generation radioligand therapies in 2026.-
Omics in Hepatocellular
-
Landscape of T-cell exhaustion heterogeneity and HBV integration in virus-related HCC revealed by whole-exome, transcriptome, and single-cell sequencing
JHEP Rep. 2025 Jul 10;7(11):101518. doi: 10.1016/j.jhepr.2025.101518. eCollection 2025 Nov.ABSTRACTBACKGROUND & AIMS: To enhance our understanding of the tumor immune microenvironment (TIME) in hepatocellular carcinoma (HCC), we investigated the heterogeneity of T-cell exhaustion and its association with HBV integrations and direct oncogenic potential in HCC.METHODS: We conducted a multi-omics analysis, including single-cell RNA sequencing, whole-exome sequencing, whole-transcriptome sequenc
Landscape of T-cell exhaustion heterogeneity and HBV integration in virus-related HCC revealed by whole-exome, transcriptome, and single-cell sequencing
JHEP Rep. 2025 Jul 10;7(11):101518. doi: 10.1016/j.jhepr.2025.101518. eCollection 2025 Nov.
ABSTRACT
BACKGROUND & AIMS: To enhance our understanding of the tumor immune microenvironment (TIME) in hepatocellular carcinoma (HCC), we investigated the heterogeneity of T-cell exhaustion and its association with HBV integrations and direct oncogenic potential in HCC.
METHODS: We conducted a multi-omics analysis, including single-cell RNA sequencing, whole-exome sequencing, whole-transcriptome sequencing, and next-generation sequencing (NGS)-based HBV integration analysis, in eight patients with virus-related HCC. For validation, bulk RNA sequencing and NGS-based HBV integration analysis were performed in an independent cohort (n = 106).
RESULTS: Based on the expression scores of exhaustion markers in effector CD8+ T cells, patients were classified into high (n = 2) and low (n = 6) exhaustion groups (p <0.001). The high-exhaustion group exhibited higher clonal expansion (Gini index: 0.83 vs. 0.48, p = 0.006) and sharing of CD8+ T effector memory and cycling T cells with elevated exhaustion markers. This group also showed increased clonal expansion of CD4+ regulatory T cells and follicular helper T cells (p <0.001) with higher PDCD1 expression. In addition, the high-exhaustion group had higher TP53 mutation rates and signature scores for proliferation subtypes compared with the low-exhaustion group, who predominantly harbored TERT mutations. Moreover, the high-exhaustion group demonstrated more pronounced HBV integrations with elevated intrahepatic covalently closed circular DNA (cccDNA) and pregenomic (pg)RNA levels. Similarly, in the validation cohort, the high-exhaustion group (n = 28) demonstrated stronger proliferation subtype signatures (p <0.001), along with higher HBV integrations, S-fusion transcripts, and an increased intrahepatic viral reservoir (cccDNA/pgRNA) (p <0.05) compared with the low-exhaustion group (n = 78).
CONCLUSIONS: Our study revealed the heterogeneity in T-cell exhaustion in the TIME of HCC, along with differences in HBV integrations and molecular subtypes. These findings provide insight into the intricate relationship between high exhaustion, proliferation subtype, increased HBV integrations, and enhanced HBV-induced oncogenic potential in virus-related HCC.
IMPACT AND IMPLICATIONS: This study provides a comprehensive immune landscape of T-cell exhaustion using multi-omics analysis, offering critical insights into T cell heterogeneity in virus-related HCC. It establishes a strong association between higher HBV integration, enhanced oncogenic potential, T-cell exhaustion, and proliferation subtypes in HCC. Our results also establish a basis for personalized therapies tailored to the immune-exhaustion status within the TIME of each patient with HCC.
PMID:41113120 | PMC:PMC12529496 | DOI:10.1016/j.jhepr.2025.101518
-
Omics In Lung
-
Single-cell multi-omics analysis reveals cancer regulatory elements of transcriptional programs and clinical implications
Cell Death Dis. 2025 Oct 21;16(1):746. doi: 10.1038/s41419-025-08060-7.ABSTRACTThe regulatory mechanisms governing transcriptional programs in the cancer genome remain elusive, particularly those concerning cell-type specificity. We carefully curated single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data from eight distinct carcinoma tissues, including breast, skin, colon, endometrium, lung, ovary, liver, and kidney. Using s
Single-cell multi-omics analysis reveals cancer regulatory elements of transcriptional programs and clinical implications
Cell Death Dis. 2025 Oct 21;16(1):746. doi: 10.1038/s41419-025-08060-7.
ABSTRACT
The regulatory mechanisms governing transcriptional programs in the cancer genome remain elusive, particularly those concerning cell-type specificity. We carefully curated single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data from eight distinct carcinoma tissues, including breast, skin, colon, endometrium, lung, ovary, liver, and kidney. Using single-cell multi-omics analysis, we identified extensive open chromatin regions and constructed peak-gene link networks, which can reveal distinct cancer gene regulation and genetic risks. We further explored conserved epigenetic regulation across cell types within cancer and elucidated their functional implications. Moreover, we identified cell-type-associated transcription factors (TFs) that regulate key cellular functions, such as the TEAD family of TFs, which widely control cancer-related signaling pathways in tumor cells. In colon cancer, we further identified tumor-specific TFs that are more highly activated in tumor cells than in normal epithelial cells, including CEBPG, LEF1, SOX4, TCF7, and TEAD4, which are pivotal in driving malignant transcriptional programs and represent potential therapeutic targets, as corroborated by single-cell sequencing data from multiple sources and in vitro experiments. Our findings provide a comprehensive understanding of the regulatory dynamics underlying carcinomas and offer valuable insights into potential therapeutic interventions.
PMID:41120274 | PMC:PMC12541060 | DOI:10.1038/s41419-025-08060-7