❌

Reading view

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
  •  

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.
  •  

Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI

arXiv:2509.13345v2 Announce Type: replace-cross Abstract: Hallucination in generative AI is often treated as a technical failure to produce factually correct output. Yet this framing underrepresents the broader significance of hallucinated content in language models, which may appear fluent, persuasive, and contextually appropriate while conveying distortions that escape conventional accuracy checks. This paper critically examines how regulatory and evaluation frameworks have inherited a narrow view of hallucination, one that prioritises surface verifiability over deeper questions of meaning, influence, and impact. We propose a layered approach to understanding hallucination risks, encompassing epistemic instability, user misdirection, and social-scale effects. Drawing on interdisciplinary sources and examining instruments such as the EU AI Act and the GDPR, we show that current governance models struggle to address hallucination when it manifests as ambiguity, bias reinforcement, or normative convergence. Rather than improving factual precision alone, we argue for regulatory responses that account for languages generative nature, the asymmetries between system and user, and the shifting boundaries between information, persuasion, and harm.
  •  

Cancer-Associated Fibroblasts: Heterogeneity, Cancer Pathogenesis, and Therapeutic Targets

MedComm (2020). 2025 Jul 11;6(7):e70292. doi: 10.1002/mco2.70292. eCollection 2025 Jul.

ABSTRACT

Cancer-associated fibroblasts (CAFs) are functionally diverse stromal regulators that orchestrate tumor progression, metastasis, and therapy resistance through dynamic crosstalk within the tumor microenvironment (TME). Recent advances in single-cell multiomics and spatial transcriptomics have identified conserved CAF subtypes with distinct molecular signatures, spatial distributions, and context-dependent roles, highlighting their dual capacity to promote immunosuppression or restrain tumor growth. However, therapeutic strategies struggle to reconcile this functional duality, hindering clinical translation. This review systematically categorizes CAF subtypes by origin, biomarkers, and TME-specific functions, focusing on their roles in chemoresistance, maintenance of stemness, and formation of immunosuppressive niches. We evaluate emerging targeting approaches, including selective depletion of tumor-promoting subsets (e.g., fibroblast activation protein+ CAFs), epigenetic reprogramming toward antitumor phenotypes, and inhibition of CXCL12/CXCR4 or transforming growth factor-beta signaling pathways. Spatial multiomics-driven combinatorial therapies, such as the synergistic use of CAFs and immune checkpoint inhibitors, are highlighted as strategies to overcome microenvironment-driven resistance. By integrating CAF biology with translational advances, this work provides a roadmap for developing subtype-specific biomarkers and precision stromal therapies, directly informing efforts to disrupt tumor-stroma coevolution. Key concepts include spatial transcriptomics, stromal reprogramming, and tumor-stroma coevolution, offering actionable insights for both mechanistic research and clinical innovation.

PMID:40656546 | PMC:PMC12246558 | DOI:10.1002/mco2.70292

  •  
❌