❌

Normal view

Cognition on Graph: Navigating Massive Knowledge Space via Cognitive Cycles and Bidirectional Graph-Text Synergy

arXiv:2609.12791v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has empowered Large Language Models (LLMs) to tackle knowledge-intensive tasks. However, navigating global, heterogeneous knowledge bases (large-scale knowledge graphs and text corpora) for complex reasoning remains a challenge. Existing methods typically employ reactive, graph-driven exploration strategies, which blindly follow graph topology without adapting to the question context or evolving exploration progress, and lack deep bidirectional synergy between graph and text. To address these limitations, we propose CoG (Cognition on Graph), a cognitive-inspired, training-free framework for adaptive knowledge exploration. Drawing inspiration from human problem-solving, CoG performs a continuous plan-explore-reflect cycle, where it proactively formulates investigation plans, performs dual-source retrieval, and dynamically reflects on progress to adjust strategies. Crucially, it establishes deep bidirectional synergy between structured graph and unstructured text, where entities extracted from text dynamically guide graph exploration to bridge knowledge gaps. Extensive experiments on seven multi-hop QA benchmarks demonstrate that CoG significantly outperforms state-of-the-art methods while achieving superior exploration efficiency. Our code and datasets are available at https://github.com/zhougengxian/CoG.

EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding

arXiv:2609.09728v1 Announce Type: cross Abstract: Source-level analysis of interictal epileptiform discharges (IEDs) is relevant to presurgical evaluation and treatment planning because it helps characterize where epileptiform activity is likely to arise. Beyond detecting whether an IED is present, this setting requires assigning IED-positive activity to clinically meaningful brain-region categories. This setting is challenging because source-region evidence in short electroencephalography (EEG) windows can be subtle, partial, and affected by subject variability, class imbalance, and imperfect multimodal context. We present EEGBind, an EEG-centric multimodal binding framework for five-class source-level IED classification. EEGBind treats EEG as the primary modality and binds synchronized video-context features around an EEG-centric representation. Instead of relying on early or overly strong multimodal fusion, which may perturb the source-sensitive EEG representation, EEGBind uses video context as auxiliary evidence for robust classification. A view-consistent repair stage is further used to improve hidden-set robustness while preserving the learned source-class boundary. On the NeuroMM 2026 Grand Challenge Track 3 NMM-Source-IED benchmark, EEGBind achieves 0.8395 on weighted-F1 and outperforms strong competitors. These results support EEG-centric multimodal binding as a practical strategy for source-level IED classification. The open-source code is available at https://github.com/HKUSTGZ-ML4Health-Lab/NeuroMM2026_IED_Detection.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Generative AI for Analysts Jian Xue · Qian Zhang · Wu Zhu
    arXiv:2512.19705v2 Announce Type: replace-cross Abstract: We study how generative artificial intelligence (GenAI) reshapes financial analysts' information production. Using the 2023 integration of GenAI into FACTSET as a plausibly exogenous change in AI access, we find that FACTSET-associated reports become markedly richer--featuring 26% more distinct information sources, 24% broader topical coverage, and 21% more analytical methods--while also improving timeliness. However, these gains do not
     

Generative AI for Analysts

10 September 2026 at 12:00
arXiv:2512.19705v2 Announce Type: replace-cross Abstract: We study how generative artificial intelligence (GenAI) reshapes financial analysts' information production. Using the 2023 integration of GenAI into FACTSET as a plausibly exogenous change in AI access, we find that FACTSET-associated reports become markedly richer--featuring 26% more distinct information sources, 24% broader topical coverage, and 21% more analytical methods--while also improving timeliness. However, these gains do not uniformly improve decision quality: relative forecast accuracy declines when analysts face greater information-processing demands. Yet, a machine-learning benchmark processing the same observable inputs shows no analogous deterioration, pointing to a human processing constraint rather than poorer underlying information. Placebo tests using other data vendors make a common platform-wide technology trend unlikely. Overall, GenAI relaxes information-acquisition constraints while making human attention a more important bottleneck.

Integrated single-cell and bulk RNA sequencing reveals novel biomarkers of invasive adenocarcinoma subtypes in lung adenocarcinoma

Transl Cancer Res. 2026 Apr 30;15(4):314. doi: 10.21037/tcr-2025-aw-2503. Epub 2026 Mar 20.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) is one of the most common lung cancer subtypes worldwide, and its aggressive subtype invasive adenocarcinoma (IAC) has low survival rates. The precise identification of IAC is vital for the clinical diagnosis and treatment. The purpose of this study is to identify novel biomarkers for LUAD using single-cell and bulk RNA sequencing, so as to provide theoretical basis and practical support for the diagnosis, treatment and prognosis evaluation of lung invasive adenocarcinoma.

METHODS: We employed a combination of transcriptomic analysis and single-cell analysis to investigate the molecular characteristics and immune microenvironment of four subtypes of LUAD, including atypical adenomatous hyperplasia (AAH), adenocarcinoma in situ (AIS), minimally invasive adenocarcinoma (MIA), and IAC, with the aim of screening for biomarkers to differentiate pre-invasive lesions from invasive lesions.

RESULTS: Transcriptomic and single-cell analyses revealed that IAC subtypes demonstrated the most substantial molecular differences, particularly in immune cell infiltration and immune-related gene expression. Three genes-CD27, TIGIT, and TNFRSF18-that were significantly upregulated in IAC, predominantly expressed in immune cells and closely linked to immune regulatory pathways. We further analyzed T cell subpopulations in the IAC subtype and explored the expression of transcription factors (TFs) corresponding to these three genes, revealing their critical roles in immune cell function. Additionally, communication between T cells and other cells showed significantly enhanced signaling pathways, particularly those related to immune co-stimulatory molecules and inflammation pathways. Immunohistochemical validation of clinical samples showed that these three genes have high diagnostic value in IAC subtypes. These findings establish a crucial biological foundation for diagnosis, classification, and immunotherapy of LUAD, which contributes to the development of individualized treatment strategies.

CONCLUSIONS: This study identifies a three-gene signature (CD27, TIGIT, and TNFRSF18) that not only distinguishes invasive from pre-invasive LUAD with high precision by capturing the immune checkpoint disequilibrium characteristic of IAC, but also provides a clinically actionable biomarker panel for preoperative diagnosis and personalized immunotherapy strategies.

PMID:42180871 | PMC:PMC13190665 | DOI:10.21037/tcr-2025-aw-2503

Bibliometric analysis of lung cancer organoid research: trends and emerging areas of study

J Thorac Dis. 2026 Apr 30;18(4):406. doi: 10.21037/jtd-2026-0547. Epub 2026 Apr 27.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide, posing a substantial global health burden. Despite advances in early detection, molecular profiling, and targeted therapies, patient outcomes remain unsatisfactory due to tumor heterogeneity, therapeutic resistance, and the lack of reliable preclinical models. In recent years, lung cancer organoids (LCOs), patient-derived three-dimensional (3D) culture systems, have demonstrated the ability to preserve the histological architecture and genomic features of primary tumors more faithfully than conventional models, making them a promising platform for translational research and precision medicine. This study aims to quantitatively evaluate the global research output, identify major contributors and collaboration patterns, and systematically uncover research hotspots and emerging trends in the field of LCOs through bibliometric analysis.

METHODS: A systematic bibliometric analysis was conducted using publications on LCOs retrieved from the Web of Science Core Collection (WoSCC). Articles published between 2015 and 2024 were included. A total of 356 publications were analyzed. Publication outputs, country and institutional contributions, collaboration networks, and keyword co-occurrence were evaluated using Bibliometrix (R package), VOSviewer, and CiteSpace.

RESULTS: The number of publications on LCOs has increased steadily over the past decade, reflecting growing research interest and technological advancement. China and the United States were identified as the leading contributors, accounting for the majority of publications, while Germany, South Korea, and Japan also demonstrated strong research capacity and active collaboration. Keyword and thematic analyses revealed several major research hotspots, including personalized medicine, drug response and resistance mechanisms, tumor microenvironment modeling, and immune-related interactions. Burst keyword analysis further identified emerging trends, such as co-culture systems, immunotherapy evaluation, and the integration of LCOs with high-throughput screening and multi-omics approaches.

CONCLUSIONS: LCOs have evolved into a versatile platform bridging basic research and clinical applications in lung cancer. This study provides a comprehensive overview of the current research landscape and highlights emerging directions in the field. Future research should focus on methodological standardization, optimization of organoid construction and evaluation, integration with multi-omics and immune models, and strengthened international collaboration to facilitate clinical translation and improve patient outcomes.

PMID:42182656 | PMC:PMC13190222 | DOI:10.21037/jtd-2026-0547

Bibliometric analysis of lung cancer organoid research: trends and emerging areas of study

25 May 2026 at 18:00

J Thorac Dis. 2026 Apr 30;18(4):406. doi: 10.21037/jtd-2026-0547. Epub 2026 Apr 27.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide, posing a substantial global health burden. Despite advances in early detection, molecular profiling, and targeted therapies, patient outcomes remain unsatisfactory due to tumor heterogeneity, therapeutic resistance, and the lack of reliable preclinical models. In recent years, lung cancer organoids (LCOs), patient-derived three-dimensional (3D) culture systems, have demonstrated the ability to preserve the histological architecture and genomic features of primary tumors more faithfully than conventional models, making them a promising platform for translational research and precision medicine. This study aims to quantitatively evaluate the global research output, identify major contributors and collaboration patterns, and systematically uncover research hotspots and emerging trends in the field of LCOs through bibliometric analysis.

METHODS: A systematic bibliometric analysis was conducted using publications on LCOs retrieved from the Web of Science Core Collection (WoSCC). Articles published between 2015 and 2024 were included. A total of 356 publications were analyzed. Publication outputs, country and institutional contributions, collaboration networks, and keyword co-occurrence were evaluated using Bibliometrix (R package), VOSviewer, and CiteSpace.

RESULTS: The number of publications on LCOs has increased steadily over the past decade, reflecting growing research interest and technological advancement. China and the United States were identified as the leading contributors, accounting for the majority of publications, while Germany, South Korea, and Japan also demonstrated strong research capacity and active collaboration. Keyword and thematic analyses revealed several major research hotspots, including personalized medicine, drug response and resistance mechanisms, tumor microenvironment modeling, and immune-related interactions. Burst keyword analysis further identified emerging trends, such as co-culture systems, immunotherapy evaluation, and the integration of LCOs with high-throughput screening and multi-omics approaches.

CONCLUSIONS: LCOs have evolved into a versatile platform bridging basic research and clinical applications in lung cancer. This study provides a comprehensive overview of the current research landscape and highlights emerging directions in the field. Future research should focus on methodological standardization, optimization of organoid construction and evaluation, integration with multi-omics and immune models, and strengthened international collaboration to facilitate clinical translation and improve patient outcomes.

PMID:42182656 | PMC:PMC13190222 | DOI:10.21037/jtd-2026-0547

MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding

arXiv:2604.00513v2 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language models (MLLMs) have driven significant progress in product understanding, they are typically employed as feature extractors that implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. Therefore, we argue that leveraging the reasoning capabilities of MLLMs to explicitly model fine-grained product attributes holds significant potential. Nevertheless, achieving this goal remains non-trivial due to several key challenges: (i) long-context reasoning tends to dilute the model's attention to salient information in the raw input; (ii) supervised fine-tuning (SFT) primarily encourages rigid imitation, limiting the exploration of effective reasoning strategies; and (iii) fine-grained details are progressively attenuated during forward propagation. To address these issues, we propose MOON3.0, the first reasoning-aware MLLM-based model for product representation learning. Our method (1) employs a multi-head modality fusion module to adaptively integrate raw signals; (2) incorporates a joint contrastive and reinforcement learning framework to autonomously explore more effective reasoning strategies; and (3) introduces a fine-grained residual enhancement module to progressively preserve local details throughout the network. Additionally, we release a large-scale multimodal e-commerce benchmark MBE3.0. Experimentally, our model demonstrates state-of-the-art zero-shot performance across various downstream tasks on both our benchmark and public datasets.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

arXiv:2511.12449v2 Announce Type: replace-cross Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of the intrinsic alignment relationships among visual and textual information within a product; and (iii) limited handling of noise in e-commerce multimodal data. To address these, we propose MOON2.0, a dynamic modality-balanced MultimOdal representation learning framework for e-commerce prOduct uNderstanding. It comprises: (1) a Modality-driven Mixture-of-Experts (MoE) that adaptively processes input samples by their modality composition, enabling Multimodal Joint Learning to mitigate the modality imbalance; (2) a Dual-level Alignment method to better leverage semantic alignment properties inside individual products; and (3) an MLLM-based Image-text Co-augmentation strategy that integrates textual enrichment with visual expansion, coupled with Dynamic Sample Filtering to improve training data quality. We further release MBE2.0, a co-augmented Multimodal representation Benchmark for E-commerce representation learning and evaluation at https://huggingface.co/datasets/ZHNie/MBE2.0. Experiments show that MOON2.0 delivers state-of-the-art zero-shot performance on MBE2.0 and multiple public datasets. Furthermore, attention-based heatmap visualization provides qualitative evidence of improved multimodal alignment of MOON2.0.

TRACE: A Multi-Agent System for Autonomous Physical Reasoning in Seismological

arXiv:2603.21152v2 Announce Type: replace-cross Abstract: Inferring the physical mechanisms that govern earthquake sequences from indirect geophysical observations remains difficult, particularly across tectonically distinct environments where similar seismic patterns can reflect different underlying processes. Current interpretations rely heavily on the expert synthesis of catalogs, spatiotemporal statistics, and candidate physical models, limiting reproducibility and the systematic transfer of insight across settings. Here we present TRACE (Trans-perspective Reasoning and Automated Comprehensive Evaluator), a multi-agent system that combines large language model planning with formal seismological constraints to derive auditable, physically grounded mechanistic inference from raw observations. Applied to the 2019 Ridgecrest sequence, TRACE autonomously identifies stress-perturbation-induced delayed triggering, resolving the cascading interaction between the Mw 6.4 and Mw 7.1 mainshocks; in the Santorini-Kolumbo case, the system identifies a structurally guided intrusion model, distinguishing fault-channeled episodic migration from the continuous propagation expected in homogeneous crustal failure. By providing a generalizable logical infrastructure for interpreting heterogeneous seismic phenomena, TRACE advances the field from expert-dependent analysis toward knowledge-guided autonomous discovery in Earth sciences.

Variational Learning of Gaussian Process Latent Variable Models through Stochastic Gradient Annealed Importance Sampling

arXiv:2408.06710v3 Announce Type: replace-cross Abstract: Gaussian Process Latent Variable Models (GPLVMs) have become increasingly popular for unsupervised tasks such as dimensionality reduction and missing data recovery due to their flexibility and non-linear nature. An importance-weighted version of the Bayesian GPLVMs has been proposed to obtain a tighter variational bound. However, this version of the approach is primarily limited to analyzing simple data structures, as the generation of an effective proposal distribution can become quite challenging in high-dimensional spaces or with complex data sets. In this work, we propose an Annealed Importance Sampling (AIS) approach to address these issues. By transforming the posterior into a sequence of intermediate distributions using annealing, we combine the strengths of Sequential Monte Carlo samplers and VI to explore a wider range of posterior distributions and gradually approach the target distribution. We further propose an efficient algorithm by reparameterizing all variables in the evidence lower bound (ELBO). Experimental results on both toy and image datasets demonstrate that our method outperforms state-of-the-art methods in terms of tighter variational bounds, higher log-likelihoods, and more robust convergence.

ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue

arXiv:2603.02216v1 Announce Type: cross Abstract: Effective information seeking in multi-turn medical dialogues is critical for accurate diagnosis, especially when dealing with incomplete information. Aligning Large Language Models (LLMs) for these interactive scenarios is challenging due to the uncertainty inherent in user-agent interactions, which we formulate as a Hierarchical Markov Decision Process (H-MDP). While conventional Reinforcement Learning (RL) methods like Group Relative Policy Optimization (GRPO) struggle with long-horizon credit assignment and Proximal Policy Optimization (PPO) suffers from unstable value estimation in this context, we propose a novel uncertainty-aware Adaptive Tree Policy Optimization (ATPO) algorithm. Our method adaptively allocates the rollout budget to states with high uncertainty, quantified by a composite metric of Bellman error and action-value variance. This strategy enables more accurate value estimation, while fostering more efficient and diverse exploration. To mitigate the high computational cost of tree-based RL, we introduce two key optimizations: an uncertainty-guided pruning mechanism to minimize the number of rollouts, and an asynchronous search architecture that leverages KV cache reuse to maximize inference throughput. Extensive experiments on three public medical dialogue benchmarks demonstrate that our algorithm significantly outperforms several strong baselines, culminating in Qwen3-8B model surpassing the much larger GPT-4o ($+0.92\%$ accuracy).

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

arXiv:2509.15927v4 Announce Type: replace-cross Abstract: Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidding (AIGB), which learns a conditional generative planner from offline data, achieves superior performance compared to typical offline reinforcement learning (RL)-based auto-bidding methods. However, existing AIGB methods still face a performance bottleneck due to their inherent inability to explore beyond the static dataset with feedback. To address this, we propose \textbf{AIGB-Pearl} (\emph{\textbf{P}lanning with \textbf{E}valu\textbf{A}tor via \textbf{RL}}), a novel method that integrates generative planning and policy optimization. The core of AIGB-Pearl lies in constructing a trajectory evaluator to assess the quality of generated scores and designing a provably sound KL-Lipschitz-constrained score-maximization scheme to ensure safe and efficient exploration beyond the offline dataset. A practical algorithm that incorporates the synchronous coupling technique is further developed to ensure the model regularity required by the proposed scheme. Extensive experiments on both simulated and real-world advertising systems demonstrate the state-of-the-art performance of our approach.

RAIR: A Rule-Aware Benchmark Uniting Challenging Long-Tail and Visual Salience Subset for E-commerce Relevance Assessment

arXiv:2512.24943v2 Announce Type: replace-cross Abstract: Search relevance plays a central role in web e-commerce. While large language models (LLMs) have shown significant results on relevance task, existing benchmarks lack sufficient complexity for comprehensive model assessment, resulting in an absence of standardized relevance evaluation metrics across the industry. To address this limitation, we propose Rule-Aware benchmark with Image for Relevance assessment(RAIR), a Chinese dataset derived from real-world scenarios. RAIR established a standardized framework for relevance assessment and provides a set of universal rules, which forms the foundation for standardized evaluation. Additionally, RAIR analyzes essential capabilities required for current relevance models and introduces a comprehensive dataset consists of three subset: (1) a general subset with industry-balanced sampling to evaluate fundamental model competencies; (2) a long-tail hard subset focus on challenging cases to assess performance limits; (3) a visual salience subset for evaluating multimodal understanding capabilities. We conducted experiments on RAIR using 14 open and closed-source models. The results demonstrate that RAIR presents sufficient challenges even for GPT-5, which achieved the best performance. RAIR data are now available, serving as an industry benchmark for relevance assessment while providing new insights into general LLM and Visual Language Model(VLM) evaluation.

Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy

arXiv:2508.03346v2 Announce Type: replace Abstract: Large Language Models (LLMs) using Chain-of-Thought (CoT) prompting excel at complex reasoning but generate verbose thought processes with considerable redundancy, leading to increased inference costs and reduced efficiency. We introduce a novel CoT compression framework based on step entropy, a metric that quantifies \emph{the informational contribution of individual reasoning steps} to identify redundancy. Through theoretical analysis and extensive empirical validation on mathematical reasoning benchmarks, we demonstrate that steps with low entropy are indeed highly redundant. Our experiments reveal that an astonishing 80\% of low-entropy intermediate steps can be pruned with minor degradation in the final answer accuracy across DeepSeek-R1-7B, 14B and Qwen3-8B. This finding sharply contrasts with random or high-entropy pruning, which severely impairs reasoning performance. Building on this, we propose a novel two-stage training strategy combining Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) reinforcement learning. This approach enables LLMs to autonomously learn to generate compressed COTs during inference by strategically incorporating [SKIP] tokens. Our method significantly improves LLM inference efficiency while preserving accuracy, paving the way for more scalable LLM deployments and a better understanding of their internal reasoning. The code and data are released in https://github.com/staymylove/COT_Compresstion_via_Step_entropy.
❌