Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables
arXiv:2604.03660v1 Announce Type: new Abstract: Structured tables are essential for conveying high-density information in professional domains such as finance, healthcare, and scientific research. Despite the progress in Multimodal Large Language Models (MLLMs), reasoning performance remains limited for complex tables with hierarchical layouts. In this paper, we identify a critical Perception Bottleneck through quantitative analysis. We find that as task complexity scales, the number of involve
-
cs.AI, q-bio.NC updates on arXiv.org
-
Memory Intelligence Agent
arXiv:2604.04503v2 Announce Type: new Abstract: Deep research agents (DRAs) integrate LLM reasoning with external tools. Memory systems enable DRAs to leverage historical experiences, which are essential for efficient reasoning and autonomous evolution. Existing methods rely on retrieving similar trajectories from memory to aid reasoning, while suffering from key limitations of ineffective memory evolution and increasing storage and retrieval costs. To address these problems, we propose a novel
Memory Intelligence Agent
-
cs.AI, q-bio.NC updates on arXiv.org
-
Learning Additively Compositional Latent Actions for Embodied AI
arXiv:2604.03340v1 Announce Type: cross Abstract: Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn latent actions without structural priors that encode the additive, compositional structure of physical motion. As a result, latents often entangle irrelevant scene details or information about future observations with true state changes and miscalibrate motion magnitude. W
Learning Additively Compositional Latent Actions for Embodied AI
-
cs.AI, q-bio.NC updates on arXiv.org
-
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
arXiv:2604.04875v1 Announce Type: cross Abstract: Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio-visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross-level multimodal orchestration to achieve professional-grade fluidity, resulting in disjointed sequences with abrupt visual transi
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
-
cs.AI, q-bio.NC updates on arXiv.org
-
TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
arXiv:2601.22776v2 Announce Type: replace Abstract: Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (RL) frameworks for search-augmented reasoning predominantly rely on sparse outcome-level rewards, leading to a "Double Homogenization Dilemma." This manifests as (1) Process homogenization, where the thinking, reasoning, and tooling involved in generation are ignored.
TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
arXiv:2604.02368v3 Announce Type: replace Abstract: As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authen
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving
arXiv:2603.13842v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving is typically built upon imitation learning (IL), yet its performance is constrained by the quality of human demonstrations. To overcome this limitation, recent methods incorporate reinforcement learning (RL) through sequential fine-tuning. However, such a paradigm remains suboptimal: sequential RL fine-tuning can introduce policy drift and often leads to a performance ceiling due to its dependence on the pre
Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving
-
cs.AI, q-bio.NC updates on arXiv.org
-
ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
arXiv:2604.02006v1 Announce Type: new Abstract: Reinforcement Learning (RL) significantly enhances the reasoning abilities of large language models (LLMs), yet applying it to multi-turn agentic tasks remains challenging due to the long-horizon nature of interactions and the stochasticity of environmental feedback. We identify a structural failure mode in agentic exploration: suboptimal actions elicit noisy observations into misleading contexts, which further weaken subsequent decision-making, m
ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
arXiv:2511.18123v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have become indispensable for multimodal reasoning, yet their representations often encode and amplify demographic biases, resulting in biased associations and misaligned predictions in downstream tasks. Such behavior undermines fairness and distorts the intended alignment between vision and language. Recent post-hoc approaches attempt to mitigate bias by replacing the most attribute-correlated embedding coo
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
-
Journal of Medical Internet Research
-
Accuracy of Radiomics-Based Machine Learning for Predicting Risk of Recurrence in Non–Small Cell Lung Cancer: Systematic Review and Meta-Analysis
Background: During the diagnosis and treatment of non–small cell lung cancer (NSCLC), detecting the risk of its recurrence in an early phase is still challenging. Recent studies have investigated the radiomics-based machine learning (ML) models for detecting the risk of recurrence in NSCLC. However, there is still insufficient systematic evidence to prove its efficiency. Objective: This study is designed to systematically evaluate the effectiveness of radiomics-based ML in predicting the risk of
Accuracy of Radiomics-Based Machine Learning for Predicting Risk of Recurrence in Non–Small Cell Lung Cancer: Systematic Review and Meta-Analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
arXiv:2603.25158v3 Announce Type: replace Abstract: Equipping Large Language Model (LLM) agents with domain-specific skills is critical for tackling complex tasks. Yet, manual authoring creates a severe scalability bottleneck. Conversely, automated skill generation often yields fragile or fragmented results because it either relies on shallow parametric knowledge or sequentially overfits to non-generalizable trajectory-local lessons. To overcome this, we introduce Trace2Skill, a framework that
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
-
cs.AI, q-bio.NC updates on arXiv.org
-
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
arXiv:2512.08829v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly tasked with ultra-long multimodal understanding. While linear architectures offer constant computation and memory footprints, they often struggle with high-frequency visual perception compared to standard Transformers. To bridge this gap, we introduce \textbf{InfiniteVL}. We first develop a hybrid base model called \textbf{InfiniteVL-Base} that interleaves a small fraction of Full Attention
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
arXiv:2603.22386v1 Announce Type: new Abstract: Large language model (LLM)-based systems are becoming increasingly popular for solving tasks by constructing executable workflows that interleave LLM calls, information retrieval, tool use, code execution, memory updates, and verification. This survey reviews recent methods for designing and optimizing such workflows, which we treat as agentic computation graphs (ACGs). We organize the literature based on when workflow structure is determined, whe
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Symbolic Graph Networks for Robust PDE Discovery from Noisy Sparse Data
arXiv:2603.22380v1 Announce Type: cross Abstract: Data-driven discovery of partial differential equations (PDEs) offers a promising paradigm for uncovering governing physical laws from observational data. However, in practical scenarios, measurements are often contaminated by noise and limited by sparse sampling, which poses significant challenges to existing approaches based on numerical differentiation or integral formulations. In this work, we propose a Symbolic Graph Network (SGN) framework
Symbolic Graph Networks for Robust PDE Discovery from Noisy Sparse Data
-
cs.AI, q-bio.NC updates on arXiv.org
-
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
arXiv:2603.23269v1 Announce Type: cross Abstract: Large Language Models(LLMs) are widely deployed, yet are vulnerable to jailbreak prompts that elicit policy-violating outputs. Although prior studies have uncovered these risks, they typically treat all tokens as equally important during prompt mutation, overlooking the varying contributions of individual tokens to triggering model refusals. Consequently, these attacks introduce substantial redundant searching under query-constrained scenarios,
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
arXiv:2602.07023v2 Announce Type: replace-cross Abstract: Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. In
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
-
npj Digital Medicine
-
Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework
npj Digital Medicine, Published online: 24 March 2026; doi:10.1038/s41746-026-02560-2Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework
Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework
npj Digital Medicine, Published online: 24 March 2026; doi:10.1038/s41746-026-02560-2
Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Unannotated noncoding transcripts as a source of intratumor heterogeneity in malignant cell states
Sci China Life Sci. 2026 Mar 16. doi: 10.1007/s11427-025-3273-6. Online ahead of print.ABSTRACTPhenotypic diversity of malignant cells within a tumor underlies intratumor heterogeneity (ITH), a key determinant of cancer metastasis and treatment failure. However, the molecular mechanisms driving this heterogeneity are poorly understood. Here, we curated and analyzed a cohort of 3' tag-based single-cell RNA-seq covering 12 common cancer types. We identified thousands of poly(A) site (PAS) peaks re
Unannotated noncoding transcripts as a source of intratumor heterogeneity in malignant cell states
Sci China Life Sci. 2026 Mar 16. doi: 10.1007/s11427-025-3273-6. Online ahead of print.
ABSTRACT
Phenotypic diversity of malignant cells within a tumor underlies intratumor heterogeneity (ITH), a key determinant of cancer metastasis and treatment failure. However, the molecular mechanisms driving this heterogeneity are poorly understood. Here, we curated and analyzed a cohort of 3' tag-based single-cell RNA-seq covering 12 common cancer types. We identified thousands of poly(A) site (PAS) peaks representing the 3' ends of previously unannotated transcripts, whose expression is widely associated with diverse malignant cellular states. By integrating multi-omics data, we characterized the expression patterns and epigenetic landscape of these unannotated PAS peak-associated transcripts (UPTs). The expression heterogeneity of UPTs was supported by multi-region sampling bulk RNA-seq data and recapitulated within cancer cell lines. As proof of principle validation, functional experiments confirmed that two noncoding UPTs promoted the proliferation and migration of lung cancer cells. Our results suggest that epigenetic activation of unannotated noncoding transcripts might represent a previously unrecognized mechanism contributing to transcriptomic ITH.
PMID:41870780 | DOI:10.1007/s11427-025-3273-6
-
Omics In Lung
-
Unannotated noncoding transcripts as a source of intratumor heterogeneity in malignant cell states
Sci China Life Sci. 2026 Mar 16. doi: 10.1007/s11427-025-3273-6. Online ahead of print.ABSTRACTPhenotypic diversity of malignant cells within a tumor underlies intratumor heterogeneity (ITH), a key determinant of cancer metastasis and treatment failure. However, the molecular mechanisms driving this heterogeneity are poorly understood. Here, we curated and analyzed a cohort of 3' tag-based single-cell RNA-seq covering 12 common cancer types. We identified thousands of poly(A) site (PAS) peaks re
Unannotated noncoding transcripts as a source of intratumor heterogeneity in malignant cell states
Sci China Life Sci. 2026 Mar 16. doi: 10.1007/s11427-025-3273-6. Online ahead of print.
ABSTRACT
Phenotypic diversity of malignant cells within a tumor underlies intratumor heterogeneity (ITH), a key determinant of cancer metastasis and treatment failure. However, the molecular mechanisms driving this heterogeneity are poorly understood. Here, we curated and analyzed a cohort of 3' tag-based single-cell RNA-seq covering 12 common cancer types. We identified thousands of poly(A) site (PAS) peaks representing the 3' ends of previously unannotated transcripts, whose expression is widely associated with diverse malignant cellular states. By integrating multi-omics data, we characterized the expression patterns and epigenetic landscape of these unannotated PAS peak-associated transcripts (UPTs). The expression heterogeneity of UPTs was supported by multi-region sampling bulk RNA-seq data and recapitulated within cancer cell lines. As proof of principle validation, functional experiments confirmed that two noncoding UPTs promoted the proliferation and migration of lung cancer cells. Our results suggest that epigenetic activation of unannotated noncoding transcripts might represent a previously unrecognized mechanism contributing to transcriptomic ITH.
PMID:41870780 | DOI:10.1007/s11427-025-3273-6
-
cs.AI, q-bio.NC updates on arXiv.org
-
LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction
arXiv:2603.12647v1 Announce Type: cross Abstract: Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth supervision, while the rich scene information contained in point clouds, such as reflectance, and the complementarity between LiDAR and RGB have not been fully exploited, leading to degradatio