Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables
arXiv:2604.03660v1 Announce Type: new Abstract: Structured tables are essential for conveying high-density information in professional domains such as finance, healthcare, and scientific research. Despite the progress in Multimodal Large Language Models (MLLMs), reasoning performance remains limited for complex tables with hierarchical layouts. In this paper, we identify a critical Perception Bottleneck through quantitative analysis. We find that as task complexity scales, the number of involve
-
cs.AI, q-bio.NC updates on arXiv.org
-
Safe Decentralized Operation of EV Virtual Power Plant with Limited Network Visibility via Multi-Agent Reinforcement Learning
arXiv:2604.03278v1 Announce Type: cross Abstract: As power systems advance toward net-zero targets, behind-the-meter renewables are driving rapid growth in distributed energy resources (DERs). Virtual power plants (VPPs) increasingly coordinate these resources to support power distribution network (PDN) operation, with EV charging stations (EVCSs) emerging as a key asset due to their strong impact on local voltages. However, in practice, VPPs must make operational decisions with only partial vi
Safe Decentralized Operation of EV Virtual Power Plant with Limited Network Visibility via Multi-Agent Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
arXiv:2603.29318v1 Announce Type: new Abstract: Smartphone GUI agents execute tasks by operating directly on app interfaces, offering a path to broad capability without deep system integration. However, real-world smartphone use is highly personalized: users adopt diverse workflows and preferences, challenging agents to deliver customized assistance rather than generic solutions. Existing GUI agent benchmarks cannot adequately capture this personalization dimension due to sparse user-specific d
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
-
Nature - Issue - nature.com science feeds
-
Expansion of outer cortical CUX2 neurons requires adaptations for DNA repair
Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10290-4The transcription factor ATF4 is shown to regulate double-stranded DNA repair within vulnerable CUX2+ upper-layer 2/3 cortical neurons, enabling their survival during development.
Expansion of outer cortical CUX2 neurons requires adaptations for DNA repair
Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10290-4
The transcription factor ATF4 is shown to regulate double-stranded DNA repair within vulnerable CUX2+ upper-layer 2/3 cortical neurons, enabling their survival during development.-
cs.AI, q-bio.NC updates on arXiv.org
-
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
arXiv:2510.14967v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based agents are increasingly trained with reinforcement learning (RL) to enhance their ability to interact with external environments through tool use, particularly in search-based settings that require multi-turn reasoning and knowledge acquisition. However, existing approaches typically rely on outcome-based rewards that are only provided exclusively upon generating the final answer. This reward sparsity bec
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
arXiv:2603.19643v2 Announce Type: replace-cross Abstract: Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization to complex scenes, complicated pipeline, and efficient inference. To tackle these problems, we propose OmniDiT, an omni Virtual Try-On framework based on the Diffusion Transformer, which combines try-on and try-off tasks into one unified model. Specifically, w
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
-
Omics In Lung
-
Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer
Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.ABSTRACTPURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell
Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer
Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.
ABSTRACT
PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.
EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.
RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3α and reduced circulating α4β7+ regulatory T cells.
CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.
PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153
-
Cell
-
Tuning the sensitivity of mechanosensory receptors through histidine scanning
Histidine scanning represents a broadly applicable technique for the identification of critical interaction sites within TCRs and other mechanosensory receptors to enhance receptor signaling strength and augment therapeutic efficacy via the catch bond mechanism.
Tuning the sensitivity of mechanosensory receptors through histidine scanning
-
Journal of Medical Internet Research
-
Effect of a Digital-Driven Physician-Pharmacist Collaborative Model for Diabetes in Primary Health Care: Cluster Randomized Trial
Background: Evidence-based physician-pharmacist collaborative clinics have demonstrated significant short-term benefits for patients with type 2 diabetes (T2D), but their long-term effectiveness remains unclear, especially in primary health care settings. Objective: This study aimed to explore the long-term effectiveness and cost-effectiveness of a novel, digital-driven, multifaceted physician-pharmacist collaborative model for managing patients with T2D in underresourced settings. Methods: We c
Effect of a Digital-Driven Physician-Pharmacist Collaborative Model for Diabetes in Primary Health Care: Cluster Randomized Trial
-
cs.AI, q-bio.NC updates on arXiv.org
-
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
arXiv:2603.08090v1 Announce Type: cross Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit critical limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across different subject d
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Real-Time Aligned Reward Model beyond Semantics
arXiv:2601.22664v3 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for aligning large language models (LLMs) with human preferences, yet it is susceptible to reward overoptimization, in which policy models overfit to the reward model, exploit spurious reward patterns instead of faithfully capturing human intent. Prior mitigations primarily relies on surface semantic information and fails to efficiently address the misalignment between th
Real-Time Aligned Reward Model beyond Semantics
-
cs.AI, q-bio.NC updates on arXiv.org
-
To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models
arXiv:2602.12566v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level performance in some specific domains via RLVR, such as coding or math. When a general multi-domain expert-level model is required, we need to carefully consider the collaboration of RLVR across different domains. The current state-of-the-art models mainly employ two
To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
arXiv:2505.07365v2 Announce Type: replace-cross Abstract: We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over diverse acoustic scenes. We describe the dataset composition (from marine mammal calls to soundscapes and complex real-world clips), the evalua
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
arXiv:2506.17252v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly dependent on the quality of the underlying human preference data. To address this bottleneck, prior work has explored various data selection strategies, but these methods often overlook the impact of the evolving states of the language model during the optimization
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards Personalized Deep Research: Benchmarks and Evaluations
arXiv:2509.25106v3 Announce Type: replace-cross Abstract: Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended deep research benchmarks remain scarce and typically neglect personalized scenarios. To bridge this gap, we introduce Personalized Deep Research Bench (PDR-Bench), the first benchmark for evaluating p
Towards Personalized Deep Research: Benchmarks and Evaluations
-
cs.AI, q-bio.NC updates on arXiv.org
-
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
arXiv:2602.08354v2 Announce Type: replace Abstract: Recent advancements in large reasoning models (LRMs) have greatly improved their capabilities on complex reasoning tasks through Long Chains of Thought (CoTs). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. Recent studies show that longer reasoning chains are frequently uncorrelated with correctness and can even be detrimental to accur