Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
PANDO: Efficient Multimodal AI Agents via Online Skill Distillation
arXiv:2605.24785v2 Announce Type: new Abstract: Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raises a central question: can a web agent become more efficient as it accumulates experience, rather than more expensive? We first analyze trajectories from VisualWebArena and identify three recurring sources of inefficiency: repeat-action loops, hidden d
-
cs.AI, q-bio.NC updates on arXiv.org
-
L2IR: Revealing Latent Intent in Graph Fraud Detection
arXiv:2605.26040v1 Announce Type: new Abstract: Graph fraud detection has long depended on Graph Neural Networks (GNNs) to propagate and aggregate information across relational data. A critical obstacle in practice, however, is that fraudsters frequently disguise themselves by forging numerous connections with benign users, causing fraud signals to be progressively diluted during neighborhood aggregation and undermining detection reliability. While recent efforts have used Large Language Models
L2IR: Revealing Latent Intent in Graph Fraud Detection
-
cs.AI, q-bio.NC updates on arXiv.org
-
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
arXiv:2605.24675v1 Announce Type: cross Abstract: Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce domains. Although Large Vision-Language Models (LVLMs) have advanced multimodal understanding, applying them to Web image translation remains challenging due to the visual representation gap: standard encoders often prioritize high-level semantics over the fine-grained vis
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
arXiv:2605.24681v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine-tuning LLMs with parallel corpora presents major challenges, namely parameter interference. To address these issues, we propose Mix-MoE, a mixed Mixture-of-Experts framework designed to train LLMs for multilingual MT. Our framework operates in two distinct stages: (1) post-pretraining with MoE on
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
-
cs.AI, q-bio.NC updates on arXiv.org
-
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
arXiv:2605.24786v1 Announce Type: cross Abstract: Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention, leaving unused a signal computed on every decoding step: the model's current uncertainty. We introduce CONF-KV, a KV-cache manager that converts the next-token distribution into a scalar confidence score and uses it
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
-
cs.AI, q-bio.NC updates on arXiv.org
-
Agent Learning via Early Experience
arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents from experience data with reinforcement learning remains difficult in many environments, which either lack verifiable rewards (e.g., websites) or require inefficient long-horizon rollouts (e.g., multi-turn tool use). As a result, most current agents rely on supervised f
Agent Learning via Early Experience
-
cs.AI, q-bio.NC updates on arXiv.org
-
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
arXiv:2605.02900v2 Announce Type: replace-cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act un
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
-
cs.AI, q-bio.NC updates on arXiv.org
-
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
arXiv:2605.20251v4 Announce Type: replace-cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluate
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
-
Nature Biotechnology - Issue - nature.com science feeds
-
A framework for building a synthetic cell from the SynCell Asia Initiative
Nature Biotechnology, Published online: 26 May 2026; doi:10.1038/s41587-026-03153-wBuilding a living cell from scratch requires overcoming a bottleneck that has remained unresolved despite decades of progress: orchestrating the spatiotemporal integration of core functional modules. To tackle this barrier, the SynCell Asia Initiative outlines a strategy for developing core functional modules followed by their systems-level integration through the establishment of a centralized, artificial intelli
A framework for building a synthetic cell from the SynCell Asia Initiative
Nature Biotechnology, Published online: 26 May 2026; doi:10.1038/s41587-026-03153-w
Building a living cell from scratch requires overcoming a bottleneck that has remained unresolved despite decades of progress: orchestrating the spatiotemporal integration of core functional modules. To tackle this barrier, the SynCell Asia Initiative outlines a strategy for developing core functional modules followed by their systems-level integration through the establishment of a centralized, artificial intelligence (AI)-driven biofoundry.-
Omics in Gastric
-
Integrative multi-omics analysis identifies stromal-immune crosstalk as a determinant of immunotherapy efficacy and establishes a prognostic signature in gastric cancer
Comput Biol Chem. 2026 Apr 23;124(Pt 1):109095. doi: 10.1016/j.compbiolchem.2026.109095. Online ahead of print.ABSTRACTImmune checkpoint inhibitors like pembrolizumab exhibit variable efficacy in metastatic gastric cancer (GC). This study aimed to identify molecular drivers of pembrolizumab response, explore mechanisms of immune checkpoint inhibitors (ICIs) efficacy, and develop a prognostic signature. Transcriptomic analysis of pembrolizumab-treated GC (TIGER database) identified 165 response-a
Integrative multi-omics analysis identifies stromal-immune crosstalk as a determinant of immunotherapy efficacy and establishes a prognostic signature in gastric cancer
Comput Biol Chem. 2026 Apr 23;124(Pt 1):109095. doi: 10.1016/j.compbiolchem.2026.109095. Online ahead of print.
ABSTRACT
Immune checkpoint inhibitors like pembrolizumab exhibit variable efficacy in metastatic gastric cancer (GC). This study aimed to identify molecular drivers of pembrolizumab response, explore mechanisms of immune checkpoint inhibitors (ICIs) efficacy, and develop a prognostic signature. Transcriptomic analysis of pembrolizumab-treated GC (TIGER database) identified 165 response-associated differentially expressed genes (DEGs). Functional annotation and single-cell RNA sequencing (scRNA-seq) data from the Gene Expression Omnibus (GEO) revealed that responder-upregulated genes (R-DEGs) were enriched in immune activation pathways and mainly localized to CD8 + T/NK cells. In contrast, non-responder-upregulated genes (D-DEGs) were linked to extracellular matrix (ECM) remodeling and mainly expressed in fibroblasts/endothelial cells. CellChat analysis demonstrated that key DEGs mediate immune-stromal crosstalk via MHC-I and collagen/laminin signaling. A prognostic signature (Lasso-StepCox[forward] Riskscore; LSR: APOD, APOH, BATF2, GJA1, MAGED1, SLC5A1, SLCO2A1, VWF, VCAN) was derived and validated in four independent GC cohorts from the GEO and Cancer Genome Atlas (TCGA) database. Multi-omics analyses showed that LSR-high tumors exhibited aggressive clinicopathological features, increased stromal components, reduced cytotoxic immune infiltration, diminished tumor mutational burden (TMB), and poorer prognosis. Immunohistochemistry (IHC) and spatial transcriptomics in GC showed that stromal VWF/VCAN expression correlates with reduced CD8⁺ T cell granzyme B expression, suggesting T cell dysfunction. High VWF expression in GC predicted poor survival, and a combined VWF/VCAN score showed enhanced prognostic stratification. This study highlights stromal-immune crosstalk as a driver of pembrolizumab resistance and provides a signature as a clinical tool for prognosis and personalized therapy in metastatic GC.
PMID:42068630 | DOI:10.1016/j.compbiolchem.2026.109095
-
Nature - Issue - nature.com science feeds
-
Autonomous closed-loop framework for reproducible perovskite solar cells
Nature, Published online: 14 April 2026; doi:10.1038/s41586-026-10482-yAutonomous closed-loop framework for reproducible perovskite solar cells
Autonomous closed-loop framework for reproducible perovskite solar cells
Nature, Published online: 14 April 2026; doi:10.1038/s41586-026-10482-y
Autonomous closed-loop framework for reproducible perovskite solar cells-
Nature - Issue - nature.com science feeds
-
Mummified early Permian reptile reveals ancient amniote breathing apparatus
Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10307-yA mummified fossil of the early Permian reptile Captorhinus reveals the potential ancestral amniote breathing mechanism and its impact on terrestrial vertebrate evolution.
Mummified early Permian reptile reveals ancient amniote breathing apparatus
Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10307-y
A mummified fossil of the early Permian reptile Captorhinus reveals the potential ancestral amniote breathing mechanism and its impact on terrestrial vertebrate evolution.-
cs.AI, q-bio.NC updates on arXiv.org
-
ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
arXiv:2604.04426v1 Announce Type: new Abstract: Existing research on LLM agent security mainly focuses on prompt injection and unsafe input/output behaviors. However, as agents increasingly rely on third-party tools and MCP servers, a new class of supply-chain threats has emerged, where malicious behaviors are embedded in seemingly benign tools, silently hijacking agent execution, leaking sensitive data, or triggering unauthorized actions. Despite their growing impact, there is currently no com
ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
arXiv:2604.04901v1 Announce Type: cross Abstract: Coworking AI agents operating within local file systems are rapidly emerging as a paradigm in human-AI interaction; however, effective personalization remains limited by severe data constraints, as strict privacy barriers and the difficulty of jointly collecting multimodal real-world traces prevent scalable training and evaluation, and existing methods remain interaction-centric while overlooking dense behavioral traces in file-system operations
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
-
cs.AI, q-bio.NC updates on arXiv.org
-
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
arXiv:2505.21605v3 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit advancing capabilities in complex tasks, such as reasoning and graduate-level question answering, yet their resilience against misuse, particularly involving scientifically sophisticated risks, remains underexplored. Existing safety benchmarks typically focus either on instructions requiring minimal knowledge comprehension (e.g., ``tell me how to build a bomb") or utilize prompts that are relatively l
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
-
cs.AI, q-bio.NC updates on arXiv.org
-
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
arXiv:2602.01554v2 Announce Type: replace-cross Abstract: Unified multimodal large language models (MLLMs) aim to unify image understanding and image generation within a single framework, where a shared visual tokenizer serves as the sole interface that maps high-dimensional images into a limited token budget for downstream multimodal reasoning and synthesis. However, existing shared-token designs are largely architecture-driven and lack an explicit criterion for what information should be pres
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
arXiv:2601.04823v5 Announce Type: replace Abstract: Mixture-of-Experts (MoE) has become a prominent paradigm for scaling Large Language Models (LLMs). Parameter-efficient fine-tuning methods, such as LoRA, are widely adopted to adapt pretrained MoE LLMs to downstream tasks. However, existing approaches typically assign identical LoRA ranks to all expert modules, ignoring the heterogeneous specialization of pretrained experts. This uniform allocation leads to a resource mismatch: task-relevant e
DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
arXiv:2604.00830v2 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is an adaptation policy that updates the actor policy based on experience from previous episodes, thereby improving future behavior. Existing methods rely on fixed, hand-crafted adaptation policies rather than optimizing them for downstream improvement. We argue
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
-
cs.AI, q-bio.NC updates on arXiv.org
-
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
arXiv:2603.29211v1 Announce Type: new Abstract: In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of how general multimodal models can be developed into
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
arXiv:2603.29025v1 Announce Type: cross Abstract: Large language models systematically fail when a salient surface cue conflicts with an unstated feasibility constraint. We study this through a diagnose-measure-bridge-treat framework. Causal-behavioral analysis of the ``car wash problem'' across six models reveals approximately context-independent sigmoid heuristics: the distance cue exerts 8.7 to 38 times more influence than the goal, and token-level attribution shows patterns more consistent