Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
Toward Executable Repository-Level Code Generation via Environment Alignment
arXiv:2604.03622v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on code generation, but existing methods still struggle with repository-level code generation under executable validation. Under this evaluation setting, success is determined not by the plausibility of isolated code fragments, but by whether a generated multi-file repository can be successfully installed, have its dependencies and internal references resolved, be launched, and be val
-
cs.AI, q-bio.NC updates on arXiv.org
-
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
arXiv:2604.03632v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved substantial progress in repository-level code generation. However, solving the same repository-level task often requires multiple attempts, while existing methods still optimize each attempt in isolation and do not preserve or reuse task-specific state across attempts. In this paper, we propose LiveCoder, a novel framework for repository-level code generation based on cross-attempt knowledge optimizatio
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
LightThinker++: From Reasoning Compression to Memory Management
arXiv:2604.03679v1 Announce Type: cross Abstract: Large language models (LLMs) excel at complex reasoning, yet their efficiency is limited by the surging cognitive overhead of long thought traces. In this paper, we propose LightThinker, a method that enables LLMs to dynamically compress intermediate thoughts into compact semantic representations. However, static compression often struggles with complex reasoning where the irreversible loss of intermediate details can lead to logical bottlenecks
LightThinker++: From Reasoning Compression to Memory Management
-
cs.AI, q-bio.NC updates on arXiv.org
-
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
arXiv:2604.02368v3 Announce Type: replace Abstract: As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authen
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
-
cs.AI, q-bio.NC updates on arXiv.org
-
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
arXiv:2604.01621v1 Announce Type: cross Abstract: Large language model (LLM) inference increasingly depends on multi-GPU execution, yet existing inference parallelization strategies require layer-wise inter-rank synchronization, making end-to-end performance sensitive to workload imbalance. We present DWDP (Distributed Weight Data Parallelism), an inference parallelization strategy that preserves data-parallel execution while offloading MoE weights across peer GPUs and fetching missing experts
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
-
cs.AI, q-bio.NC updates on arXiv.org
-
Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion
arXiv:2604.01669v1 Announce Type: cross Abstract: Embodied perception systems face severe challenges of dynamic environment distribution drift when they continuously interact in open physical spaces. However, the existing domain incremental awareness methods often rely on the domain id obtained in advance during the testing phase, which limits their practicability in unknown interaction scenarios. At the same time, the model often overfits to the context-specific perceptual noise, which leads t
Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion
-
cs.AI, q-bio.NC updates on arXiv.org
-
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
arXiv:2601.10611v4 Announce Type: replace-cross Abstract: Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from them, or do not disclose their training data or recipe. As a result, the open-source community lacks the foundations needed to improve on the state-of-the-art video (and image) language models. Crucially, many downstream applications require more than just hig
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
-
cs.AI, q-bio.NC updates on arXiv.org
-
D-SPEAR: Dual-Stream Prioritized Experience Adaptive Replay for Stable Reinforcement Learning in Robotic Manipulation
arXiv:2603.27346v2 Announce Type: replace-cross Abstract: Robotic manipulation remains challenging for reinforcement learning due to contact-rich dynamics, long horizons, and training instability. Although off-policy actor-critic algorithms such as SAC and TD3 perform well in simulation, they often suffer from policy oscillations and performance collapse in realistic settings, partly due to experience replay strategies that ignore the differing data requirements of the actor and the critic. We
D-SPEAR: Dual-Stream Prioritized Experience Adaptive Replay for Stable Reinforcement Learning in Robotic Manipulation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Owl-AuraID 1.0: An Intelligent System for Autonomous Scientific Instrumentation and Scientific Data Analysis
arXiv:2603.29828v1 Announce Type: new Abstract: Scientific discovery increasingly depends on high-throughput characterization, yet automation is hindered by proprietary GUIs and the limited generalizability of existing API-based systems. We present Owl-AuraID, a software-hardware collaborative embodied agent system that adopts a GUI-native paradigm to operate instruments through the same interfaces as human experts. Its skill-centric framework integrates Type-1 (GUI operation) and Type-2 (data
Owl-AuraID 1.0: An Intelligent System for Autonomous Scientific Instrumentation and Scientific Data Analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
MindCube: Spatial Mental Modeling from Limited Views
arXiv:2506.21458v2 Announce Type: replace Abstract: Can Vision-Language Models (VLMs) imagine the full scene from just a few views, like humans do? Humans form spatial mental models naturally, internal representations of unseen space, to reason about layout, perspective, and motion. Our MindCube benchmark with 21,154 questions across 3,268 images exposes this critical gap, where existing VLMs exhibit near-random performance. Using MindCube, we systematically evaluate how well VLMs build robust
MindCube: Spatial Mental Modeling from Limited Views
-
cs.AI, q-bio.NC updates on arXiv.org
-
Incorporating LLM Embeddings for Variation Across the Human Genome
arXiv:2509.20702v2 Announce Type: replace-cross Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated annotations from FAVOR, ClinVar, and the GWAS Catalog, we construct functional text descriptions for 8.9 billion possible
Incorporating LLM Embeddings for Variation Across the Human Genome
-
cs.AI, q-bio.NC updates on arXiv.org
-
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
arXiv:2603.19979v2 Announce Type: replace-cross Abstract: Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realis
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
-
cs.AI, q-bio.NC updates on arXiv.org
-
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
arXiv:2603.23481v1 Announce Type: cross Abstract: Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long-horizon tasks through visual reasoning, they remain limited in contact-rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Metaphor-based Jailbreak Attacks on Text-to-Image Models
arXiv:2512.10766v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these mechanisms and induce T2I models to produce sensitive content, revealing critical safety vulnerabilities. However, existing attack methods implicitly assume that the attacker knows the type of deployed defenses, which limi
Metaphor-based Jailbreak Attacks on Text-to-Image Models
-
Omics in Gastric
-
Mechanisms of Xinwei Tang in stress-induced gastric dysmotility: evidence from rat and In Vitro models
In Vitro Cell Dev Biol Anim. 2026 Mar 18. doi: 10.1007/s11626-026-01151-5. Online ahead of print.ABSTRACTStress is a key trigger of gastric dysmotility, partly via mitochondrial dysfunction and disordered gut-brain hormonal signaling. Xinwei Tang (XWT) is a multi-herb formula used empirically for upper gastrointestinal symptoms, but its mechanisms remain unclear. This study aimed to determine whether XWT alleviates water-immersion restraint stress (WIRS)-induced gastric dysmotility and to deline
Mechanisms of Xinwei Tang in stress-induced gastric dysmotility: evidence from rat and In Vitro models
In Vitro Cell Dev Biol Anim. 2026 Mar 18. doi: 10.1007/s11626-026-01151-5. Online ahead of print.
ABSTRACT
Stress is a key trigger of gastric dysmotility, partly via mitochondrial dysfunction and disordered gut-brain hormonal signaling. Xinwei Tang (XWT) is a multi-herb formula used empirically for upper gastrointestinal symptoms, but its mechanisms remain unclear. This study aimed to determine whether XWT alleviates water-immersion restraint stress (WIRS)-induced gastric dysmotility and to delineate underlying mitochondrial and metabolic pathways using integrated in vivo, in vitro and multi-omics approaches. Male rats underwent 7-d WIRS and received vehicle, domperidone (3 mg/kg) or XWT (3, 6, 12 g/kg). Gastric emptying, serum motilin/gastrin, oxidative stress indices and PINK1/Parkin-LC3/p62 proteins were assessed, and H₂O₂-injured GES-1 cells were treated with XWT-medicated serum. Gastric antra from MOD and XWT-H rats were analyzed by RNA-seq and DIA proteomics (n = 3/group). WIRS reduced gastric emptying by roughly half and lowered motilin/gastrin, increased ROS/MDA and disrupted PINK1/Parkin-LC3/p62 profiles; XWT dose-dependently reversed these changes, with XWT-H approximating domperidone. Omics revealed XWT-associated downregulation of inflammatory/protease and acute-phase genes/proteins and enrichment of oxidative phosphorylation, tricarboxylic-acid cycle and other metabolic pathways, without global activation of canonical autophagy/mitophagy gene sets. These preclinical data indicate that XWT ameliorates stress-induced gastric dysmotility via mitochondria- and metabolism-centred protection with selective tuning of mitophagy-related proteins.
PMID:41851413 | DOI:10.1007/s11626-026-01151-5
-
cs.AI, q-bio.NC updates on arXiv.org
-
Scaling Generalist Data-Analytic Agents
arXiv:2509.25084v3 Announce Type: replace-cross Abstract: Data-analytic agents are emerging as a key catalyst for automated scientific discovery and for the vision of Innovating AI. Current approaches, however, rely heavily on prompt engineering over proprietary models, while open-source models struggle to face diverse-format, large-scale data files and long-horizon, multi-step reasoning that real-world analytics demands. This paper introduces DataMind, a scalable data synthesis and agent train
Scaling Generalist Data-Analytic Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
arXiv:2510.23571v2 Announce Type: replace-cross Abstract: The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to reproduce. As policies expand in scope and complexity, these barriers only intensify, since defining "success" in robotics often hinges on nuanced human j
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Overcoming the Curvature Bottleneck in MeanFlow
arXiv:2511.23342v2 Announce Type: replace-cross Abstract: MeanFlow offers a promising framework for one-step generative modeling by directly learning a mean-velocity field, bypassing expensive numerical integration. However, we find that the highly curved generative trajectories of existing models induce a noisy loss landscape, severely bottlenecking convergence and model quality. We leverage a fundamental geometric principle to overcome this: mean-velocity estimation is drastically simpler alo
Overcoming the Curvature Bottleneck in MeanFlow
-
Nature - Issue - nature.com science feeds
-
The real story behind China’s technology triumph
Nature, Published online: 16 March 2026; doi:10.1038/d41586-026-00841-0An exemplary book gets much right, but misses the point that the nation’s tech dominance has more to do with globalization than with policy choices.
The real story behind China’s technology triumph
Nature, Published online: 16 March 2026; doi:10.1038/d41586-026-00841-0
An exemplary book gets much right, but misses the point that the nation’s tech dominance has more to do with globalization than with policy choices.-
Oncogene - Issue - nature.com science feeds
-
LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma
Oncogene, Published online: 14 March 2026; doi:10.1038/s41388-026-03714-yLINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma
LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma
Oncogene, Published online: 14 March 2026; doi:10.1038/s41388-026-03714-y
LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma