❌

Normal view

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

arXiv:2609.11318v2 Announce Type: replace Abstract: Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr. LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr. LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

Advances in understanding the mechanisms underlying acquired resistance to third-generation tyrosine kinase inhibitors in non-small cell lung cancer

Front Cell Dev Biol. 2026 Aug 24;14:1867246. doi: 10.3389/fcell.2026.1867246. eCollection 2026.

ABSTRACT

Acquired resistance to third-generation epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors (TKIs) presents a formidable challenge in the treatment of non-small cell lung cancer (NSCLC). Despite the remarkable efficacy of these agents, resistance inevitably develops, typically within approximately 10 months of treatment initiation. This review elucidates the multifaceted mechanisms driving this resistance, broadly categorized into on-target EGFR-dependent alterations and off-target EGFR-independent bypass pathway activations. On-target mechanisms include the emergence of tertiary EGFR mutations, most notably C797S, which disrupts TKI binding. Off-target mechanisms encompass the activation of alternative signaling pathways such as MET and HER2/HER3 amplification, as well as histological transformations and complex changes within the tumor microenvironment. Furthermore, recent discoveries highlight the role of epigenetic dysregulation and metabolic reprogramming in fostering resistance. To counter this pervasive adaptability, advanced diagnostic methodologies, including liquid biopsy and high-resolution omics technologies, are crucial for real-time molecular profiling. The field is actively exploring emerging combination therapeutic strategies to circumvent these diverse resistance pathways, aiming to prolong clinical benefits and improve patient outcomes. The persistent emergence of resistance underscores that current targeted therapies, while revolutionary, are primarily disease-modifying rather than curative, necessitating continuous innovation to overcome the inherent biological challenge of tumor adaptability and heterogeneity.

PMID:42707604 | PMC:PMC13547778 | DOI:10.3389/fcell.2026.1867246

Learn to Relax with Large Language Models: Solving Constraint Optimization Problems via Bidirectional Coevolution

arXiv:2509.12643v3 Announce Type: replace Abstract: Large Language Model (LLM)-based optimization has recently shown promise for autonomous problem solving, yet most approaches still cast LLMs as passive constraint checkers rather than proactive strategy designers, limiting their effectiveness on complex Constraint Optimization Problems (COPs). To address this, we present AutoCO, an end-to-end Automated Constraint Optimization method that tightly couples operations-research principles of constraint relaxation with LLM reasoning. A core innovation is a unified triple-representation that binds relaxation strategies, algorithmic principles, and executable codes. This design enables the LLM to synthesize, justify, and instantiate relaxation strategies that are both principled and executable. To navigate fragmented solution spaces, AutoCO employs a bidirectional global-local coevolution mechanism, synergistically coupling Monte Carlo Tree Search (MCTS) for global relaxation-trajectory exploration with Evolutionary Algorithms (EAs) for local solution intensification. This continuous exchange of priors and feedback explicitly balances diversification and intensification, thus preventing premature convergence. Extensive experiments on three challenging COP benchmarks validate AutoCO's consistent effectiveness and superior performance, especially in hard regimes where current methods degrade. Results highlight AutoCO as a principled and effective path toward proactive, verifiable LLM-driven optimization.

Integrated transcriptomic and proteomic analyses elucidate the stress tolerance network of <em>Saccharomyces boulardii</em> under gastrointestinal challenge

Food Funct. 2026 Mar 31. doi: 10.1039/d5fo04958j. Online ahead of print.

ABSTRACT

The probiotic yeast Saccharomyces boulardii is renowned for its clinical efficacy, which is intrinsically linked to its exceptional ability to survive the harsh gastrointestinal (GI) environment. However, a comprehensive understanding of the molecular mechanisms and regulatory pathways underlying the stress tolerance of S. boulardii remains limited. This study employed an integrated transcriptomic and proteomic approach to systematically map the dynamic responses of S. boulardii to simulated GI transit. Our analysis revealed that the intestinal phase posed a significantly greater challenge than the gastric phase, triggering extensive molecular reprogramming. A core adaptive strategy was the marked upregulation of the central carbon metabolism, particularly glycolysis, as evidenced by the concerted overexpression of key enzymes at both transcriptional and translational levels, indicating a heightened demand for energy to fuel stress defence mechanisms. Furthermore, significant enrichment was observed in the pathways related to nitrogen and fatty acid metabolism. Integration of the multi-omics datasets highlighted the complexity of the regulatory response, with frequent discordance between mRNA and protein abundance underscoring the importance of post-transcriptional regulation. This study provides a detailed molecular profile of the stress tolerance network in S. boulardii, elucidating the strategic metabolic rewiring and multi-layered regulation that underpin its probiotic resilience. The findings offer valuable insights and a foundational resource for the future development of enhanced probiotic therapies.

PMID:41914832 | DOI:10.1039/d5fo04958j

Integrated transcriptomic and proteomic analyses elucidate the stress tolerance network of <em>Saccharomyces boulardii</em> under gastrointestinal challenge

Food Funct. 2026 Mar 31. doi: 10.1039/d5fo04958j. Online ahead of print.

ABSTRACT

The probiotic yeast Saccharomyces boulardii is renowned for its clinical efficacy, which is intrinsically linked to its exceptional ability to survive the harsh gastrointestinal (GI) environment. However, a comprehensive understanding of the molecular mechanisms and regulatory pathways underlying the stress tolerance of S. boulardii remains limited. This study employed an integrated transcriptomic and proteomic approach to systematically map the dynamic responses of S. boulardii to simulated GI transit. Our analysis revealed that the intestinal phase posed a significantly greater challenge than the gastric phase, triggering extensive molecular reprogramming. A core adaptive strategy was the marked upregulation of the central carbon metabolism, particularly glycolysis, as evidenced by the concerted overexpression of key enzymes at both transcriptional and translational levels, indicating a heightened demand for energy to fuel stress defence mechanisms. Furthermore, significant enrichment was observed in the pathways related to nitrogen and fatty acid metabolism. Integration of the multi-omics datasets highlighted the complexity of the regulatory response, with frequent discordance between mRNA and protein abundance underscoring the importance of post-transcriptional regulation. This study provides a detailed molecular profile of the stress tolerance network in S. boulardii, elucidating the strategic metabolic rewiring and multi-layered regulation that underpin its probiotic resilience. The findings offer valuable insights and a foundational resource for the future development of enhanced probiotic therapies.

PMID:41914832 | DOI:10.1039/d5fo04958j

Scalable single-cell total RNA sequencing unifies coding and noncoding transcriptomics

Nature Biotechnology, Published online: 31 March 2026; doi:10.1038/s41587-026-03068-6

Simultaneous profiling of adenylated and non-adenylated RNAs reveals regulatory programs across diverse cell types.

Integrated Machine Learning and Multi-Omics Identifies a Novel Molecular Signature for Improving the Prognosis of Hepatocellular Carcinoma

J Hepatocell Carcinoma. 2026 Mar 11;13:574690. doi: 10.2147/JHC.S574690. eCollection 2026.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) exhibits significant molecular heterogeneity and complex immune microenvironment, which to some extent limits the accuracy of prognosis assessment and the formulation of individualized treatment strategies. This study aims to identify immune-derived molecular signatures based on multi-omics data and machine learning methods for the prognosis prediction and risk stratification of HCC.

METHODS: Based on weighted gene co-expression network analysis(WGCNA) and differential gene analysis,immune-derived molecular signature (IDMS) were screened in both single-cell and bulk transcriptomes. Prognostic model was constructed by multi-machine learning approachs. Subsequently, we investigated the differences in mutations, biological functions, and immune cell infiltration within the tumor microenvironment between the high- and low-risk groups.In addition, we comprehensively analyzed the drug sensitivity of IDMS and predicted potential drugs.

RESULTS: We identified seven hub genes at the single-cell and bulk transcriptome levels. Based on multiple machine learning, we constructed a prognostic model that demonstrated excellent performance in predicting overall survival for patients with HCC. IDMS -integrated normograms provide a promising and quantitative tool for clinical risk management.Notably, a significant difference in microsatellite instability (MSI) was observed between the high- and low-risk groups. This indicates that patients in the high-risk group might have a better response to immunotherapy. Additionally, we predicted potential drugs targeting to these risk subgroups.

CONCLUSION: Our research developed an IDMS that could serve as an effective tool for patient stratification management and prognosis prediction. This signature could provide a reference for immunotherapy for patients with HCC and improve their prognosis.

PMID:41847219 | PMC:PMC12991065 | DOI:10.2147/JHC.S574690

Human-specific features of the cerebellum and ZP2-regulated synapse development

Human-specific transcriptomic and regulatory features are present in the cerebellum, with ZP2 playing a key role in synapse regulation. ZP2 expression is induced by pontine mossy fibers, leading to decreased synaptic proteins and neuronal activity, which provides insights into the evolutionary development of the human cerebellum.

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

arXiv:2603.07648v1 Announce Type: cross Abstract: Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or skills. These challenges present significant barriers for existing VLA models, which use monolithic action decoders trained on aggregated data, resulting in poor scalability. To address these challenges, we propose AtomicVLA, a unified planning-and-execution framework that jointly generates task-level plans, atomic skill abstractions, and fine-grained actions. AtomicVLA constructs a scalable atomic skill library through a Skill-Guided Mixture-of-Experts (SG-MoE), where each expert specializes in mastering generic yet precise atomic skills. Furthermore, we introduce a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning. We validate our approach through extensive experiments. In simulation, AtomicVLA outperforms $\pi_{0}$ by 2.4\% on LIBERO, 10\% on LIBERO-LONG, and outperforms $\pi_{0}$ and $\pi_{0.5}$ by 0.22 and 0.25 in average task length on CALVIN. Additionally, our AtomicVLA consistently surpasses baselines by 18.3\% and 21\% in real-world long-horizon tasks and continual learning. These results highlight the effectiveness of atomic skill abstraction and dynamic expert composition for long-horizon and lifelong robotic tasks. The project page is \href{https://zhanglk9.github.io/atomicvla-web/}{here}.
❌