❌

Reading view

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.
  •  

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal composition. We introduce EvoRS, a self-evolving RL framework that evolves the reward system from on-policy experience, representing it as an executable Reward-DAG. Specifically, an agentic designer updates this system from on-policy rollouts and reward traces to maintain train-time reliability. Across writing and roleplay, EvoRS achieves the best quality under all three judges, outperforming the policy by \(2.107\) and \(4.767\) points, respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations confirm that a comprehensive fixed reward system cannot remain reliable in open-ended tasks and must evolve throughout training.
  •  

PRXL2B facilitates the progression of hepatocellular carcinoma and the therapeutic efficacy of oncolytic adenovirus H101 through the PI3K/AKT/PD-L1 axis

Biosci Trends. 2026 May 21. doi: 10.5582/bst.2026.01000. Online ahead of print.

ABSTRACT

Oncolytic adenovirus H101 has shown antitumor activity in hepatocellular carcinoma (HCC), but the molecular determinants of treatment response remain unclear. In this study, a Hepa1-6 subcutaneous tumor model was established in C57BL/6 mice and treated with intratumoral H101, followed by integrated transcriptomic and proteomic analyses to identify candidate genes associated with H101 response. PRXL2B was selected for further investigation using public multi-omics datasets, tissue microarray-based immunohistochemistry, in vitro functional assays, mechanistic analyses, and in vivo validation experiments. Integrated multi-omics analyses identified PRXL2B as a candidate gene downregulated after H101 treatment. Public datasets and tissue-based validation further showed that PRXL2B was upregulated in HCC tissues. In MHCC97H and HCCLM3 cells, PRXL2B knockdown inhibited proliferation, migration, and invasion, promoted apoptosis and cell-cycle arrest, and enhanced the antitumor effect of H101. Mechanistically, PRXL2B silencing reduced AKT phosphorylation and PD-L1 expression. In vivo, PRXL2B knockdown suppressed tumor growth, and the combination of PRXL2B knockdown and H101 produced the strongest antitumor effect. These findings indicate that PRXL2B promotes malignant phenotypes in HCC and may modulate H101 efficacy through the PI3K/AKT/PD-L1 axis. Targeting PRXL2B may therefore represent a potential strategy to enhance the therapeutic efficacy of oncolytic virus therapy in HCC.

PMID:42161529 | DOI:10.5582/bst.2026.01000

  •  

Characterization of dysbiosis patterns in gut microbiota of digestive system cancers: an umbrella review

Front Microbiol. 2026 Apr 28;17:1782471. doi: 10.3389/fmicb.2026.1782471. eCollection 2026.

ABSTRACT

Digestive system cancers (DSCs) represent a substantial global health burden. In recent years, the role of gut microbiota in the DSCs has garnered considerable attention, but its change pattern during tumor progression and the specific mechanisms are still not fully understood. We conducted a comprehensive systematic review to characterize patterns of gut microbiota dysbiosis across different DSC types and assess their clinical significance. We systematically searched four English and three Chinese databases up to January 2025 to identify systematic reviews focused on the dynamic characteristics of the gut microbiota during gastrointestinal tumorigenesis. Microbiota biodiversity and taxonomic composition were extracted to identify specific signatures associated with DSCs. The ROBIS tool was used to evaluate the methodological quality of the included studies. Ultimately, 59 studies involving six distinct DSC types were included. Data synthesis and comparison revealed distinct microbiota profiles across DSCs. At the phylum level, Bacillota was decreased in esophageal cancer (EC) and pancreatic ductal adenocarcinoma (PDAC), Pseudomonadota was augmented in EC but exhibited divergent trajectories in colorectal cancer (CRC) and PDAC. Genus-level analyses revealed Veillonella enrichment in EC and PDAC, and Fusobacterium outgrowth in EC, gastric cancer (GC) and CRC. Parvimonas and Streptococcus showed a concordant ascending trend in GC and CRC. Prevotella was overrepresented in EC and GC. This synthesis delineates a qualitative landscape of gut microbiota imbalances associated with various DSCs, highlighting the potential for these microbial shifts to serve as markers for early detection and targeted therapy. Multiomics integration and prospective cohort studies should be prioritized to accelerate clinical translation.

PMID:42131199 | PMC:PMC13161176 | DOI:10.3389/fmicb.2026.1782471

  •  

Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs

arXiv:2603.04099v1 Announce Type: cross Abstract: Multi-Layer Perceptron (MLP) models are the foundation of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength and limit the application of these models. In this article, we develop a two-stage abstraction and refinement (ABS-REF) view for modular feature extraction in point cloud processing. This view elucidates that whereas the early models focused on ABS stages, the more recent techniques devise sophisticated REF stages to attain performance advantages. Then, we propose a High-dimensional Positional Encoding (HPE) module to explicitly utilize intrinsic positional information, extending the ``positional encoding'' concept from Transformer literature. HPE can be readily deployed in MLP-based architectures and is compatible with transformer-based methods. Within our ABS-REF view, we rethink local aggregation in MLP-based methods and propose replacing time-consuming local MLP operations, which are used to capture local relationships among neighbors. Instead, we use non-local MLPs for efficient non-local information updates, combined with the proposed HPE for effective local information representation. We leverage our modules to develop HPENets, a suite of MLP networks that follow the ABS-REF paradigm, incorporating a scalable HPE-based REF stage. Extensive experiments on seven public datasets across four different tasks show that HPENets deliver a strong balance between efficiency and effectiveness. Notably, HPENet surpasses PointNeXt, a strong MLP-based counterpart, by 1.1% mAcc, 4.0% mIoU, 1.8% mIoU, and 0.2% Cls. mIoU, with only 50.0%, 21.5%, 23.1%, 44.4% of FLOPs on ScanObjectNN, S3DIS, ScanNet, and ShapeNetPart, respectively. Source code is available at https://github.com/zouyanmei/HPENet_v2.git.
  •  
❌